Seminario de Estudiantes de Doctorado en Estadística

El propósito de estos seminarios es conocer los proyectos de investigación en los que han participado los y las estudiantes del programa de Doctorado en Estadística en la modalidad de ponencia. Se extiende la invitación a participar a toda la comunidad UC.


2026-06-02
16:10hrs.
Bladimir Morales Torrez. Pontificia Universidad Católica de Chile
Analysis of Massive Spatial Datasets: A Block-Pairwise Likelihood Approach
Sala 3, Facultad de Matemáticas
Abstract:
Likelihood-based inference for Gaussian random fields becomes computationally prohibitive for massive datasets because evaluating the full Gaussian likelihood requires dense covariance matrix factorizations. We propose a block–pairwise composite likelihood framework for scalable inference in large spatial Gaussian random fields. The method partitions the spatial domain into disjoint blocks and constructs localized pairwise composite likelihoods using within-block nearest-neighbor marginal or conditional contributions. This strategy preserves essential short-range spatial dependence while substantially reducing memory usage and computational cost.
Under standard regularity conditions, the resulting estimator is consistent and asymptotically normal within the composite likelihood framework, and uncertainty quantification can be conducted through the Godambe information matrix. Simulation studies show that the proposed method achieves a favorable balance between statistical efficiency and computational scalability, is robust to the choice of block partition strategy, and performs well across different numbers of neighbors and blocks. In particular, simple block constructions combined with a moderate number of nearest neighbors provide competitive estimation accuracy at substantially lower computational cost. Applications to large, regular, and irregular spatial datasets further illustrate that the proposed approach delivers reliable inference and prediction while remaining considerably faster than global pairwise methods and competitive Vecchia-type alternatives.
2026-05-26
16:10hrs.
Ingrid Guevara Romero. Pontificia Universidad Católica de Chile
Bayesian Multiple Comparisons with Applications to Alzheimer's data
Sala 3, Facultad de Matemáticas
Abstract:
Multiple-comparison methods are widely used in medicine, engineering, and the social sciences to determine whether groups differ. Applications range from comparing medical treatments and biological populations to evaluating environmental or technological interventions. Classical approaches, such as Tukey or Dunnett tests, remain standard tools for these problems, but they often rely on restrictive assumptions, including Gaussianity, equal variances, or specific data types. In practice, however, real-world data are frequently heterogeneous, asymmetric, multimodal, or measured on nonstandard scales, limiting the applicability of traditional methods. 
 
This project develops a new Bayesian framework for multiple comparisons that is broadly applicable across scientific disciplines. The proposed methodology formulates the problem as one of identifying which groups share the same underlying data-generating mechanism and which differ. Within this framework, competing hypotheses are represented as probabilistic models, allowing uncertainty to be naturally incorporated into the analysis while simultaneously controlling false discoveries through a novel prior distribution that penalizes overly complex group structures, encouraging parsimonious and interpretable solutions while preserving sensitivity to meaningful differences.  
 
Unlike many existing approaches, the method requires only independence across samples and can adapt automatically to complex distributional features. 
 
The methodology is motivated by applications in areas where reliable group comparisons are essential, including biomedical research and disease detection. In particular, the framework is illustrated through an application to Alzheimer's disease, where detecting subtle population differences may improve diagnostic strategies. 
2026-05-12
16:10hrs.
Bryan Andres Tobar Torres. Pontificia Universidad Católica de Chile
Spatial Segmentation Using Product Partition Models
Sala 3, Facultad de Matemáticas
Abstract:
The study applies the methodology proposed by Teixeira and Assunção in A Generative Spatial Clustering Model for Random Data Through Spanning Trees to analyze the support rate for Candidate X across districts in Santiago, Chile. The study explores the spatial structure of voting behavior by identifying groups of districts with similar support patterns while accounting for geographical connectivity.
2026-05-05
16:10hrs.
Nataly Martinez Riascos. Pontificia Universidad Católica de Chile
Introduction to repairable systems models
Sala 3, Facultad de Matemáticas
Abstract:
This presentation provides an introductory overview of repairable systems models within a point process framework. We begin with ideas from classical survival analysis as a conceptual entry point, then move to recurrent failures, which are described using counting processes, intensity functions, and cumulative intensity functions.
 
Starting from the Power Law Process (PLP) under minimal repair assumptions, we discuss how more general formulations capture different repair mechanisms, including perfect and imperfect repairs, as well as extensions that account for repair effects or multiple failure causes. This evolution leads to flexible structures such as the Generalized Power Law Process (GPLP), which encompasses several classical models as particular cases.
 
The objective is to provide a clear introduction to the modeling of repairable systems, emphasizing their interpretation, evolution, and relevance in reliability analysis.
2026-04-28
16:10hrs.
José Luis Pérez Zamorano. Pontificia Universidad Católica de Chile
Theoretical Principles of Asymmetric Item Response Theory Models
Sala 3, Facultad de Matemáticas
Abstract:
This seminar explores extensions of classical Item Response Theory (IRT) models that allow for asymmetric item characteristic curves (ICCs). Standard models, such as the normal ogive and logistic models, impose symmetry around the item difficulty parameter, which may be unrealistic in many testing contexts. To address this limitation, we examine models based on skewed link functions, particularly those derived from the skew-normal distribution.
We begin by reviewing the classical probit framework and its interpretation in terms of latent variable models. We then introduce skew-probit formulations, where asymmetry is incorporated through a shape parameter that modifies the probability of a correct response.
In addition, we discuss alternative approaches to modeling asymmetry, including the residual heteroscedastic model proposed by Molenaar (2015) and the Logistic Positive Exponent (LPE) model introduced by Samejima (2000). These models offer different perspectives on asymmetry, either through variability in the latent response process or through flexible transformations of the link function.
The seminar emphasizes both theoretical properties and practical implications, showing how asymmetric models provide a more flexible and realistic framework for modeling item responses in educational and psychological measurement.
2026-04-21
16:10hrs.
Ignacio Betancourt Peters. Pontificia Universidad Católica de Chile
A Mixed Regression and Item Response Models for Bounded Count Data
Sala 3, Facultad de Matemáticas
Abstract:
Count data are ubiquitous across disciplines such as medicine, economics, and psychology, yet their statistical modeling has historically received less attention than continuous data. While generalized linear models (GLMs) and their mixed extensions (GLMMs) provide a common framework for analyzing such data, they typically rely on distributions with unbounded support (e.g., Poisson or Negative Binomial), even in contexts where counts are naturally bounded. Although some models account for this bounded nature, few adequately address issues such as varying dispersion and the presence of extreme observations. Recent developments, such as the Beta-2-Binomial (B2B) regression model, offer greater flexibility and robustness through hierarchical structures. In parallel, Item Response Theory (IRT) models, widely used in psychometrics, often involve bounded count data but inherit similar limitations from standard modeling approaches. Motivated by these gaps, this research proposes the development of a mixed regression model for bounded count data based on the B2B distribution, along with its extension to count-based IRT models. The proposed approach aims to better reflect the intrinsic characteristics of the data while providing robustness to dispersion and outliers, and will be studied under both frequentist and Bayesian frameworks through simulations and implementations in R.
2026-03-31
16:10hrs.
Bryan Andres Tobar Torres. Pontificia Universidad Católica de Chile
New Contributions to Change Point Detection Based on Product Partition Models
Sala 3, Facultad de Matemáticas
Abstract:
Understanding how structural changes propagate across financial markets is essential for characterizing contagion and anticipating periods of instability. A standard tool in Bayesian statistics for detecting changes is the product partition model (PPM) and its multivariate extensions, which are effective at detecting abrupt shifts in distributional features, such as mean or variance. These models often rely on Gaussian assumptions, limiting their ability to detect changes only in the presence of possible symmetry or heavy tails, and implicitly assume that structural breaks occur synchronously across series. These simplifying assumptions limit their ability to capture other empirical features of financial returns, such as skewness, asymmetric relations and the gradual asynchronous transmission of shocks across interconnected markets.
 
The primary objective of this thesis is to extend the theory of multivariate change point detection within the PPM framework by introducing distributional and dependence flexibility. First, we employ the Generalized Hyperbolic (GH) distribution as the block-specific predictive distributions, enabling a representation of data dynamics that accommodates skewness and heavy-tailed behavior. Second, to capture directional contagion effects across markets, we introduce an asymmetric cross-series dependence structure using an influence matrix. Third, we extend the dependence framework to account for the delayed transmission of structural breaks by modeling the logit-transformed change point probabilities using a separable covariance structure with an explicit time-delay component. 
 
Beyond detection, the fourth objective is to develop a fully probabilistic framework for forecasting future structural changes. Building on recent advances in Bayesian online change point prediction, future change points are inferred directly from the evolving posterior distribution of the model’s own change point probabilities. This approach naturally accounts for uncertainty in both the current regime and the timing of future breaks, enabling coherent probabilistic forecasts of structural change in interconnected financial time series.
2026-03-24
16:10hrs.
Fabián Gómez Solar. Pontificia Universidad Católica de Chile
Effective sample size in random fields and sampling design under fixed domain
Sala 3, Facultad de Matemáticas
Abstract:
At present, when performing multivariate statistical analyses at a given point in time, it is common to encounter difficulties due to the high dimensionality of the data sets involved. Spatial statistics plays a fundamental role in these contexts, as it models correlation as a function of the spatial separation among sample observations. However, a major limitation of such analyses arises from the large volumes of available data. For instance, information contained in images may lead to the analysis of massive data sets, posing substantial computational challenges. These considerations motivate the idea that it is not necessary to rely on the entire sample in order to represent the effective information it contains. In this context, the effective sample size (ESS) plays a fundamental role, as it allows the amount of independent information present in the sample to be quantified. Its implementation in spatial settings has increased considerably over the last decades. Nevertheless, in the existing literature, ESS calculations typically require the full sample information, which in many cases leads to severe computational difficulties, mainly due to the inversion of high-dimensional matrices. The objective of this research consists of two main components. The first is to redefine the calculation of the ESS, making it independent of the observed sample and instead relying on the assumed correlation model and the spatial domain to obtain a new approximation of this quantity. The second objective is to define a strategy for selecting the sampling locations that will compose the sample.
2025-11-20
16:10hrs.
José Luis Pérez Zamorano. Pontificia Universidad Católica de Chile
Theoretical Principles and Recent Developments in the Identifiability of Fixed-Effects IRT Models
Sala 3, Facultad de Matemáticas
Abstract:
The identifiability of a model constitutes a fundamental aspect of its specification. According to Fisher (1922), specification is one of the three essential problems that must be addressed when working with statistical models, along with the problems of estimation and distribution. In the field of Psychometrics, Item Response Theory (IRT) models have become one of the most widely used tools for the analysis of assessment and measurement data. Among these, the Rasch model has been one of the most influential, establishing the theoretical and methodological foundations that have guided much of the subsequent development in this area.
This dissertation examines the problem of identifiability in various IRT models, particularly in the one-, two-, and three-parameter logistic models (1PL, 2PL, and 3PL), as well as in an extension of the one-parameter model that incorporates a guessing parameter (1PL-G). All of these models are analyzed under a fixed-effects specification. Furthermore, a recent research advancement is presented concerning the identifiability analysis of the one-parameter logistic positive exponential model (1PL-LPE) with fixed effects, developed by González et al. (2025), which provides a new perspective on the parametric structure and the conditions for the identification of this type of models.
2025-10-20
14:50hrs.
Lucas Osses Camblor. Pontificia Universidad Católica de Chile
Online Estimation Methods for Irregular Autoregressive Models (iAR)
Sala 3, Facultad de Matemáticas
Abstract:
Over the last few decades, due to significant technological growth, it has become increasingly common for a collection of temporal data to accumulate in vast amounts rapidly. This scenario presents an opportunity to extract valuable information by estimating increasingly precise models. However, at the same time, it imposes the challenge of continuously updating the models as new data becomes available. Currently available methods for addressing this problem, known as online learning methods, utilize current parameter estimations and novel data to update the estimators. These approaches avoid using the complete raw data and speed up the computations. In this work, we consider three online learning algorithms for parameter estimation in the context of time series models. In particular, the methods implemented are gradient descent, Newton-step, and Kalman filter recursions. These algorithms are applied to the recently developed irregularly observed autoregressive (iAR) model. The estimation accuracy of the proposed methods is evaluated through Monte Carlo experiments. The results obtained show that the proposed online estimation methods enable precise estimation of the parameters that generate the data, both for regularly and irregularly observed time series. These online approaches are numerically efficient, allowing substantial computational time savings. Moreover, we demonstrate that the proposed methods can quickly adapt the parameter estimates when the time series behavior changes, unlike batch estimation methods.
2025-10-06
14:50hrs.
Fabián Gómez Solar. Pontificia Universidad Católica de Chile
Survival Analysis Driven Prediction of Construction Phase Transitions with Fully Censored Data
Sala 3, Facultad de Matemáticas
Abstract:
Accurately predicting the duration of each construction phase is critical for project management, yet exact transition times are unknown and observed only within intervals, i.e., censored data, defined by periodic site inspections or georeferenced imagery. This paper estimates phase durations under fully censored conditions by fitting semiparametric Cox proportional hazards models adapted for interval-censored data and parametric accelerated failure time models with six candidate survival distributions. AIC guides covariate selection, and we compared the baseline hazard functions across parametric and semiparametric frameworks. We also propose a modification of the Cox–Snell residual plots to handle fully censored observations and assess goodness of fit. The resulting models are implemented within a predictive tool that provides phase-completion forecasts, which could integrate into decision-support systems for contractors, suppliers, and investors in Chile’s construction industry.
 
2025-09-02
16:10hrs.
Mauricio Alejandro Toro Cea. Pontificia Universidad Católica de Chile
Decomposing the Average Treatment Effect: Distributional Foundations and Partial Identification
Sala 3, Facultad de Matemáticas
Abstract:
This paper decomposes the average treatment effect (ATE) into the average gain and loss from treatment under a nonparametric partial identification framework. We introduce the treatment effect heterogeneity coefficient, DC, along with its sharp bounds. We also define the Cumulative Treatment Effect (CTE) and characterize its relationship to both ATE and the Quantile Treatment Effect (QTE). To address potential violations of ignorability in observational studies, we propose the 4-ε and 4-η assumptions, which allow bounded deviations in conditional potential outcomes distributions. These assumptions enable a sensitivity analysis and define a sensitivity index—the maximum level of confounding under which the sign of the ATE remains identified. An application to Card and Krueger (1994)’s minimum wage study reveals that their conclusions are highly sensitive to modest assumption relaxations.
2025-06-17
16:10hrs.
Ingrid Guevara Romero. Pontificia Universidad Católica de Chile
Bayesian regression and model selection for planar shapes
Sala 3, Facultad de Matemáticas
Abstract:
The interest in analyzing, comparing, and studying shapes is prevalent in various disciplines, such as medicine, biology, and chemistry. Applications range from disease diagnosis and the study of biological variation to facial recognition. Motivated by these broad uses, we present a methodology to model a set of landmarks representing two-dimensional shapes. Our work presents a Bayesian regression framework that incorporates covariates into the modal shape of a complex Watson distribution. Furthermore, we incorporate a spike-and-slab prior distribution for model selection to assess whether relevant factors contribute to shape differences. We demonstrate the utility of the proposed method through an application involving shapes of the midline section of ape skulls, aiming to determine whether cranial shape differences are associated with sex and/or species.
2025-06-10
16:10hrs.
Francisco Antonio Segovia Godoy. Pontificia Universidad Católica de Chile
A Flexible continuous and binary Bayesian model selection method
Sala 3, Facultad de Matemáticas
Abstract:
Regression analysis aims to explore the relationship between a response variable and predictors. A key aspect of regression analysis is variable selection, which enables researchers to identify the most relevant predictors while adhering to the principle of parsimony. The standard frequentist strategy is to explore the model space using, for instance, a Stepwise strategy. Alternatively, a popular Bayesian strategy is the spike-and-slab methodology, which assigns a prior to the predictor coefficients conditional on a latent binary vector that indicates which variables are relevant. This strategy incorporates a prior on the binary vector to penalize overly complex models. This work presents a general Bayesian nonparametric approach to model selection for continuous and binary regression models, employing the spike-and-slab strategy combined with a data augmentation technique. The posterior distribution admits closed-form expressions. Guidelines for the inference implementation are detailed with a supporting Monte Carlo simulation study.
2025-06-03
16:10hrs.
Cristian Capetillo Constela. Pontificia Universidad Católica de Chile
Introducción a los métodos de Inferencia Variacional: una alternativa a MCMC
Sala 3, Facultad de Matemáticas
Abstract:
La inferencia Bayesiana se basa únicamente en la distribución a posteriori ya sea de los parámetros o de observaciones "futuras". Así, una vez calculada esta distribución, la inferencia se realiza mediante funcionales de ella (media, varianza, cuantiles, etc.). Sin embargo, así como aparenta ser una gran ventaja, también conlleva una importante dificultad; y es que en la mayoría de modelos Bayesianos, la distribución a posteriori no está disponible en forma cerrada. Más aún, si lo estuviese, podrían existir funcionales difíciles de calcular. Esta razón ha invitado a vertir innumerables esfuerzos en desarrollar métodos computacionales que permitan aproximar la distribución a posteriori.
 
Si bien tales esfuerzos se han concentrado mayoritariamente en los métodos MCMC (Markov Chain Monte Carlo), existen muchas otras estrategias de aproximar la distribución a posteriori. En esta discusión nos enfocaremos en el artículo "Variational Inference: A review for Statisticians" escrito por Blei, Kucukelbir y McAuliffe en el año 2017, el cuál trata, como su nombre lo indica, de una revisión de los métodos de Inferencia Variacional (VI, por sus siglas en inglés), una alternativa para el cálculo aproximado de la distribución a posteriori. El artículo expone las ideas fundamentales detrás de estos métodos, realizando una aplicación ilustrativa en un modelo de Mezclas Gaussianas, así como aplicaciones realizadas en la literatura reciente, extensiones, resultados teóricos y algunos problemas abiertos sobre estos métodos.
 
La invitación es a discutir y reflexionar sobre los métodos de Inferencia Variacional, y en general sobre aquellos enfoques que buscan aproximar distribuciones, tomando este artículo como punto de referencia.
2025-05-27
16:10hrs.
Jesus Enrique Achire Quispe. Pontificia Universidad Católica de Chile
Objective Bayesian inference for the Exponential-Logarithmic distribution
Sala 3, Facultad de Matemáticas
Abstract:
This study explores an objective Bayesian inference approach for parameter estimation in the Exponential-Logarithmic (EL) distribution. Initially, we establish the necessary and sufficient conditions under which improper priors yield proper posterior distributions for the EL distribution. Additionally, we provide sufficient conditions to ensure the finiteness of posterior moments. These theoretical results are specifically applied to Jeffreys' prior, the maximal data information prior, and reference priors, demonstrating that such improper priors indeed generate proper posterior distributions. To assess the impact of the proposed priors on posterior estimation, we employ Markov Chain Monte Carlo methods and conduct extensive numerical simulations, comparing Bayesian estimators with the maximum likelihood estimators in terms of bias, mean squared error, and coverage probability.
2025-05-20
16:10hrs.
Mauricio Alejandro Toro Cea. Pontificia Universidad Católica de Chile
Causal inference and logically possible worlds: Partial Identification and statistical methods for numerical outcomes in finite populations
Sala 3, Facultad de Matemáticas
Abstract:
This work proposes a new methodological framework for causal inference. The perspective is based on the assumption of a finite sample space, where the data correspond to the population of interest or to a subset of statistical units from a larger finite population, without assuming the existence of an underlying data generation process. A formalization of Partial Identification is proposed, distinct from Manski’s approach, based on logically possible worlds. Additionally, causal inference methods are proposed on the probability space of the set of logically possible counterfactual worlds. The proposed statistical methods, based on a logicist interpretation of probability, are valid for both dichotomous and numerical outcomes, without the need to assume continuity in the outcomes. The methodology is illustrated by analyzing data on the impact of minimum wage increases on unemployment rates.
2025-05-13
16:10hrs.
Jose Alejandro Ordoñez Cuastumal. Pontificia Universidad Católica de Chile
Penalized complexity priors for the skewness parameter of power links
Sala 3, Facultad de Matemáticas
Abstract:

The choice of a prior distribution is a key aspect of the Bayesian method. However, in many cases, such as the family of power links, this is not trivial. In this article, we introduce a penalized complexity prior (PC prior) of the skewness parameter for this family, which is useful for dealing with imbalanced data. We derive a general expression for this density and show its usefulness for some particular cases such as the power logit and the power probit links. A simulation study and a real data application are used to assess the efficiency of the introduced densities in comparison with the Gaussian and uniform priors. Results show improvement in point and credible interval estimation for the considered models when using the PC prior in comparison to other well-known standard priors.

2025-05-06
16:10hrs.
Nixon Andrés Jerez Lillo. Pontificia Universidad Católica de Chile
Beyond the Power Law: Estimation, Goodness-of-Fit, and a Semiparametric Extension in Complex Networks
Sala 1, edificio Rolando Chuaqui
Abstract:

Scale-free networks play a fundamental role in the study of complex networks and various applied fields due to their ability to model a wide range of real-world systems. A key characteristic of these networks is their degree distribution, which often follows a power-law distribution, where the probability mass function is proportional to $x^{-\alpha}$, with $\alpha$ typically ranging between $2 < \alpha < 3$. In this talk, we introduce Bayesian inference methods to obtain more accurate estimates than those obtained using traditional methods, which often yield biased estimates, and precise credible intervals. Through a simulation study, we demonstrate that our approach provides nearly unbiased estimates for the scaling parameter, enhancing the reliability of inferences. We also evaluate new goodness-of-fit tests to improve the effectiveness of the Kolmogorov-Smirnov test, commonly used for this purpose. Our findings show that the Watson test offers superior power while maintaining a controlled type I error rate, enabling us to better determine whether data adheres to a power-law distribution. Finally, we propose a piecewise extension of this model to provide greater flexibility, evaluating the estimation and its goodness-of-fit features as well. In the complex networks field, this extension allows us to model the full degree distribution, instead of just focusing on the tail, as is commonly done. We demonstrate the utility of these novel methods through applications to two real-world datasets, showcasing their practical relevance and potential to advance the analysis of power-law behavior.

2025-04-22
16:10hrs.
Luz Marina Ramos Quispe. Pontificia Universidad Católica de Chile
Scale mixture of a multivariate normal distribution with a Birnbaum-Saunders mixing distribution
Sala 3, Facultad de Matemáticas
Abstract:

In the search for multivariate distributions that provide greater flexibility in modeling data characterized by high levels of skewness, kurtosis, and the presence of outliers, new families of multivariate distributions have emerged, among which multivariate normal mixture distributions stand out. In this context, we introduce a multivariate normal mixture distribution based on the Birnbaum-Saunders distribution and examine some of its key properties. To estimate the parameters of this normal scale mixture distribution, we propose a maximum likelihood approach implemented via the EM algorithm. To support inferential analyses, we derive the Fisher information matrix. Additionally, we formulate a linear hypothesis on the parameter vector of interest and evaluate it using the likelihood ratio, Wald, score, and gradient statistics. Finally, we illustrate the application of the proposed methodology to real datasets, complementing the analysis with a simulation study to assess its performance.