Research
High-dimensional statistics and implicit regularization
High-dimensional procedures regularize not only through explicit penalties, but also through sampling geometry, spectral truncation, and finite-time optimization. I study when these mechanisms attenuate recoverable signal, how to quantify the corresponding variance price, and when deattenuation or anti-shrinkage improves prediction.
Research Highlight · arXiv preprint
Beyond Negative-Ridge Endpoints: Mixed-Sign Spectral Regularization via Negative-Shifted Gradient Descent
Finite-time negative-shifted gradient descent crosses a spectral barrier that constrains every stable negative-ridge endpoint, enabling head anti-shrinkage and lower-spectrum control in one path.
Why high-dimensional prediction behaves differently
When the number of parameters exceeds the sample size, fitting the training data does not by itself determine test performance. Prediction depends on the sample covariance spectrum: a small leading head can carry the recoverable signal, while many weak directions determine how the fitted inverse behaves.
Notation. Let \(X\in\mathbb R^{n\times p}\) and \(y=X\beta^\star+\varepsilon\) denote the observed design and response, and let \(\Sigma=\mathbb E(xx^\top)\) be the population feature covariance. In the population eigenbasis, split the observed design as \(X=[X_H\;X_T]\), where \(H\) is the signal-bearing head and \(T\) is the weak tail. We use \(\lambda\) for population covariance eigenvalues—\(\lambda_h\) for the head and \(\lambda_T\) for the flat tail—and \(\mu\) for eigenvalues of the observed Gram matrix.
\[K=\frac{XX^\top}{n}=K_H+K_T,\qquad K_H=\frac{X_HX_H^\top}{n},\quad K_T=\frac{X_TX_T^\top}{n}\approx aI_n,\quad a=\frac{\operatorname{tr}(\Sigma_T)}{n}.\]
A high-effective-rank tail can therefore be weak coordinate by coordinate yet create a substantial sample-space floor. The head is learned through \(K_H+aI_n\), so ridgeless regression experiences implicit positive shrinkage on signal-bearing directions. Negative ridge has the corrective sign, but its stable endpoint is limited by a pole and a tail-heavy filter shape.
Start with the common-spike model. If a flat tail has level \(\lambda_T\) and aspect ratio \(\gamma_T=d_T/n\), its many weak directions create the effective floor
The shift \(\nu_\star\) therefore lies beyond the admissible stable-endpoint range \(0\leq\nu<\widehat\mu_{\min}^{+}\). This is the Marchenko–Pastur pole barrier.
- 1Weak tail directions create an implicit ridge floor.
Collectively, the high-effective-rank tail shrinks the signal-bearing head as if a positive ridge of size \(a\) were present.
- 2Stable negative ridge meets a spectral barrier.
The endpoint filter \(A_\nu(\mu)=\mu/(\mu-\nu)\) has a pole and amplifies smaller stable eigenvalues most—the wrong shape for head correction with tail control.
- 3Finite time removes the endpoint restriction.
Stopping before convergence makes the singularity removable and lets the displacement \(f_{\nu,t}(\mu)-1\) cross sign once, on a data-adaptive leading spectral prefix.
Schematic normalization: orange is the admissible endpoint envelope as \(\nu\uparrow\widehat\mu_{\min}^{+}\); blue uses the selected \((\nu,t)\).
\[f_{\nu,t}(\mu)=\frac{\mu}{\mu-\nu}\left\{1-e^{-t(\mu-\nu)}\right\},\qquad f_{\nu,t}(\nu)=\nu t.\]
Networks and relational data
Many modern datasets describe relationships rather than independent observations. My work develops latent-space, covariate-assisted, dynamic, and spatial models for networks, with an emphasis on uncertainty quantification, robustness, and interpretable structure.
Scalable Bayesian inference
I develop variational and approximate Bayesian methods for models where conventional posterior computation is too costly. The goal is to pair computational scalability with rigorous statistical guarantees.
Multivariate and nonparametric Bayes
I am also interested in shrinkage priors, graphical models, dimension reduction, and Bayesian models that adapt to complex dependence without imposing unnecessarily rigid parametric structure.