Abstract: Insufficient sample size is a primary obstacle to precise inference, especially when observations pair text covariates with scalar outcomes. Unlike structured co-variates with prespecified coordinates and interpretable perturbations, text lacks a canonical low-dimensional representation, and routine edits need not preserve outcome-relevant semantics. Modern large language models can generate source-conditioned texts, creating a new possibility for addressing this difficulty. Yet distribution shift and within-source dependence can make naive use biased and over-confident. We introduce Model-assisted Semantic Inference (MaSI), which uses generated texts to construct a more precise confidence interval for the parameter of interest while controlling the inferential effects of LLM-induced distribution shift under the stated calibration conditions. Specifically, the first set is generated under a scheme that preserves source content while varying expression style together with real labeled data, it is used to learn outcome-relevant content coordinates and complementary style coordinates. A separate evaluation set constructs a calibrated auxiliary statistic in those coordinates. Density-ratio calibration ad-dresses distribution shift, real-label residual correction preserves the population target, and uncertainty is estimated across original sources. Under explicit representation, calibration, and covariance-control conditions, MaSI provides consistent variance estimation and asymptotically valid confidence intervals. Simulations and empirical-data experiments show shorter intervals with near-nominal coverage in the primary settings, establishing a principled route to more precise inference from scarce text-outcome pairs.
个人简介:张新雨,中国科学院数学与系统科学研究院研究员。长期从事统计和计量经济学理论与应用方面的研究工作,与合作者解决了模型平均研究中的多个难题,并将模型平均与迁移学习、随机森林、库存管理等融合提出了新方法,同时将所提出的预测方法应用于实际问题为相关部门的决策提供了参考依据。先后主持国家自然科学基金委青年基金A类及其延续项目等,曾获中国青年科技奖。