publications
2026
- Distribution-Free Fair Federated Learning with Small SamplesQichuan Yin, Zexian Wang, Junzhou Huang, Huaxiu Yao, and Linjun ZhangIn ASA Statistical Learning and Data Science (SLDS), 2026
As federated learning gains increasing importance in real-world applications due to its capacity for decentralized data training, addressing fairness concerns across demographic groups becomes critically important. However, most existing machine learning algorithms for ensuring fairness are designed for centralized data environments and generally require large-sample and distributional assumptions, underscoring the urgent need for fairness techniques adapted for decentralized and heterogeneous systems with finite-sample and distribution-free guarantees. To address this issue, this paper introduces FedFaiREE, a post-processing algorithm developed specifically for distribution-free fair learning in decentralized settings with small samples. Our approach accounts for unique challenges in decentralized environments, such as client heterogeneity, communication costs, and small sample sizes. We provide rigorous theoretical guarantees for both fairness and accuracy, and our experimental results further provide robust empirical validation for our proposed method.
@inproceedings{yin2026fedfairee, title = {Distribution-Free Fair Federated Learning with Small Samples}, author = {Yin, Qichuan and Wang, Zexian and Huang, Junzhou and Yao, Huaxiu and Zhang, Linjun}, booktitle = {ASA Statistical Learning and Data Science (SLDS)}, year = {2026}, status = {journal}, } - Differentially Private Model MergingQichuan Yin, Manzil Zaheer, and Tian LiIn Advances in Neural Information Processing Systems (NeurIPS), 2026
In machine learning, privacy requirements at inference or deployment time often evolve due to changing policies, regulations, or user preferences. In this work, we aim to construct a magnitude of models to satisfy any target differential privacy (DP) requirement without additional training, given a set of existing models trained on the same dataset with different privacy/utility tradeoffs. We propose two post-processing techniques, namely random selection and linear combination, to generate final private models satisfying any target privacy parameter. We provide privacy accounting of these approaches from the lens of Rényi DP and privacy loss distributions on general problems, as well as on private mean estimation, where we precisely characterize the privacy/utility tradeoffs and compare the two mechanisms. Empirically, we demonstrate the effectiveness of our approaches and validate our analyses on several models and both synthetic and real-world datasets.
@inproceedings{yin2026dpmerging, title = {Differentially Private Model Merging}, author = {Yin, Qichuan and Zaheer, Manzil and Li, Tian}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS)}, year = {2026}, status = {conference}, } - Overcoming the Incentive Collapse ParadoxQichuan Yin*, Ziwei Su*, and Shuangning LiIn International Conference on Machine Learning (ICML), 2026
AI-assisted task delegation is increasingly common, yet human effort in such systems is costly and typically unobserved. Recent work by Bastani and Cachon (2025); Sambasivan et al. (2021) shows that accuracy-based payment schemes suffer from incentive collapse: as AI accuracy improves, sustaining positive human effort requires unbounded payments. We study this problem in a budget-constrained principal-agent framework with strategic human agents whose output accuracy depends on unobserved effort. We propose a sentinel-auditing payment mechanism that enforces a strictly positive and controllable level of human effort at finite cost, independent of AI accuracy. Building on this incentive-robust foundation, we develop an incentive-aware active statistical inference framework that jointly optimizes (i) the auditing rate and (ii) active sampling and budget allocation across tasks of varying difficulty to minimize the final statistical loss under a single budget. Experiments demonstrate improved cost-error tradeoffs relative to standard active learning and auditing-only baselines.
@inproceedings{yin2026incentive, title = {Overcoming the Incentive Collapse Paradox}, author = {Yin, Qichuan and Su, Ziwei and Li, Shuangning}, booktitle = {International Conference on Machine Learning (ICML)}, year = {2026}, status = {conference}, } - Synthetic but Not Infinite: How Much LLM-Generated Data to Use in Market ResearchQichuan Yin and Linwei XinWorking paper, 2026Major revision at Management Science
Large Language Models (LLMs) have transformed data generation by enabling scalable synthetic data augmentation. This paper studies a fundamental question in market research: how much synthetic data should be used? In the existing literature, the proportion of LLM-generated synthetic data is usually treated as a hyperparameter, tuned by ad hoc experimentation. In addition, a more important yet unexplored question is how the optimal amount of synthetic data should vary across different population segments, e.g., between younger and older respondents. In this work, we address these challenges by deriving a closed-form expression for this highly heterogeneous hyperparameter vector that governs the proportion of synthetic data. This prescriptive characterization arises from our analysis of a logit choice model that integrates real and synthetic data, under which we establish a finite-sample complexity bound for the estimation error of the resulting hybrid estimator. Specifically, our expression for the heterogeneous proportion of synthetic data corresponds to the one that minimizes this upper bound. We further demonstrate the strong empirical performance of our closed-form expression using both synthetic data and a real-world vaccine preference dataset. Overall, our results suggest that the closed-form expression serves as a practical and reliable empirical rule of thumb for determining how much LLM-generated synthetic data to incorporate into market research studies.
@article{yin2026synthetic, title = {Synthetic but Not Infinite: How Much {LLM}-Generated Data to Use in Market Research}, author = {Yin, Qichuan and Xin, Linwei}, journal = {Working paper}, year = {2026}, note = {Major revision at Management Science}, status = {working}, }
- Token Design for Favor Trading with Dynamic MembershipXing Hu, Zhixi Wan, and Qichuan YinManagement ScienceForthcoming. Authors in alphabetical order.
The existing literature on trading-favor systems, or scrip systems, generally assumes a fixed set of participants. We study the design of a trading-favor system that allows participants to join and exit. We examine how the dynamics of participation create challenges to participants’ strategies and, consequently, the design of the system. We introduce a token exchange mechanism that can be implemented by simple smart contracts. Such mechanism can regulate token distribution, evaluate token value and fairly adjust the total number of tokens in circulation. We show that the mechanism can induce participants to adopt an always-trade strategy in equilibrium, resulting in robust and highly efficient resource sharing in the decentralized system.
@article{hu_token, title = {Token Design for Favor Trading with Dynamic Membership}, author = {Hu, Xing and Wan, Zhixi and Yin, Qichuan}, journal = {Management Science}, note = {Forthcoming. Authors in alphabetical order.}, status = {journal}, }