About
I am a Staff Research Scientist at Meta AI, where I lead data curation and pretraining
for the open-source on-device MobileLLM and MobileMoE model series.
My research focuses on data-centric training of foundation models:
deciding what data a model should see, in what proportion, and in what order,
so that a fixed compute budget buys the most capability.
Research interests: data curation and mixture optimization,
data-influence estimation, synthetic data generation, large-scale pretraining and
midtraining, on-device and streaming models, multilingual and low-resource NLP.
Open-Source Models
- MobileLLM-R1, on-device reasoning model (data curation and pretraining lead)
- MobileLLM-Pro, 1B on-device LLM; the language backbone selected for Meta’s on-device VLM
- MobileMoE, sparse mixture-of-experts variant for on-device deployment
Selected Publications
-
MobileLLM-R1: Exploring the Limits of Sub-Billion Language Model Reasoners with Open Training Recipes
C. Zhao*, E. Chang*, Z. Liu*, et al.
ICLR 2026 Data curation lead
-
MobileMoE: Scaling On-Device Mixture of Experts
Y. Chen, H. Huang, E. Chang, et al.
arXiv, 2026
-
Automixer: Checkpoint Artifacts as Automatic Data Mixers
E. Chang, Y. Li, P. Huber, V. Vogeti, D. Kant, Y. Shi, V. Chandra
ACL 2025 First author
-
MobileLLM-Pro Technical Report
P. Huber*, E. Chang*, W. Wen*, I. Fedorov*, et al.
arXiv, 2025 Data curation lead
-
Scaling Parameter-Constrained Language Models with Quality Data
E. Chang, M. Paltenghi, Y. Li, et al.
EMNLP 2024 First author
-
DART: A Lightweight Quality-Suggestive Data-to-Text Annotation Tool
E. Chang, J. Caplinger, A. Marin, X. Shen, V. Demberg
COLING 2020 · Best System Paper Award First author
* Equal contribution
Full list on Google Scholar.