About
Sheng Zha 查晟
I lead pretraining research for Amazon Nova, focusing on principled scaling, model architecture, optimization, and novel pretraining objectives.
I design model algorithms together with the systems needed to train them at scale. The goal is to improve model capability, reduce its cost, and make it useful to more people. I built my team from zero, beginning with distributed training and shared representations, and grew it into a group supporting foundation models used across AWS.
Before that, I helped shape the open-source AI ecosystem as VP and PMC Chair of Apache MXNet, co-authored the Gluon interface, founded GluonNLP, served on the ONNX Steering Committee, and co-founded the Python Data API Standards Consortium.
I want AI to expand what people can decide and create for themselves. I use that test when choosing research problems and leading teams.
Selected work
Foundation models · 2024–present
Amazon Nova and foundation model stack
Leading research on principled scaling, model architecture, optimization, and pretraining objectives.
Read the technical report ↗Team building · 2018–2024
Foundation models for AWS AI services
Built the team and models from zero, supporting Amazon Q, Titan, Lex, Comprehend, and Kendra.
Related platform ↗Systems · 2018–2023
Distributed training infrastructure
Contributed core technology for scalable, fault-resilient training infrastructure including SageMaker HyperPod.
Related platform ↗Open source · 2016–2023
Apache MXNet
VP and PMC Chair. Co-authored the Gluon API and led project maintenance, releases, and community engagement.
Project archive ↗Open source · 2018–2023
GluonNLP
Founded a deep-learning NLP toolkit that reproduced BERT with record-setting training speeds.
View source ↗Standards
ONNX Steering Committee
Project site ↗Standards
Python Data API Standards founding member
Read the standard ↗Machine learning systems · 2013–2015
Fraud detection platform
Designed scalable machine-learning and graph systems for fraud and abuse detection.
Selected publications
DEM: Distribution Edited Model for Training with Mixed Data Distributions
Dhananjay Ram, Aditya Rawal, Momchil Hardalov, Nikolaos Pappas, Sheng Zha · EMNLP 2024 Special Theme Paper Award
Differentially Private Bias-term Fine-tuning of Foundation Models
Zhiqi Bu, Yu-Xiang Wang, Sheng Zha, George Karypis · ICML 2024 · NeurIPS 2022 TSRML Workshop Outstanding Paper Award
HyTrel: Hypergraph-enhanced Tabular Data Representation Learning
Pei Chen, Soumajyoti Sarkar, Leonard Lausen, Balasubramaniam Srinivasan, Sheng Zha, Ruihong Huang, George Karypis · NeurIPS 2023 Spotlight
Zero Redundancy Distributed Learning with Differential Privacy
Zhiqi Bu, Justin Chiu, Ruixuan Liu, Sheng Zha, George Karypis · MLSys 2026
Meta-learning via Language Model In-context Tuning
Yanda Chen, Ruiqi Zhong, Sheng Zha, George Karypis, He He · ACL 2022
Automatic Clipping: Differentially Private Deep Learning Made Easier and Stronger
Zhiqi Bu, Yu-Xiang Wang, Sheng Zha, George Karypis · NeurIPS 2023
Differentially Private Optimization on Large Model at Small Cost
Zhiqi Bu, Yu-Xiang Wang, Sheng Zha, George Karypis · ICML 2023
The Amazon Nova Family of Models: Technical Report and Model Card
Technical report · 2025
Large Language Models of Code Fail at Completing Code with Potential Bugs
Tuan Dinh, Jinman Zhao, Samson Tan, Renato Negrinho, Leonard Lausen, Sheng Zha, George Karypis · NeurIPS 2023
Sequence-level Large Language Model Training with Contrastive Preference Optimization
Zhili Feng, Dhananjay Ram, Cole Hawkins, Aditya Rawal, Jinman Zhao, Sheng Zha · Findings of NAACL 2025
GluonCV and GluonNLP: Deep Learning in Computer Vision and Natural Language Processing
Jian Guo, He He, Tong He, Leonard Lausen, Mu Li, Haibin Lin, Xingjian Shi, Chenguang Wang, Junyuan Xie, Sheng Zha, et al. · JMLR 2020
Talks and interviews
Education
MS in Computer Science, University of Maryland. BS, Shanghai Jiao Tong University.
ElsewhereXLeave anonymous feedback