<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>Sheng Zha</title><description>AI researcher and builder writing about model training, systems, open source, and organizations.</description><link>https://szha.ai/</link><item><title>Compute-Optimal Is Not Cluster-Optimal</title><link>https://szha.ai/blog/compute-optimal-is-not-cluster-optimal/</link><guid isPermaLink="true">https://szha.ai/blog/compute-optimal-is-not-cluster-optimal/</guid><description>MOSAIC jointly selects a sparse-MoE architecture, token budget, and parallel layout under a fixed cluster and training window.</description><pubDate>Wed, 12 Aug 2026 00:00:00 GMT</pubDate><category>pretraining</category><category>scaling-laws</category><category>moe</category><category>systems</category></item><item><title>Research Problems in Pretraining</title><link>https://szha.ai/blog/research-problems-in-llm-pretraining/</link><guid isPermaLink="true">https://szha.ai/blog/research-problems-in-llm-pretraining/</guid><description>A practitioner&apos;s account of what pretraining research can predict, where current methods break, and which questions remain open.</description><pubDate>Wed, 03 Jun 2026 00:00:00 GMT</pubDate><category>pretraining</category><category>scaling-laws</category><category>optimization</category><category>muP</category><category>research</category></item><item><title>Your Org Has the Same Scaling Problem as a Badly Tuned Training Run</title><link>https://szha.ai/blog/badly-tuned-training-run/</link><guid isPermaLink="true">https://szha.ai/blog/badly-tuned-training-run/</guid><description>AI raised individual throughput but coordination overhead stayed fixed. For many product-engineering orgs, the bottleneck flipped from compute-bound to communication-bound.</description><pubDate>Wed, 15 Apr 2026 00:00:00 GMT</pubDate><category>management</category><category>ai</category><category>scaling</category><category>organizations</category></item><item><title>On Assessing the Value of a Project</title><link>https://szha.ai/blog/on-assessing-the-value-of-a-project/</link><guid isPermaLink="true">https://szha.ai/blog/on-assessing-the-value-of-a-project/</guid><description>A practical framework for comparing research projects by probability of success, effect size, and weighted reach.</description><pubDate>Tue, 20 May 2025 00:00:00 GMT</pubDate><category>research</category><category>decision-making</category></item><item><title>Determining Model Size and Training Horizon through Scaling Laws</title><link>https://szha.ai/blog/model-size-and-training-horizon-scaling-laws/</link><guid isPermaLink="true">https://szha.ai/blog/model-size-and-training-horizon-scaling-laws/</guid><description>Deriving model size and training tokens from a fitted scaling law, then extending the calculation to repeated data, inference demand, and cluster efficiency.</description><pubDate>Mon, 02 Dec 2024 00:00:00 GMT</pubDate><category>scaling</category><category>training</category><category>compute-optimal</category><category>chinchilla</category></item><item><title>GluonNLP — Deep Learning Toolkit for Natural Language Processing</title><link>https://szha.ai/blog/gluonnlp-deep-learning-toolkit-for-nlp/</link><guid isPermaLink="true">https://szha.ai/blog/gluonnlp-deep-learning-toolkit-for-nlp/</guid><description>Why we built GluonNLP to make NLP experiments easier to reproduce, maintain, and reuse.</description><pubDate>Tue, 24 Jul 2018 00:00:00 GMT</pubDate><category>nlp</category><category>deep-learning</category><category>mxnet</category><category>open-source</category></item></channel></rss>