Technical write-ups on distributed systems, ML infrastructure, and lessons from building AI at scale.
Deep dive into how gradient synchronization works in PyTorch Distributed Data Parallel and why it bottlenecks at scale.
Lessons learned from building a team of specialized LLM agents that collaboratively review code.
A practical guide to diagnosing and fixing common failures in distributed ML training pipelines.