Deep technical posts from the team building CNLab — schedulers, GPU virtualization, cloud-native infrastructure, and the occasional war story.
A deep look at the per-tenant economics of GPU sharing — why MIG isn't fine-grained enough for course-scale or dev-scale workloads, and how CNLab's software slicer fills the gap without sacrificing isolation.
How we move a running CUDA context from cluster A to cluster B in under 800ms — without checkpointing or user-visible interruption.
Step-by-step guide: parallel runtime, gradual cutover, a fair-share mapping that keeps faculty happy.
How a small LSTM trained on queue depth predicts saturation early enough to provision cloud capacity before users feel any wait.
AMD MI300 series is now first-class. ROCm 6.1 + heterogeneous scheduling. Performance numbers vs H100 inside.
A deep dive into the trade-offs of self-hosting vs hyperscaler. The TCO math is more favorable than most teams realize.
What the people building CNLab do day-to-day. Pager rotations, customer escalations, and the cluster room espresso machine.
One email per week, deep technical content only. No marketing.