Choose the cluster topology that matches your scale today. Expand without rewrites tomorrow. Every CNLab deployment shares the same scheduler, the same APIs, and the same admin console — only the resource boundary changes.
Auto-detection of deep-learning GPU resource allocation timing — across every environment your researchers already use.
Open a notebook, import torch, call .cuda() — CNLab routes the request to a sliced or full GPU on the closest cluster. No code changes, no kernel surprises.
Cursor, Windsurf, JetBrains — the CNLab agent runs as a remote runtime. Workspace files, secrets, and SSH keys are forwarded transparently.
cnlab run train.py --gpu h100-25 queues your script on the next available 25%-block H100. Logs stream back; the GPU releases on exit.
Choose by org size, GPU pool, and security posture. Every CNLab tier is forward-compatible — Single grows into Multi grows into Hybrid.
| Dimension | On-Premise Single | Multi On-Premise | Hybrid Cloud |
|---|---|---|---|
| Recommended for | Labs, small AI teams | Multi-campus universities, mid-sized AI | Enterprise, AI scale-ups |
| Typical pool size | 4–32 GPUs | 32–256 GPUs | 128+ GPUs + cloud burst |
| Initial cost | Lowest | Medium | Medium + cloud opex |
| TCO | Lowest | Low | −50%+ vs. cloud-only |
| Failover | Single cluster | Cross-cluster live migration | Cross-cluster + cloud |
| Heterogeneous GPU | — | NVIDIA + AMD | NVIDIA + AMD + cloud |
| Data residency | 100% on-prem | 100% on-prem | On-prem first, encrypted egress |
| Min. allocation unit | 1% Block | 1% Block | 1% Block |
| Setup time | Days | Weeks | Weeks |
Not sure which fits? Talk to an engineer →
Achieve 5%→95%+ utilization without replacing or expanding GPU servers. Maximize ROI without additional CAPEX.
Consolidate distributed GPU resources under a single control plane. 2× pool expansion, eliminated single-cluster SPOF.
My GPU First — utilize on-prem 100% before cloud. Auto-burst to AWS / GCP / Azure / Naver / KT when local capacity fills.
Every CNLab deployment runs the same orchestration primitives. Switching topology is a configuration change, not a rewrite.
1% Block + MIG. 100 simultaneous tenants per H100.
Priority + deadline queues. Proactive memory control.
Zero-downtime workload hand-off between any two pools.
User & project-level access control with full audit.
Book a 30-minute architecture review. We'll map your current GPU footprint, model the savings against your cloud spend, and recommend the right tier.