Solutions  ›  On-Premise Single

Make every GPU you already own pull its weight.

1% Block partitioning, MIG slicing, and intelligent scheduling — running on the GPUs you already own. No replacement, no expansion, no rebuild. Just utilization.

5% → 95%+
GPU utilization
100
tenants per H100
0
new GPUs required
< 30s
job start latency
Recommended For

Built for labs and small AI teams.

If your GPU pool is between 4 and 32 cards and your team is between 5 and 50 researchers, the single-cluster topology is the right starting point. CNLab installs as a Kubernetes operator on your existing nodes — there is no rip-and-replace.

  • • University AI labs running mixed graduate and undergraduate workloads
  • • Small AI companies with 1–4 servers and shared experiments
  • • Educational institutions teaching deep-learning courses with bursty demand
  • • Shared compute clusters within a single business unit

Typical environment

GPU pool4 to 32 GPUs
GPU mixH100 / A100 / L40S
Users5 to 50 active
Existing OSUbuntu / RHEL
K8s1.27+ (or installed by CNLab)
StorageLocal NVMe + optional NFS

Isolated Resource Precision

A four-layer architecture: tenants on top, the orchestration layer in the middle, the GPU virtualization layer below it, and the physical GPU servers at the bottom. Every tenant request flows through the same scheduler.

Single-Cluster Architecture Diagram

On-Premise Single Cluster — full architecture diagram

Click the diagram to view full resolution.

  ┌──────────────────────────────────────────────────────────────┐
  │              MULTI-TENANT USERS                              │
  │   Jupyter · VS Code · CLI · Project APIs                     │
  └──────────────────────────────────┬───────────────────────────┘
                                     │
  ┌──────────────────────────────────▼───────────────────────────┐
  │     INTELLIGENT ORCHESTRATION LAYER                          │
  │  ┌─────────┐  ┌──────────────┐  ┌────────────────┐           │
  │  │Scheduler│  │Resource Mgr. │  │  Monitoring    │           │
  │  │priority │  │quota & MIG   │  │  Prometheus    │           │
  │  │+ deadline│ │assignment    │  │  + Grafana     │           │
  │  └────┬────┘  └──────┬───────┘  └────────┬───────┘           │
  └───────┼──────────────┼───────────────────┼───────────────────┘
          │              │                   │
  ┌───────▼──────────────▼───────────────────▼───────────────────┐
  │         GPU VIRTUALIZATION & SHARING LAYER                   │
  │   MIG slices  │  1% Blocks  │  Memory pinning  │  IPC isol.  │
  └──────────────────────────────────┬───────────────────────────┘
                                     │
  ┌──────────────────────────────────▼───────────────────────────┐
  │     GPU SERVERS  (H100 / A100 / L40S)                        │
  └──────────────────────────────────────────────────────────────┘
GPU Partitioning

Two slicing strategies. Mixed-tenant safe.

CNLab combines NVIDIA's hardware MIG partitioning with our software 1% Block partitioning, choosing automatically per workload.

🔧

MIG (Multi-Instance GPU)

For H100 / A100 cards. Hardware-isolated partitions: 1g, 2g, 3g, 4g, 7g profiles. Use when workloads need full memory bandwidth and hard isolation, e.g. inference servers and large-batch training.

  • • Up to 7 partitions per H100
  • • Full memory bandwidth within each slice
  • • Hardware-isolated faults
📐

1% Block Partitioning

CNLab's software-level slicing: minimum allocation = 1% of CUDA cores + VRAM. 100 simultaneous tenants per GPU. Use for development, lightweight experiments, course assignments — anywhere a MIG slice is too coarse.

  • • 100 blocks per GPU, software-enforced
  • • Dynamic resize without job restart
  • • Combines with MIG on the same node
Workflow

Five steps. Fully automatic.

01

Request

User submits via Jupyter, VS Code, or CLI.

02

Validate

Quota check + priority assignment by RBAC role.

03

Schedule

Scheduler picks the best slice across the cluster.

04

Allocate

GPU partitioned and assigned. Mounts attached.

05

Reclaim

Auto-reclaim on idle threshold or job completion.

No replacement. No CAPEX. Real ROI.

Zero hardware refresh

CNLab installs as a Kubernetes operator on the GPUs you already have. Your physics professor's A100 from 2021 stays useful.

Capacity on day one

First-week measurements show 3–5× throughput on the same iron. The change is in the scheduler, not the silicon.

Forward-compatible

When you add a second cluster or burst to cloud, the same APIs and admin console keep working — no migration.

FAQ

Five questions before you book a demo.

What's the minimum cluster size?

A single GPU on a single node. We've installed against a single L40S in a research closet. Below 4 GPUs, you'll see less benefit; above 32, consider Multi On-Premise.

Which GPUs and drivers are supported?

NVIDIA H100, H200, A100, A40, L40S, L4, RTX 6000 Ada — driver 535+. CUDA 12.x. AMD MI250 / MI300 supported in Multi On-Premise tier.

How is multi-user safety enforced?

Hardware MIG isolation when used; namespace + cgroup isolation otherwise; encrypted IPC; and per-tenant memory pinning. Audit log captures every allocation.

Can users see their own usage and cost?

Yes. Every authenticated user has a personal dashboard with GPU-hours, storage GB-days, and projected cost (if billing is configured). Admins see fleet-wide rollups.

What's the upgrade path to Multi or Hybrid?

Configuration-only. Multi adds a federated control-plane node and a peering link; Hybrid adds cloud credentials in the admin panel. No re-install, no scheduler change.

Get Started Now

Free 30-minute technical Q&A and demo. We'll bring an architecture diagram for your specific cluster.