Solutions  ›  Hybrid Cloud Bursting

On-prem first. Cloud only when it pays.

My GPU First — utilize on-prem 100% before reaching for cloud. When local capacity fills, auto-burst to AWS, GCP, Azure, Naver, or KT Cloud. Return automatically when the job completes. Spend cloud money on demand, never by default.

−70%
cloud cost vs cloud-only
−50%+
total TCO
burst capacity
< 60s
on-prem → cloud handoff

My GPU First.

A simple operating principle. Run jobs on the GPUs you own — they're already paid for. Reach for cloud only when local utilization crosses a threshold you set. The moment a local slot opens up, the cloud workload returns.

PRINCIPLE 01

Local before remote

Every job tries on-prem first. Only when no slot fits within the SLO does the scheduler reach for cloud.

PRINCIPLE 02

Spot before on-demand

For cloud, prefer Spot. Fall back to On-Demand only when spot eviction risk exceeds tolerance for that workload.

PRINCIPLE 03

Auto-return

A cloud-running workload migrates back the instant a matching local slice frees up. The cloud bill stops the moment your GPU does.

Predictive Capacity Management

CNLab's scheduler doesn't just react — it predicts. By learning your workload patterns, it pre-provisions cloud capacity 90 seconds before local saturation hits, avoiding any user-visible queue time.

Auto Burst-out: 5-Step Logic

01

On-Premises First

Jobs queue on local clusters. 100% utilization before cloud is touched.

02

Threshold Detection

When utilization > 90% sustained, the scheduler arms burst-out.

03

Auto Cloud Provisioning

Spot or On-Demand instance auto-provisioned on the configured cloud.

04

Workload Migration

Pending or running job placed on the cloud node. Live migration if running.

05

Auto Return

Workload migrates back when local frees up. Cloud node auto-terminated.

Zero-Downtime Elastic Scaling

Live migration applies between on-prem and cloud the same way it works between two on-prem clusters. The CUDA context, memory pages, and process state move via RDMA-direct transfer (on-prem → on-prem) or encrypted TCP (on-prem ↔ cloud). The user sees nothing.

Cost Optimization

Cloud cost down 70%. Total TCO down 50%+.

Spot / On-Demand auto-selection

CNLab tracks per-AZ spot interruption rates. Below 5% interruption: prefer Spot. Above: On-Demand for SLO-bound jobs, Spot for resilient ones.

Immediate auto-return

Cloud nodes terminate the moment training completes. No hourly waste, no orphaned instances, no Slack post-mortems.

Egress-aware scheduler

For data-heavy training, scheduler picks the cloud region closest to your S3/GCS bucket. Cuts egress costs by 60% vs naive placement.

Reservation pooling

If you have AWS Reserved or GCP Committed-Use discounts, CNLab routes burst-out to those reservations first. Effective rate beats Spot.

Full Hybrid Cloud topology.

The complete picture: on-prem clusters as the primary pool, predictive burst-out, encrypted egress to AWS / GCP / Azure / Naver / KT, and auto-return when local capacity frees up.

Hybrid Cloud — burst-out architecture diagram

Click the diagram to view full resolution.

ROI Calculator

Run the numbers.

Estimate your annual savings from My-GPU-First. Assumes 70% cloud-cost reduction on the workload that today runs cloud-only.

Estimated annual savings
$84,000

Based on 70% reduction × monthly spend × 12. Real savings depend on workload mix; we'll model your actual numbers in a 30-min consult.

Security Strategy

Sensitive data stays on-prem.

Pinned workloads

Mark a project, dataset, or model as on-prem-only. Scheduler never bursts those workloads to cloud, regardless of capacity.

Encrypted egress

All on-prem ↔ cloud traffic goes over TLS 1.3 + WireGuard VPN, with optional cloud-provider PrivateLink/Interconnect.

RBAC enforcement

User-level access control applies in both planes. A user blocked from cloud bursts simply queues on-prem; no policy bypass.

Audit logs on-prem

Audit trail (who provisioned what, when, where) lives on the on-prem control plane — never in the cloud, even for cloud-running jobs.

Supported Clouds

Five clouds. One scheduler.

AWS

EC2 G5/P5, Spot & On-Demand

GCP

A3/G2 Compute, Spot VMs

Azure

ND/NC series, Spot VMs

Naver Cloud

G2/G3 GPU servers

KT Cloud

G-Class GPU instances

Start Your My GPU First Strategy

Free 30-min consult. Tell us your monthly cloud bill and we'll model the on-prem-first save in your inbox within one business day.