Solutions

Seven practices. One engineering team. Nothing quoted before it is measured.

Cloud Natives designs, builds and operates compute infrastructure for Australian organisations that have run out of headroom. We start with your job on our bench, not with a reference architecture. If the measurement says you do not need what you came to buy, that is the advice you get.

Capability

Pick the practice that owns your bottleneck.

The seven practices share one team, one set of instrumentation and one review board. Most engagements touch three of them, because a training run that misses its window is rarely a GPU problem alone.

First Constraint

The bottleneck is almost never the accelerator.

Five things decide whether a compute estate performs, and silicon selection is the last of them. We measure all five before recommending anything, because getting the order wrong is how organisations buy hardware they cannot power.

01

Power and thermal envelope

A rack you cannot cool is a rack you cannot use. We model kilowatts per rack, return-water temperature and containment strategy before choosing silicon, because the building usually decides the architecture. Where the facility cannot take the density, we say so at week one rather than at commissioning.

02

Memory bandwidth

Most jobs described as compute-bound are bandwidth-bound. A roofline analysis tells us whether more cores will help or whether the answer is high-bandwidth memory and a different node shape. It takes about a day to find out and routinely changes the bill of materials by six figures.

03

The storage path

Between the media and the application there are a dozen places to lose a gigabyte per second: the file system client, the page cache, the protocol, the switch buffer, the queue depth. We instrument each hop instead of trusting an aggregate figure from a data sheet.

04

Queue and scheduling policy

High utilisation means nothing if the jobs that matter wait three days. Fair-share weighting, preemption rules and reservation policy are governance decisions with technical consequences, so they belong in the architecture document rather than in an operations ticket six months later.

05

Fabric topology

Collective operations expose topology mistakes that point-to-point tests never will. We test with the collectives your code issues, at the message sizes it issues them, then size oversubscription accordingly. Only after that does the choice of accelerator become an interesting question.

How We Engage

Five gates, and you can stop at any of them.

Each gate produces an artefact you keep, whether or not you continue. There is no discovery phase that exists only to justify the next phase.

  1. Workload assessment

    We take the real job: the mesh, the model checkpoint, the market data capture, the sequencing run. Then we profile where the wall-clock time actually goes. No questionnaire, no maturity matrix, no capability heat map. Typically one to two weeks, depending on how easily the workload can be shared.

  2. Benchmark on our bench

    The workload runs on representative hardware in our lab and you get the raw numbers, including the runs that disappointed. If a cheaper configuration wins, we will tell you, and the engagement can reasonably end there. We would rather lose the order than defend a benchmark we did not believe.

  3. Architecture and costing

    You receive a topology, a bill of materials, a power and cooling budget and a three-year total cost of ownership, with every assumption listed so your team can argue with it. Where a public-cloud option is genuinely cheaper for your duty cycle, the model shows that too.

  4. Build, burn-in, commission

    Systems are assembled and soak-tested in-country, then commissioned on site against the benchmark from gate two. Acceptance is a measurement with a pass threshold, not a signature on a delivery docket. Anything that fails soak testing does not ship.

  5. Operate and review

    Monitoring, on-call and change control from an Australian operations centre. Every quarter we review capacity, spend and incidents, and state plainly what we would design differently with what we now know about your workload.

Delivery Models

On-premise, hybrid, or sovereign cloud.

This is a per-workload decision, not a per-organisation one. Most estates we build end up mixed: steady-state training on owned hardware, seasonal peaks in a sovereign region, and one air-gapped enclave that never touches either.

Indicative characteristics only. Figures require verification before publication.
Decision factor On-premise Hybrid Sovereign cloud
Data residency Your facility, your jurisdiction. Physical custody is demonstrable to an assessor on a site visit. Split by classification. Requires a documented data-flow map and enforced egress controls. Australian regions with Australian-operated support. Contractual and physical residency, no offshore administration path.
Cost shape Capital up front, low marginal cost per job. Wins where sustained utilisation stays above roughly 60 per cent. Capital for the steady state, operating spend for the peaks. Needs a commitment strategy or you pay twice. Operating spend only. Predictable monthly, and more expensive per FLOP once load is continuous.
Time to first workload Roughly 8 to 16 weeks, gated by power, floor space and component lead times. Two to six weeks for the cloud portion, with the owned portion following. Days to weeks, subject to accreditation paperwork rather than engineering.
Elasticity Fixed. Burst capacity must be planned and funded a financial year ahead. Elastic at the edges, fixed at the core. The usual answer for research and rendering peaks. Elastic within regional capacity. Accelerator availability, not contract terms, is the practical ceiling.
Classification ceiling Up to and including fully air-gapped environments with no external route. PROTECTED where the boundary is documented, tested and enforced. PROTECTED under assessed control sets. Confirm the current assessment status in writing.
Key custody You hold the hardware security module and the keys. We never need them. Split custody with a documented break-glass procedure and dual control. Customer-managed keys, held in-country, revocable by you without a support ticket.
Typical fit Sustained model training, defence enclaves, sub-microsecond trading paths. Research with seasonal demand, media rendering, platforms growing faster than their forecast. Agencies without a data-centre programme, regulated startups, disaster recovery estates.

The Bench

We benchmark before we quote, and we show the failures.

The bench is a small lab with representative silicon across the practices, kept deliberately heterogeneous so comparisons mean something. Your workload runs on more than one configuration and you see all the results, including the ones that argue against the more expensive option.

  • Current and prior-generation GPU nodes, air and liquid cooled
  • InfiniBand and 400 GbE fabric segments for collective testing
  • Parallel file system test rigs on NVMe with tape recall behind them
  • Hardware timestamping and a GNSS-locked grandmaster for latency work
Facility photography — benchmark lab, instrumented rack. 1600×1200
2 wks Median time from workload received to benchmark report
1 in 5 Assessments that conclude no new hardware is required
3 Configurations tested per workload, minimum

Procurement

The questions contracting teams ask first.

Answers written for the person who has to defend the purchase, not the person who wants the hardware.

Gate One

Tell us which of the five constraints you have already ruled out. We will measure the rest.

A workload assessment is a fixed-price piece of engineering with a written report at the end. You keep the data and the harness whatever you decide to do next.

New engagements
hello@cloudnatives.example
Direct line
+61 0 0000 0000
Tender and panel enquiries
hello@cloudnatives.example

Capability statement and referee list available under a mutual non-disclosure agreement.