Tier 1 — Operations analyst
On shift, in-country. Owns acknowledgement, classification, runbook execution and the incident record. Instructed to escalate rather than improvise.
Holds the incident end to end
Seven practices, one engineering team. We design, build, benchmark and operate the compute layer beneath your hardest problems — on-premise, hybrid or sovereign cloud.
A genomics pipeline and a trading engine are not the same machine. We start from your regulatory perimeter, your data gravity and your deadlines — then design backwards.
Managed Services
Operations is where infrastructure value is either realised or quietly lost. We monitor, patch, tune and cost-control the estates we build — from a network operations centre staffed in Australia, by engineers who were on the design review.
The Constraint We Design Around First
Every managed services proposal promises 24/7. The question that decides the outcome is who picks up. And whether they can already see your architecture, or are reading it for the first time while the cluster is down.
A build ends. Operations does not. The value of a well-specified cluster is realised across the four years after commissioning, or lost in them one week at a time.
Most degradation is not an outage. It is a queue that drains a little slower each week. A storage tier whose p99 has doubled since March. A GPU clocked down by a failing fan, still billing at full rate. None of that trips a reachability check, and none of it appears in a monthly uptime figure.
So we staff operations with the people who did the design. Tier three escalation reaches an engineer who holds the as-built topology in their head, not a triage script written by someone who has never seen your fabric. That is the whole differentiator. It is also why we will not operate an estate we were not permitted to document.
The responder already knows the architecture, because we built it.
Operations Centre
In-country staffing, four escalation tiers, and a monitoring scope that goes well past reachability. We manage alert quality, never alert volume.
Sovereign NOC. Shift roster, console wall and escalation board.
Every shift is worked from Australia. There is no follow-the-sun handover to an offshore desk at 11 p.m., because a handover is where context dies and where an incident stops making progress.
Monitoring telemetry stays in the same jurisdiction as the estate it describes. We do not ship metrics to a shared offshore observability tenancy, because for government and defence work that single decision undoes the rest of the design.
On shift, in-country. Owns acknowledgement, classification, runbook execution and the incident record. Instructed to escalate rather than improvise.
Holds the incident end to end
Scheduler, fabric, storage and hypervisor depth. Holds root on your estate and the change authority to use it inside an agreed window.
Diagnosis and remediation
The engineer who specified the build. Reached by name from your escalation card, not through a queue. Joins any P1 by default rather than on request.
Architectural authority
Silicon vendor support contracts, colocation remote hands and the data centre's own operations desk. We drive those cases and keep the incident open on our side.
We hold the thread, not you
Telemetry
These are the series we alarm on, and the reason each one matters weeks before it becomes an outage. Thresholds are published to you and any of them can be vetoed.
Slurm queue depth by partition, pending reason codes, fair-share drift between projects, and job failure rate by node. A queue that lengthens for three straight days is a capacity decision, not an incident — and it should reach a human as exactly that.
Leading indicator — capacity exhaustion
Corrected and uncorrected ECC counts, retired pages, NVLink and PCIe replay errors, and SM clock measured against the thermal and power caps. A throttled accelerator costs the same per hour as a healthy one and finishes nothing on time.
Leading indicator — silent performance loss
InfiniBand symbol-error and link-downed counters, congestion and credit-stall metrics, RDMA retransmit rates, and port flap history per leaf. One marginal cable degrades every collective operation that crosses it, and the job log blames the application.
Leading indicator — one bad optic or cable
Per-target latency percentiles across Lustre and BeeGFS, NVMe-oF namespace queue depth, metadata operation rates, and rebuild or scrub progress. We alarm on p99 and p99.9. An average read latency has never once told anyone the truth.
Leading indicator — tail latency before collapse
Rack inlet temperature and delta-T, coolant distribution unit supply temperature and flow rate, PDU phase balance and branch-circuit headroom, fan and pump duty cycles. Headroom is the number that predicts a February outage in November.
Leading indicator — summer capacity limits
PTP offset and path-delay variance, PPS signal integrity, grandmaster holdover state, plus BGP session health and egress volume per flow. Trading and timestamping estates fail on time discipline long before they fail on compute.
Leading indicator — timestamp and audit risk
A NOC that pages on everything trains its own staff to ignore it. Within a month the operator who acknowledges fastest is the one who reads least. That failure mode is cultural, it is predictable, and it is created by the monitoring configuration rather than by the people.
So every alarm we create carries three things: a runbook, a named owner, and a review date. If an alarm fires three times without a decision attached to it, the alarm is wrong. We rewrite it or delete it, and the change appears in your monthly service report with the reasoning.
Engagement Models
Monitor if you have capable engineers and want a second set of eyes. Operate if you want the work done. Fully Managed if you want one accountable party for the platform, its documentation and its spend.
| Service dimension | Monitor | Operate | Fully Managed |
|---|---|---|---|
| Coverage hours | 24/7 monitoring, response in business hours (AEST/AEDT) | 24/7/365 monitoring and response | 24/7/365, with tier three on call |
| P1 response target | 30 minute acknowledgement, advisory only | 15 minute acknowledgement, engineer engaged within 30 minutes | 15 minute acknowledgement, engineer engaged within 15 minutes |
| P2 response target | Next business day | 4 business hours | 2 hours, around the clock |
| Remediation authority | You act, we advise and document | We act under runbooks you have approved | We act, including unscripted diagnosis and recovery |
| Patching and firmware | Advisory notices, you schedule the work | Quarterly OS updates, monthly security patching in your window | Monthly OS, quarterly firmware and BIOS baseline, canary node first |
| Capacity review cadence | Quarterly written report | Monthly written review of utilisation and queue trend | Monthly review plus a quarterly capacity and commitment plan |
| On-site attendance | By the hour, on request | Next business day for hardware faults | 4 hour metro target, next available flight regional, local spares held |
| Named engineer | Shared operations queue | Named platform engineer | Named platform engineer and named design engineer |
| Runbook ownership | You write them, we review annually | We write and maintain them | We write, maintain and game-day test them quarterly |
| Change execution | Yours entirely | Ours, in your window, with minuted change approval | Ours, against a change freeze calendar we hold on your behalf |
| Reporting | Monthly availability and alert summary | Monthly service report with incident register | Monthly report plus a quarterly service review with your executive |
Deployment & Handover
Commissioning is not the moment it powers on. It is a written acceptance test against the numbers in the proposal, run in front of your team, reporting the misses as plainly as the passes.
Measured result beside proposed result, line by line.
Acceptance testing re-runs the benchmark suite named in the proposal, on your data where licence and privacy allow. It prints the measured number beside the proposed one. Where we fall short we state by how much and why, before you sign.
A restore is tested during commissioning, not described. Until a file has come back from the backup tier and been checksummed, you do not have a backup — you have an expense with optimistic documentation.
Cloud Optimisation
We have yet to review an estate where the largest recoverable line was a badly negotiated price. It is nearly always compute nobody switched off — development, staging, and a training cluster from a project that finished last year.
We take 90 days of CPU, memory, GPU and IOPS telemetry per workload and size to the measured p95, with headroom you agree to in writing. Instance families shift generation to generation, so we re-test rather than assume the newest silicon is cheaper for your particular shape. Memory-bound workloads routinely get worse on a faster core count.
Reserved instances win where the shape is fixed and the term is survivable. Savings plans win where the shape moves but the spend does not. We model both against your own twelve-month usage curve, buy in tranches rather than one decision, and hold a deliberate on-demand margin. A commitment you have grown out of becomes a floor you are paying to ignore.
Egress is the cost that never appears in the sizing spreadsheet. We map every cross-region, cross-account and internet-bound flow, then put the chatty ones behind a cache or a private interconnect. After that we move the workload to the data, rather than the data to the workload. For estates with a sovereignty obligation, the same map doubles as evidence of where the bytes actually go.
Class policy comes from access telemetry, not habit. Hot tier for the working set, infrequent access past thirty days, archive on a retention rule, deletion on a legal one. We test the retrieval before we trust the tier. An archive you cannot restore inside your recovery window is a deletion with extra paperwork.
In most reviews, non-production is the biggest single recoverable line. Orphaned volumes, load balancers with no healthy target, GPU nodes held in case, and environments running 168 hours a week to serve about forty. We schedule them off, tag every survivor with an owner, and hand you the list of resources nobody was willing to claim. That conversation is awkward and it is where the money is.
34%
Median reduction in monthly cloud invoice across optimisation engagements, measured against the trailing quarter before review.
61%
Share of recovered spend attributable to idle development, staging and abandoned experiment environments rather than production rightsizing.
0rebuilds
Optimisation work delivered without asking engineering teams to refactor an application. Rewrites are a separate decision with a separate business case.
Advisory
A block of engineering hours for the decisions that are expensive to reverse: fabric topology, storage class, scheduler policy. Or whether the thing you are about to buy is the thing you need.
Bring the vendor's proposed topology and we will mark it up. Oversubscription ratios, failure domains, rebuild time at the proposed drive size, the single power feed nobody costed, and the cost of growing it by half. We are not bidding on that hardware, which is the point.
We instrument the real job — an OpenFOAM case, a Nextflow pipeline, a vLLM serving path — and report where the wall-clock time goes. It is usually not where the procurement assumed. The finding often changes the shopping list more than the budget.
A written assessment of monitoring coverage, runbook completeness, restore evidence and patch currency — mapped to the Essential Eight where that is your obligation. You get a prioritised remediation list with effort estimates, whether or not you then engage us to do the work.
Retainer hours can be spent on someone else's architecture, including one that is currently failing. We will tell you what we would change, and whether the honest answer is a change of architecture or a change of expectation. Sometimes the design is sound and the deadline was never real.
Incident Practice
Written down in advance, because the middle of an incident is a poor time to invent a process. Every interval below is a design target for planning, not a contractual commitment.
An alarm fires with a runbook already attached, a synthetic transaction fails, or you telephone the NOC. All three paths open the same incident record with the same severity rules. There is no slower, second-class path for a fault a human noticed first.
The analyst on shift classifies severity against a published matrix and acknowledges to a person, not to a ticket queue. P1 means the platform is unusable, losing data, or breaching a regulatory obligation. Nothing else earns it, and inflating severity to get attention is treated as a process defect.
Tier one holds the incident and executes the runbook. If the runbook does not resolve it within the first cycle, we escalate rather than retry the same step with more conviction. The design engineer at tier three joins any P1 by default — that is the commitment the rest of this page depends on.
One incident channel, one named incident commander, and a written update every thirty minutes whether or not there is news. Silence is the most common complaint made about managed service providers, and it is entirely avoidable. An update saying we still do not know is a legitimate update.
Restoring service and finding the cause are separate jobs, done in that order. We will fail over, drain a node or roll a firmware level back to get you working. The diagnosis becomes harder as a result, and we accept that. Where evidence would be destroyed, we capture it first and say so in the log.
Closed means measured. The queue draining at its normal rate, the latency distribution back inside its envelope, failed jobs re-queued or formally accounted for. We verify against telemetry rather than against the absence of alarms, because an alarm can be silenced by the same fault that caused it.
Within five business days you receive a written review: timeline, root cause, contributing factors, the changes we have made, and what we got wrong. If the cause was ours it says so in the first paragraph. We do not issue an unexpected combination of factors as a root cause, and we will not describe a capacity decision as an unforeseeable event.
Commercial Detail
The questions procurement asks after the technical evaluation is finished. Answered here so they are not a surprise in the contract negotiation.
Four things that matter: where availability is measured, what counts as a response, what is excluded, and what the remedy is. We measure at the service boundary you care about: the scheduler accepting jobs, the filesystem mounting, the inference endpoint answering. Not a ping to a management interface.
Exclusions are named, not implied: scheduled maintenance in an agreed window, change you made outside the process, and confirmed vendor firmware defects while the vendor case is open. Service credits are a remedy, not a business model. If you are collecting them regularly, the architecture is wrong and we will say so in the quarterly review.
Ninety days' notice either direction, no exit fee, and no penalty for terminating a service that has not performed. On exit you receive everything in open, machine-readable formats. The monitoring configuration, the alarm catalogue with thresholds, every runbook, the as-built documentation, the incident history and the capacity data.
Your infrastructure-as-code lives in your repository for the whole engagement, not ours. Nothing needs to be extracted from a proprietary console at the end, because nothing important was ever only in one.
Operations staff need telemetry and console access, not your records. Access is role-scoped, reached through jump hosts inside your jurisdiction, and recorded for the session. Data-plane access is separately requested, separately approved and separately logged, and most incidents never require it.
Monitoring data stays in the same jurisdiction as the estate. On PROTECTED environments, access is limited to cleared personnel. Every access event ships to your own SIEM, so the audit trail is not held solely by us.
Yes, after an operational readiness assessment. We need monitoring coverage we trust, documentation we can follow at 3 a.m., and a patch baseline we can defend. Where those are missing we quote the remediation separately and plainly.
We will not sign a 24/7 tier over a system we cannot see into. Taking that money would be selling you a response time we have no mechanism to meet.
No. Shifts are worked by our own employees in Australia, on a roster you can inspect. Vendor and facility support are subcontracted by their nature: silicon vendor engineering, colocation remote hands, the data centre's own operations desk. We drive those cases ourselves rather than forwarding you into them.
We hold the case with the vendor and keep your incident open on our side until it is genuinely resolved. You get one thread and one commander, not three. Where a firmware defect is confirmed, the published review names the version, the symptom and the workaround, so your other estates benefit from it.
Yes, per platform and per environment. A common shape is Fully Managed on the production cluster and Monitor on development. The boundary is written down, with named owners on both sides, so there is no argument at 3 a.m. about whose failure it is.
Operations Handover
No capability deck. We read the reviews, your alarm catalogue and the shape of your on-call roster. Then we write back with the gaps we can see, and which tier closes them. If the honest answer is that you do not need us yet, you will get that instead.
Sydney · Melbourne · Canberra. Shifts worked in Australia, every hour of the year.