The catalogue for this hardware reads as though it were designed to stop you understanding what you are buying.
DGX, HGX, MGX, NVL72, BlueField, InfiniBand, Spectrum-X. Every one of them arrives in the same sentence, in the same deck, from someone who is genuinely convinced you already know the difference.
You do not need to know all of it. You need to know which four decisions those names are hiding, because those are the ones you will be signing.
🔑 The short version
- DGX, HGX and MGX are not competing products. They are three ways of buying the same thing, with different amounts of responsibility.
- The real question inside the rack is how the accelerators talk to each other, not how fast each one is.
- Leaving the rack there are two roads, and picking one shapes who you can hire afterwards.
- BlueField is the piece nobody explains well, and it is the one that changes your operating model.
- The part that locks you in hardest is not hardware at all: it is the software layer and its licences.
Three names people mix up constantly
Start here, because almost every confused conversation I have witnessed starts with these three.
| Name | What it really is | When it makes sense |
|---|---|---|
| DGX | The finished appliance, built and supported by the manufacturer. You plug it in | You want it working and you are willing to pay for that certainty |
| HGX | The board with the accelerators, which a server maker builds into their machine | You already have a preferred server vendor and a relationship worth keeping |
| MGX | A modular design spec other manufacturers build to | You need flexibility of configuration or of supplier |
They are not three tiers of quality. They are three points on a scale of how much integration work is yours. DGX is the most expensive and the least ambiguous; MGX is the most flexible and asks the most from your team.
In Latin America the choice usually gets decided by a fourth factor nobody puts on the slide: which of the three anyone will actually support locally, and how long a replacement part takes to clear customs.

What happens inside the rack
Back to the kitchen analogy, because it holds.
The accelerators are the cooks. Inside a rack, what matters is not how fast each one chops — it is how quickly they can hand each other the bowl. They work in lockstep: compute, exchange, compute again. If passing the bowl is slow, having faster cooks changes nothing.
That is what a rack-scale system such as an NVL72 is selling. Not more accelerators — a shorter counter between them. Seventy-two accelerators wired so that the exchange between any two is as close to instant as the physics allows.
And that reframes the purchase honestly: you are not buying compute by the unit, you are buying a rack as one object. It is not sold in pieces, and it does not behave like a room full of servers you assembled yourself.
Leaving the rack, there are two roads
Once the traffic goes beyond one rack you face a genuine fork, and it is worth understanding before somebody chooses for you.
One road is InfiniBand: a specialised network, decades old in high-performance computing, superb at exactly this, and operated by people with a specific and scarce skill set.
The other is Ethernet built for this purpose: the network your team already understands, extended to behave well enough for the job.
BlueField, or the piece nobody explains
This one always gets a vague sentence in the presentation, and it deserves better because it changes how you operate.
A BlueField is a network card that is also a small computer. It takes over the work that used to steal cycles from the main processors — moving data, encryption, storage traffic, isolating tenants from each other.
Two consequences that matter to whoever owns the service:
- The accelerators stop doing plumbing. More of the expensive silicon spends its time on the work you actually bought it for.
- There is now a control layer separate from the machine. That is what makes it possible to rent this hardware to several tenants safely — which, if you are building a service rather than a lab, is the entire business model.
That second point is why I pay attention to it. It is not a performance component. It is the component that makes multi-tenancy possible, and multi-tenancy is the difference between owning expensive hardware and selling a service.

What is not hardware, and ties you down far more
Here is the part that gets waved through, and it is where the real commitment lives.
Above all of this sits a software layer: the libraries the frameworks call, the orchestration, the management tooling, the enterprise licences. It is genuinely good software, and it is also where the lock-in actually is.
The hardware you could, in principle, replace. The years of tuning, the operational habits, the internal tooling built around that software layer — those do not move. It is the same lesson as any observability platform: the switching cost was never the box.
So when someone tells you the hardware is a commodity decision, ask what the licences cost in year three, and what happens to them if you change supplier.
What to demand from whoever is selling it
| The question | Why it matters |
|---|---|
| DGX, HGX or MGX — and why that one for us? | If the answer is “it’s what we sell,” you are not getting advice |
| How many kW per rack does this need, and does our site have them? | This kills more of these projects than any technical factor |
| InfiniBand or Ethernet, and who operates it afterwards? | The scarce skill is the constraint, not the technology |
| What do the software licences cost in year three? | Where the lock-in lives. Get it in writing before year one |
| Who supports this locally, and how long does a spare part take? | In this region, this frequently decides the whole thing |
| Can we test with our real workload, not a demo? | One of everything always works. Scale is where the truth is |
Differences between the DGX, HGX and MGX families: HGX, DGX and MGX compared and DGX versus HGX. Rack-scale deployments and configurations: analysis of Blackwell deployments. Manufacturer ecosystem: NVIDIA and server manufacturers.
Frequently asked questions
What is the difference between DGX, HGX and MGX?
They are not competing products but three ways of buying the same capability with different amounts of responsibility. DGX is the finished appliance built and supported by the manufacturer. HGX is the board with the accelerators, which a server maker integrates into their own machine. MGX is a modular specification other manufacturers build to. DGX costs most and leaves least ambiguity; MGX is most flexible and asks most from your team.
Should I choose InfiniBand or Ethernet?
Both work, so performance is rarely the deciding factor. The better question is who will operate it in eighteen months. InfiniBand is a specialised skill that is hard to hire and expensive to retain, particularly in Latin America; purpose-built Ethernet asks less of your people and slightly more of your engineering. Choosing on a benchmark and regretting it on a hiring cycle is a common outcome.
What does a BlueField actually do?
It is a network card that is also a small computer, taking over work that would otherwise consume the main processors: moving data, encryption, storage traffic and isolating tenants. Two consequences matter — the expensive accelerators stop doing plumbing, and there is now a control layer separate from the machine. That second one is what makes it possible to rent the hardware to several tenants safely, which is the difference between owning equipment and selling a service.
Where is the real vendor lock-in in this stack?
Not in the hardware — in the software layer above it, and its licences. The libraries, orchestration and management tooling are genuinely good, and the years of tuning, operational habits and internal tooling built around them do not move if you change supplier. The hardware could in principle be replaced; the accumulated software dependency is what actually holds you. Ask what the licences cost in year three.
What decides these projects in Latin America specifically?
Two things that rarely appear on the slide: how many kilowatts a rack in the building can take today, and who supports the equipment locally — including how long a replacement part takes to clear customs. Both routinely outweigh the technical comparison, and most organizations in the region will reach this hardware through a sovereign cloud or regional colocation rather than by building it themselves.
What I take from this
The catalogue is not confusing by accident, though I do not think it is malicious either. It is confusing because it grew by accumulation, and because everyone inside that world already knows what the words mean.
But the effect is real: a buyer who cannot name the four decisions ends up signing whatever the seller’s default configuration happens to be. And the seller’s default is optimized for the seller’s business, which is a perfectly normal thing for it to be optimized for.
So the useful translation is short. How much integration work is ours, how do the accelerators talk to each other, who operates the network afterwards, and what do the licences cost in year three. Four questions. Everything else on the slide is detail.
And if the answer to any of them is a brand name rather than a reason, you have learned something more valuable than the answer. 🧩
May the systems be with you. ✦
Featured image generated with artificial intelligence.

Ethel Méndez
Senior Product Manager in B2B observability and networking, with 20 years of field work, NOC and managed services across Latin America. I write what I learned running real systems, not what I read in a course.
