AI Networking Fundamentals
Collective communication, RDMA, and rail topology explained from the operator's side, including why adding nodes to a training job can make it slower.
Pillar
The physical and platform layer beneath every model: accelerators, GPU cluster topology, RDMA and collective communication, storage for training and inference, schedulers, capacity planning, power and cooling.
Consolidated guides rather than a page per concept. Each area below is one resource, or will be, rather than a cluster of near-duplicates.
Memory capacity, memory bandwidth, and numeric formats — and why bandwidth rather than FLOPS usually sets your inference throughput.
NVLink, RDMA, InfiniBand, and Ethernet for AI. At multi-node scale the fabric, not the accelerator, is normally the limit.
Three different access patterns — dataset streaming, checkpoint bursts, and weight loading — with three different failure modes.
Rack power density, cooling, quota, and the schedulers that decide whether expensive hardware sits idle.
Resources
Every resource states its author, its review status, and whether the procedures in it were executed or only reviewed.
Collective communication, RDMA, and rail topology explained from the operator's side, including why adding nodes to a training job can make it slower.
What determines accelerator performance in production: memory capacity versus bandwidth, interconnect topology, and the checks that find a misplaced workload.
Dataset streaming, checkpoint bursts, and model loading place completely different demands on storage. Designing for one and getting the others wrong is the usual outcome.
No layer of the stack is operated in isolation. These are the sections you are most likely to need next.
Newsletter
A practical weekly briefing on AI engineering, infrastructure, production operations, and the technologies powering the AI stack.
One email a week. No sponsorship placements inside the technical sections. Unsubscribe in one click.