MetaRoCE for AI Infrastructure on Ethernet
MetaRoCE for Ethernet-based AI infrastructure: how Meta moves intelligence into the NIC, eliminates PFC, and creates a loss-resilient transport for million-GPU scale.
Highload on ThecoreGrid focuses on designing and operating systems that handle massive scale, traffic, and data under strict reliability requirements.
We explore architectures and patterns for horizontal scaling, load distribution, fault tolerance, and performance optimization in distributed environments. Topics include sharding, replication, caching strategies, queueing systems, backpressure handling, and latency reduction under peak load. We analyze real-world trade-offs between consistency, availability, and cost, along with failure scenarios and recovery strategies. Content is grounded in BigTech practices, including incident post-mortems and lessons from operating systems at global scale. You’ll find deep dives into infrastructure behavior, traffic management, autoscaling, and resilience engineering. Instead of simplified guides, the Highload tag delivers practical engineering insights for backend engineers, architects, platform teams, and SREs responsible for building and maintaining systems that must perform reliably under extreme demand.
MetaRoCE for Ethernet-based AI infrastructure: how Meta moves intelligence into the NIC, eliminates PFC, and creates a loss-resilient transport for million-GPU scale.
A central gateway for full telemetry: sizing, load testing, HPA, WAL, and GOMEMLIMIT in a production scenario—with no blind spots.
Git fetch in CI became a bottleneck. An analysis of the gitretriever architecture at Datadog, CPU reductions, relay models, and APIs for code queries.
VDR routing in ultra-dense networks: how Helmholtz-Hodge decomposition helps reduce loops, delay, and control-plane overhead.
MCPTT interoperability under high load: how jitter, QoS and carrier boundaries trigger voice failure in dense stadium networks.
Fast BFT SMR under Byzantine faults: why n ≥ 5f + 1 is optimal, how recovery works, and what trade-off n ≥ 7f + 1 simplifies
AI Moderation Platform in the Marketplace: How DoorDash Reduced Incidents, Separated the Low-Cost Layer from LLM Scoring, and Why Boolean Logic Proved to Be a Poor Choice
L.OS on AWS: how Bosch unified vehicle tracking through serverless architecture, connector layer, and provider-specific adapters for real-time visibility
Cloudflare cdnjs on R2 and Workers: how the migration of publishing and delivery simplified the architecture without losing URL, SRI hashes, and fault tolerance
Randomized LL/SC using FADD: preserving the QHI property, reducing capacity complexity, and ensuring wait-free operation without asymptotic overhead
Controls: ← → to move, ↑ to rotate, ↓ to drop.
Mobile: use buttons below.