Orbital AI Data Centers Hit Network Limits
Orbital AI data centers promise cheap energy and cooling, but network bandwidth, latency, and topology may limit large-scale LLM training in space
Architecture and Infra on ThecoreGrid covers the foundations of designing and operating scalable, reliable systems at BigTech level. This category brings together system design and infrastructure practices: distributed architectures, highload patterns, cloud-native platforms, and core layers such as compute, networking, and storage. We focus on real engineering decisions — how to balance reliability, performance, cost, and long-term system evolution. Topics include Infrastructure as Code, Kubernetes, multi-region deployments, traffic management, and platform design. Content is grounded in production experience: incident post-mortems, large-scale migrations, and lessons from operating infrastructure under heavy load. Instead of abstract theory, you get practical trade-offs, proven patterns, and insights drawn from real-world systems. Architecture & Infra is built for architects, backend and platform engineers, DevOps teams, and SREs responsible for complex distributed systems and mission-critical infrastructure.
Orbital AI data centers promise cheap energy and cooling, but network bandwidth, latency, and topology may limit large-scale LLM training in space
Securing LLM-MAS via continual learning and OOD detection: An analysis of OpenEvoShield and its architecture against dynamic attacks
MoE routing in direct topologies is limited by unpredictable traffic. An analysis of MoX shows how static routing reduces congestion and approaches switch performance. In classic ML clusters, the network is designed for regular collective operations. This works as long as the traffic is predictable. MoE routing breaks this assumption. Each token selects top-K experts, … Read more
A tight lower bound is established for the round complexity of Byzantine Agreement. An analysis of the impact of an adaptive adversary on latency and consensus architecture.
LLM serving workload turns out to be significantly more complex than is typically modeled. FineServe demonstrates how real workloads break simplified assumptions and impact architecture. Modern LLM platforms operate as always-on services with strict requirements for latency and throughput. The main issue is the unstable and heterogeneous LLM serving workload. Most studies have relied on … Read more
How in-band SDN control plane scales without state growth: an analysis of Periplus and its model of multi-controller coordination
In-band SDN with multi-controller coordination: how Periplus limits forwarding state using border-switch graphs and scales the network without increasing load
AI accelerators for scientific computing: bridging the precision, memory, and execution gap when migrating HPC workloads to NPUs –>
Primitive-level synchronization in distributed PBNR training: how to eliminate global barriers and accelerate training without sacrificing quality
Compute-communication overlap in MoE reduces latency through tiled scheduling and signaling. This directly impacts throughput and GPU utilization. Modern Mixture-of-Experts (MoE) systems are limited not by compute, but by communication. In distributed execution, each layer requires two all-to-all operations, and the second—returning results—falls into the critical path. The classical scheme only initiates it after the … Read more
Controls: ← → to move, ↑ to rotate, ↓ to drop.
Mobile: use buttons below.