• 10th Grade · State First Prize · 11th Maker Star • 4 min read

How I Cut Campus Server Latency by 73%

Table of Contents

The Problem

Our school ran an internal campus website. When students tried to access it at the same time (such as during course registration), the server slowed down significantly. Most requests finished reasonably, but a fraction suffered severe delays.

This is known as tail latency. When 99% of requests complete within 10ms but 1% take 500ms or fail, the overall average looks acceptable on paper, but the actual user experience breaks down for real students during peak traffic spikes.

High school campus core network routing and server distribution topology Figure 1: High school campus core network routing and server distribution topology.

My Approach

I set up a physical ARM testbed using Raspberry Pis to mirror the school server environment and evaluate different traffic routing schemes under simulated load.

Physical rack infrastructure and core L3 switch deployment Figure 2: Physical rack infrastructure and core L3 switch deployment.

I tested three foundational routing strategies:

  • Round Robin: Distributes requests to each server sequentially. Simple, but blind to dynamic server load.
  • Least Connections: Forwards traffic to the host with the smallest number of active sessions. Smarter, but introduces polling overhead before each routing decision.
  • Weighted Round Robin: Routes requests proportionally based on predefined host processing capacities.

Multi-container isolation and micro-node architecture on ARM testbed Figure 3: Multi-container isolation and micro-node architecture on ARM testbed.

The R-MLO Mechanism

To balance overhead and responsiveness, I designed a Resource-Monitor Load Offloading (R-MLO) mechanism. It deploys lightweight Docker containers across the testbed, continuously tracking CPU and memory pressure. When sustained load thresholds are exceeded, excess traffic is diverted dynamically to idle standby nodes.

R-MLO adaptive offload scheduling logic and state machine Figure 4: R-MLO adaptive offload scheduling logic and state machine.

Implementation

The implementation utilized Docker containers with Nginx functioning as the primary reverse proxy and load balancer.

Reverse proxy configuration for dynamic and static upstream blocks Figure 5: Reverse proxy configuration for dynamic and static upstream blocks.

Upstream cluster definitions and connection timeout parameters Figure 6: Upstream cluster definitions and connection timeout parameters.

Load Testing with Locust

To validate the configuration under load, I executed synthetic stress tests using Locust, generating concurrent user patterns against the endpoints.

High-concurrency synthetic load simulation with Locust benchmark Figure 7: High-concurrency synthetic load simulation with Locust benchmark.

The empirical results revealed key trade-offs. Least Connections minimized median response time under uniform workloads, but under high concurrency bursts, connection-polling overhead caused tail latency to spike. Weighted distribution strategies proved more resilient against queue saturation.

Tail latency comparison across multiple distribution algorithms Figure 8: Tail latency comparison across multiple distribution algorithms.

Cumulative distribution function (CDF) curve of request response times Figure 9: Cumulative distribution function (CDF) curve of request response times.

The final optimized configuration achieved a 73% reduction in P99 tail latency during simulated rush-hour scenarios, maintaining responsiveness across concurrent access spikes.

What I Learned

Real system performance is rarely just about adding raw compute. The mechanics of queue distribution, request offloading, and scheduling overhead dictate how an architecture responds under real load constraints.