- Published on
Load Balancing and Sticky Sessions: A Complete Course
- Authors

- Name
- Mehdi Akiki
Load Balancing and Sticky Sessions: A Complete Course
Build a production-ready load balancer with sticky sessions from scratch in Go. This course covers everything from basic concepts to advanced implementations.

GitHub Repository: mehdiakiki/load-balancer-sticky-sessions
What You'll Learn
By the end of this course, you'll understand:
- How load balancers work internally
- Why sticky sessions are critical for stateful applications
- How to implement WebSocket support
- Production-ready features: rate limiting, metrics, health checking
- When to use different load balancing algorithms
Chapter 1: What is Load Balancing?
The Problem
Imagine you have a popular web application. As traffic grows, a single server can't handle all the requests. You add more servers, but now you need a way to distribute traffic across them.
Without load balancing:
1,000 users → Single server → Server crashes
With load balancing:
1,000 users → Load Balancer → 3 servers → Each handles ~333 users ✅
Why Load Balancing Matters
- Scalability - Handle more requests by adding servers
- Reliability - If one server fails, others take over
- Performance - Distribute load to prevent overload
Types of Load Balancing
Layer 4 (Transport Layer):
- Routes based on IP and port
- Faster (no packet inspection)
- Examples: HAProxy, AWS Network Load Balancer
Layer 7 (Application Layer):
- Routes based on HTTP path, headers, cookies
- More intelligent routing
- Examples: Nginx, AWS Application Load Balancer
This course implements Layer 7 load balancing in Go.
Chapter 2: Load Balancing Algorithms
Round-Robin (Simplest)
Distributes requests sequentially:
// pkg/loadbalancer/loadbalancer.go
func (lb *LoadBalancer) getNextBackend() *backend.Backend {
lb.mux.Lock()
defer lb.mux.Unlock()
next := int(lb.Current+1) % len(lb.Backends)
lb.Current++
return lb.Backends[next]
}
Flow:
Request 1 → Backend 1
Request 2 → Backend 2
Request 3 → Backend 3
Request 4 → Backend 1 (cycle repeats)
Pros:
- Simple to implement
- Fair distribution
- O(1) time complexity
Cons:
- Doesn't account for server capacity
- Doesn't consider current load
Weighted Round-Robin (Better)
Distributes based on server capacity:
func (lb *LoadBalancer) getWeightedRoundRobinBackend() *backend.Backend {
var selectedBackend *backend.Backend
var totalWeight int
for _, b := range lb.Backends {
if !b.IsAlive() {
continue
}
b.CurrentWeight += b.Weight
totalWeight += b.Weight
if selectedBackend == nil || b.CurrentWeight > selectedBackend.CurrentWeight {
selectedBackend = b
}
}
if selectedBackend != nil {
selectedBackend.CurrentWeight -= totalWeight
}
return selectedBackend
}
Example:
Backend 1 (weight=5): Gets ~50% traffic
Backend 2 (weight=3): Gets ~30% traffic
Backend 3 (weight=2): Gets ~20% traffic
When to use:
- Heterogeneous servers (different CPU/RAM)
- Geographic distribution
- Gradual rollout (v2 gets higher weight)
Least Connections (Advanced)
Routes to server with fewest active connections:
func (lb *LoadBalancer) getLeastConnectionsBackend() *backend.Backend {
var selected *backend.Backend
minConnections := math.MaxInt64
for _, b := range lb.Backends {
if !b.IsAlive() {
continue
}
if b.ActiveConnections < minConnections {
minConnections = b.ActiveConnections
selected = b
}
}
return selected
}
Pros:
- Considers actual server load
- Better for heterogeneous traffic
Cons:
- Requires tracking connections
- More complex implementation
Chapter 3: Sticky Sessions (Session Affinity)
The Problem Stateful Applications Face
Many applications store session state locally:
// User logs in on Server 1
User → Server 1
Server 1: Creates session in memory { userID: 123, preferences: {...} }
// Next request goes to Server 2 (round-robin)
User → Server 2
Server 2: "I don't have this session!"
User gets logged out ❌
Stateful applications:
- In-memory session storage
- WebSocket connections
- Local caches (Redis on same server)
- File uploads in progress
What Are Sticky Sessions?
Sticky sessions ensure a client's requests always go to the same backend server:
First Request:
Client → LB (no cookie)
LB: Pick Backend 2 (round-robin)
LB: Set cookie: LB_SESSION=abc123 → Backend 2
All Future Requests:
Client → LB (with cookie: LB_SESSION=abc123)
LB: Look up "abc123" → Backend 2
LB: Route to Backend 2
Session stays intact ✅
Implementation
Session Storage (In-Memory):
type StickySession struct {
BackendID string // Which backend owns this session
ExpiresAt time.Time // When session expires
}
type LoadBalancer struct {
StickySession map[string]*StickySession
// ↑ sessionID ↑ session data
}
Session Creation:
func (lb *LoadBalancer) setStickySession(w http.ResponseWriter, r *http.Request, backendID string) {
// Generate unique session ID
sessionID := generateSessionID(r)
// Store in memory map
lb.StickySession[sessionID] = &StickySession{
BackendID: backendID,
ExpiresAt: time.Now().Add(lb.sessionTTL),
}
// Set cookie
http.SetCookie(w, &http.Cookie{
Name: "LB_SESSION",
Value: sessionID,
HttpOnly: true,
MaxAge: int(lb.sessionTTL.Seconds()),
})
}
Session Lookup:
func (lb *LoadBalancer) getStickyBackend(r *http.Request) *backend.Backend {
// Extract cookie
cookie, err := r.Cookie("LB_SESSION")
if err != nil {
return nil // New session
}
// Look up session in memory
lb.mux.RLock()
session, exists := lb.StickySession[cookie.Value]
lb.mux.RUnlock()
if !exists {
return nil // Invalid session
}
// Check expiration
if time.Now().After(session.ExpiresAt) {
delete(lb.StickySession, cookie.Value)
return nil // Expired
}
// Find backend
backend := lb.getBackendByID(session.BackendID)
if backend != nil && backend.IsAlive() {
return backend // ✅ Sticky backend found
}
return nil
}
Request Flow
func (lb *LoadBalancer) ServeHTTP(w http.ResponseWriter, r *http.Request) {
// Step 1: Check for sticky session
targetBackend := lb.getStickyBackend(r)
// Step 2: If no session, pick backend using round-robin
if targetBackend == nil {
targetBackend = lb.getNextBackend()
// Create new sticky session
lb.setStickySession(w, r, targetBackend.ServerID)
}
// Step 3: Route to backend
targetBackend.ReverseProxy.ServeHTTP(w, r)
}
Do You Need External Storage?
NO! For single load balancer instances, in-memory session storage is perfect:
Memory footprint:
Session ID: ~44 bytes
Backend ID: ~10 bytes
Timestamp: 8 bytes
Total: ~62 bytes per session
10,000 concurrent sessions = ~600KB
1,000,000 concurrent sessions = ~60MB
When would you need Redis/Database?
Only for multiple load balancer instances:
Problem: Multiple LBs, different session maps
LB-1: sessions["abc123"] = "backend-1"
LB-2: sessions["abc123"] = NOT FOUND ❌
Solution: Shared session storage
LB-1: sessions["abc123"] = "backend-1" → Redis
LB-2: reads from Redis → "backend-1" ✅
Chapter 4: WebSocket Support
Why WebSockets Need Sticky Sessions
WebSockets are long-lived, stateful connections. Once established:
- Connection stays open indefinitely
- Server maintains connection-specific state
- Connection cannot be "transferred" between servers
Without sticky sessions:
// WebSocket connection established on Backend 1
const ws = new WebSocket("ws://loadbalancer.com/ws");
// → Connected to Backend 1
// After 5 minutes, WebSocket message
ws.send("Hello");
// Load balancer routes to Backend 2 (round-robin)
// Backend 2: "I don't have this connection!" ❌
// Connection lost!
With sticky sessions:
// WebSocket connection
const ws = new WebSocket("ws://loadbalancer.com/ws");
// LB: Set cookie LB_SESSION=abc123 → Backend 1
// Connected to Backend 1 ✅
// After 5 minutes
ws.send("Hello");
// LB: Cookie abc123 → Backend 1
// Backend 1: "I have this connection!"
// Message delivered ✅
Implementation
Go's ReverseProxy handles WebSockets automatically:
import "net/http/httputil"
backend := &Backend{
ReverseProxy: httputil.NewSingleHostReverseProxy(url),
}
// When client sends:
// GET /ws HTTP/1.1
// Upgrade: websocket
// Connection: Upgrade
// ReverseProxy:
// 1. Detects Upgrade header
// 2. Hijacks connection
// 3. Upgrades to WebSocket on backend
// 4. Streams frames bidirectionally
No code changes needed! Sticky sessions handle the routing, ReverseProxy handles the WebSocket protocol.
Testing WebSocket Stickiness
# Start WebSocket backends
go run cmd/backend-ws/main.go -port 8081 -name ws-backend-1
go run cmd/backend-ws/main.go -port 8082 -name ws-backend-2
go run cmd/backend-ws/main.go -port 8083 -name ws-backend-3
# Start load balancer
go run cmd/server/main.go -config configs/config.toml
# Connect WebSocket
wscat -c ws://localhost:8080/ws
# Send messages
> Hello
< [ws-backend-1] Echo: Hello
> World
< [ws-backend-1] Echo: World
> Test
< [ws-backend-1] Echo: Test
All messages go to the SAME backend! Check backend logs:
[ws-backend-1] WebSocket connection established
[ws-backend-1] Received: Hello
[ws-backend-1] Received: World
[ws-backend-1] Received: Test
Chapter 5: Health Checking
Why Health Checking Matters
Backends can fail:
- Server crashes
- Network issues
- Application errors
Without health checking, load balancer would route to dead servers.
Implementation
Backend Health Check:
func (b *Backend) HealthCheck(timeout time.Duration) error {
client := http.Client{Timeout: timeout}
resp, err := client.Get(b.URL.String())
if err != nil {
b.SetAlive(false)
return err
}
defer resp.Body.Close()
// Consider 2xx and 3xx as healthy
alive := resp.StatusCode >= 200 && resp.StatusCode < 500
b.SetAlive(alive)
return nil
}
Background Health Monitoring:
func (lb *LoadBalancer) HealthCheck(interval, timeout time.Duration) {
// One goroutine per backend
for _, b := range lb.Backends {
go func(backend *Backend) {
ticker := time.NewTicker(interval)
for range ticker.C {
backend.HealthCheck(timeout)
}
}(b)
}
}
Failover Behavior:
When backend fails:
- Health check marks it as
Alive = false - New requests skip this backend
- Existing sticky sessions become invalid
- Clients get new backend on next request
Chapter 6: Rate Limiting
Token Bucket Algorithm
Allow bursts while maintaining average rate:
type RateLimiter struct {
limiters map[string]*rate.Limiter
rate rate.Limit // Tokens per second
burst int // Max burst
}
func (rl *RateLimiter) Middleware(next http.Handler) http.Handler {
return http.HandlerFunc(func(w http.ResponseWriter, r *http.Request) {
// Use client IP as key
key := r.RemoteAddr
// Get or create limiter for this client
limiter := rl.getLimiter(key)
// Try to consume a token
if !limiter.Allow() {
http.Error(w, "Rate limit exceeded", http.StatusTooManyRequests)
return
}
next.ServeHTTP(w, r)
})
}
Configuration:
[ratelimit]
enabled = true
requests_per_second = 100.0
burst = 200
cleanup_interval = "1m"
How it works:
- Tokens replenish at 100/second
- Bucket holds max 200 tokens
- Request consumes 1 token
- Empty bucket → HTTP 429
Chapter 7: Configuration
TOML Configuration
Easy to read and modify:
[server]
port = 8080
host = "0.0.0.0"
[loadbalancer]
algorithm = "weighted-round-robin"
session_ttl = "30m"
[[loadbalancer.backends]]
id = "backend-1"
url = "http://localhost:8081"
weight = 3
[[loadbalancer.backends]]
id = "backend-2"
url = "http://localhost:8082"
weight = 2
[logging]
level = "info"
format = "json"
[metrics]
enabled = true
port = 9090
Configuration Validation
func (c *Config) Validate() error {
if c.Server.Port < 1 || c.Server.Port > 65535 {
return fmt.Errorf("invalid port: %d", c.Server.Port)
}
if len(c.LoadBalancer.Backends) == 0 {
return fmt.Errorf("no backends configured")
}
for _, backend := range c.LoadBalancer.Backends {
if backend.Weight < 0 {
return fmt.Errorf("negative weight")
}
}
return nil
}
Chapter 8: Production Features
Prometheus Metrics
var (
TotalRequests = prometheus.NewCounter(
prometheus.CounterOpts{
Name: "lb_total_requests_total",
Help: "Total requests processed",
},
)
BackendRequests = prometheus.NewCounterVec(
prometheus.CounterOpts{
Name: "lb_requests_per_backend_total",
Help: "Requests per backend",
},
[]string{"backend_id"},
)
)
// Access metrics at /metrics endpoint
Key metrics:
lb_total_requests_total
lb_requests_per_backend_total{backend_id="backend-1"}
lb_session_count
lb_backend_health{backend_id="backend-1"}
lb_rate_limit_exceeded_total
Structured Logging
func (l *Logger) log(level LogLevel, message string, fields map[string]interface{}) {
entry := LogEntry{
Timestamp: time.Now().UTC().Format(time.RFC3339),
Level: string(level),
Message: message,
Fields: fields,
}
if l.format == "json" {
json.NewEncoder(l.writer).Encode(entry)
} else {
log.Printf("[%s] %s %v", level, message, fields)
}
}
JSON output:
{
"timestamp": "2024-01-15T10:30:00Z",
"level": "info",
"message": "Request processed",
"fields": {
"backend": "backend-1",
"duration_ms": 45,
"path": "/api/users"
}
}
Chapter 9: Testing
Proving Sticky Sessions Work
Test 1: Round-Robin (No Cookie)
for i in {1..6}; do
curl -s http://localhost:8080/ | grep "Hello from"
done
# Output:
# Hello from backend-1
# Hello from backend-2
# Hello from backend-3
# Hello from backend-1
# Hello from backend-2
# Hello from backend-3
Different backends — round-robin working! ✓
Test 2: Sticky Session (With Cookie)
curl -c cookies.txt http://localhost:8080/ > /dev/null
for i in {1..6}; do
curl -b cookies.txt -s http://localhost:8080/ | grep "Hello from"
done
# Output:
# Hello from backend-2
# Hello from backend-2
# Hello from backend-2
# Hello from backend-2
# Hello from backend-2
# Hello from backend-2
Same backend — sticky session working! ✓
Test 3: Multiple Clients
# Client 1
curl -c client1.txt http://localhost:8080/ > /dev/null
curl -b client1.txt http://localhost:8080/ | grep "Hello from"
# → backend-1, backend-1, backend-1...
# Client 2
curl -c client2.txt http://localhost:8080/ > /dev/null
curl -b client2.txt http://localhost:8080/ | grep "Hello from"
# → backend-3, backend-3, backend-3...
Different clients, different backends — sessions isolated! ✓
Chapter 10: Deployment Considerations
Single Instance Deployment
For most use cases, a single load balancer instance with in-memory sessions is sufficient:
# docker-compose.yml
services:
loadbalancer:
build: .
ports:
- "8080:8080"
- "9090:9090" # Metrics
volumes:
- ./configs:/config
command: ["./loadbalancer", "-config", "/config/config.toml"]
backend1:
image: yourapp:latest
ports:
- "8081:8081"
backend2:
image: yourapp:latest
ports:
- "8082:8082"
High Availability Deployment
For production, run multiple LB instances with shared session storage:
services:
lb1:
image: loadbalancer:latest
environment:
- REDIS_URL=redis://redis:6379
lb2:
image: loadbalancer:latest
environment:
- REDIS_URL=redis://redis:6379
redis:
image: redis:alpine
volumes:
- redis-data:/data
nginx:
image: nginx:alpine
ports:
- "80:80"
# Routes to lb1 or lb2
Performance Considerations
In-memory session storage:
- Lookup time: ~10ns
- Memory: 62 bytes per session
- 1M sessions ≈ 60MB RAM
Health checking:
- Interval: 10 seconds
- Timeout: 2 seconds
- Each check: ~5ms
Rate limiting:
- Overhead: ~50ns per request
- Memory: ~1KB per client
Complete Implementation
Get the full source code:
GitHub: mehdiakiki/load-balancer-sticky-sessions
Project Structure
loadbalancer-sticky-sessions/
├── cmd/
│ ├── server/main.go # Load balancer entry point
│ ├── backend/main.go # HTTP test backend
│ └── backend-ws/main.go # WebSocket test backend
├── pkg/
│ ├── backend/backend.go # Backend management
│ ├── config/config.go # TOML configuration
│ ├── loadbalancer/ # Core load balancer logic
│ ├── logging/logging.go # Structured logging
│ ├── metrics/metrics.go # Prometheus metrics
│ └── ratelimit/ratelimit.go # Rate limiting
└── configs/config.toml # Example configuration
Quick Start
# Clone repository
git clone https://github.com/mehdiakiki/load-balancer-sticky-sessions.git
cd load-balancer-sticky-sessions
# Start backends
go run cmd/backend/main.go -port 8081 -name backend-1 &
go run cmd/backend/main.go -port 8082 -name backend-2 &
go run cmd/backend/main.go -port 8083 -name backend-3 &
# Start load balancer
go run cmd/server/main.go -config configs/config.toml
# Test
curl http://localhost:8080/
Course Summary
What You Learned
✅ Load Balancing Fundamentals
- Round-robin and weighted round-robin algorithms
- When to use each algorithm
- Performance characteristics
✅ Sticky Sessions
- Why stateful applications need them
- Cookie-based implementation
- In-memory storage (no Redis needed!)
- Session lifecycle management
✅ WebSocket Support
- Why WebSockets require sticky sessions
- How Go's ReverseProxy handles upgrades
- Testing WebSocket stickiness
✅ Production Features
- Health checking with automatic failover
- Rate limiting (token bucket)
- Prometheus metrics
- Structured logging
- TOML configuration
✅ Testing and Deployment
- How to verify sticky sessions work
- Single-instance vs HA deployment
- Performance tuning
Next Steps
- Study the code - Read through the implementation
- Modify and experiment - Try different algorithms
- Add features - Implement IP-hash, least connections
- Deploy - Use it for learning projects
Resources
GitHub Repository: mehdiakiki/load-balancer-sticky-sessions
Related Topics:
- Nginx Load Balancing
- HAProxy Configuration
- Kubernetes Services and Ingress
- Service Mesh (Istio, Linkerd)
Course Exercises
Exercise 1: Implement IP-Hash
Add IP-hash algorithm to route based on client IP:
func (lb *LoadBalancer) getIPHashBackend(r *http.Request) *backend.Backend {
ip := r.RemoteAddr
hash := fnv.New32a()
hash.Write([]byte(ip))
index := int(hash.Sum32()) % len(lb.Backends)
return lb.Backends[index]
}
Exercise 2: Add Least Connections
Track active connections per backend:
type Backend struct {
// ... existing fields
ActiveConnections int
mux sync.Mutex
}
func (b *Backend) IncrementConnections() {
b.mux.Lock()
b.ActiveConnections++
b.mux.Unlock()
}
Exercise 3: Implement Circuit Breaker
Add circuit breaker pattern to prevent cascading failures:
type CircuitBreaker struct {
maxFailures int
timeout time.Duration
state State // Closed, Open, HalfOpen
failureCount int
lastFailTime time.Time
}
Exercise 4: Add HTTPS Support
Configure TLS in the config:
server := &http.Server{
Addr: ":443",
Handler: lb,
}
server.ListenAndServeTLS("cert.pem", "key.pem")
Conclusion
Load balancing and sticky sessions are fundamental to modern web architecture. By building one from scratch, you understand:
- How requests are distributed
- Why certain algorithms work better for different scenarios
- When sticky sessions are necessary
- What production concerns exist (health checks, rate limiting, metrics)
The complete implementation is available on GitHub with comprehensive documentation, tests, and deployment examples. Use it as a learning tool or as a foundation for more complex systems.
I build and scale reliable production systems. Open to full-time and freelance work with U.S.-based teams that value ownership and execution.
Got something in mind?
Book a Discovery Call