Why do you need a load balancer and where to start?
When your website traffic exceeds the capacity of a single server or you need high availability, the first solution that comes to mind is adding more servers. However, additional servers without a load balancer not only fail to solve the problem but also, due to unmanaged request distribution among them, cause some servers to be overloaded while others remain idle. The load balancer distributes incoming requests among several backend servers, checks their health, and if one fails, routes traffic to healthy servers.
In this article, we will set up an operational load balancer with Nginx. Nginx is a common choice for this task due to its low resource consumption and high performance. We assume you have two backend servers with addresses 10.0.0.11 and 10.0.0.12 running a web application on port 8080. The load balancer is installed on a separate server with address 10.0.0.10.
Installing and basic configuration of the load balancer
First, install Nginx on the load balancer server. On Debian/Ubuntu-based distributions:
sudo apt update
sudo apt install nginx -y
Then edit the main configuration file. It is better to create a separate file in the /etc/nginx/sites-available/ directory and link it to sites-enabled:
sudo nano /etc/nginx/sites-available/loadbalancer
Initial content for simple round-robin distribution:
upstream backend_servers {
server 10.0.0.11:8080;
server 10.0.0.12:8080;
}
server {
listen 80;
server_name example.com;
location / {
proxy_pass http://backend_servers;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
This configuration distributes requests sequentially between the two servers. Be sure to add the X-Forwarded-* headers so that your application sees the user's real IP address, not the load balancer's address. Otherwise, your application logs and traffic analysis tools will show incorrect information.
Save the file and create the symbolic link:
sudo ln -s /etc/nginx/sites-available/loadbalancer /etc/nginx/sites-enabled/
sudo nginx -t
If the output is syntax is ok, restart the service:
sudo systemctl restart nginx
Choosing the appropriate distribution algorithm
The default round-robin algorithm is suitable for most applications, but it is not always the best choice. If your servers have different processing capabilities, use weight:
upstream backend_servers {
server 10.0.0.11:8080 weight=3;
server 10.0.0.12:8080 weight=1;
}
This setting means that out of every 4 requests, 3 are sent to the first server and 1 to the second. If your application is session-based and you do not want users to be moved between servers, use ip_hash:
upstream backend_servers {
ip_hash;
server 10.0.0.11:8080;
server 10.0.0.12:8080;
}
With ip_hash, each user's requests are always sent to a specific server. This is the simplest way to manage sessions, but if the user uses a variable IP (such as on mobile), it loses its effectiveness. In that case, you should implement session stickiness at the application level, not the load balancer.
Configuring health checks to detect failed servers
The problem with the above configuration is that if one of the backend servers goes down, Nginx will still send requests to it, and users will receive 502 errors. To solve this problem, you need to enable health checks. Nginx has a passive health check by default that only checks the server at the moment of connection. However, for periodic and more accurate checks, you need to use the commercial module or alternative solutions.
The simplest free solution is to use max_fails and fail_timeout:
upstream backend_servers {
server 10.0.0.11:8080 max_fails=3 fail_timeout=30s;
server 10.0.0.12:8080 max_fails=3 fail_timeout=30s;
}
This setting means that if connecting to a server fails 3 times within a 30-second period, that server is removed from the rotation for 30 seconds. However, this method only detects connection errors, not application errors. If your application returns an HTTP 500 response, the connection is established, and Nginx considers it healthy.
For a real health check that examines the application's response, you can use the nginx-upstream-check-module or write an external script. A simpler approach is to use a dedicated endpoint in the application:
location /health {
proxy_pass http://backend_servers;
proxy_set_header Host $host;
}
location / {
proxy_pass http://backend_servers;
proxy_next_upstream error timeout http_500 http_502 http_503;
}
With proxy_next_upstream, if the first server returns a 500 or 502 error, Nginx automatically sends the request to the next server. This method is simple and effective, but it is not sufficient for periodic health monitoring of servers.
Implementing health checks with an external script
For a more complete solution, you can write a bash script that checks the health endpoint of each server every 10 seconds and removes it from the upstream if it fails. However, this method is complex and difficult to maintain. My recommendation is to use HAProxy for real health checks, but if you want to stick with Nginx, combining proxy_next_upstream and max_fails is sufficient for most applications.
Common mistake: Many people only set max_fails and think the problem is solved. But if your application returns a 200 response with error content (such as a PHP error displayed in HTML), Nginx considers it healthy. Always define a separate health endpoint in the application that returns only the real status.
SSL termination on the load balancer
One of the most important tasks of a load balancer is SSL termination. This means that HTTPS traffic from the user is encrypted up to the load balancer, but between the load balancer and backend servers, it is transmitted over HTTP. This removes the encryption load from application servers and simplifies certificate management.
First, install the SSL certificate on the load balancer server. If you do not have a certificate, get a free one with Let's Encrypt:
sudo apt install certbot python3-certbot-nginx -y
sudo certbot --nginx -d example.com -d www.example.com
Then configure the load balancer for SSL:
server {
listen 443 ssl http2;
server_name example.com;
ssl_certificate /etc/letsencrypt/live/example.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/example.com/privkey.pem;
ssl_protocols TLSv1.2 TLSv1.3;
ssl_ciphers HIGH:!aNULL:!MD5;
location / {
proxy_pass http://backend_servers;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
proxy_set_header X-Forwarded-For $proxy_add_x_forwarded_for;
proxy_set_header X-Forwarded-Proto $scheme;
}
}
server {
listen 80;
server_name example.com;
return 301 https://$host$request_uri;
}
Important note: Be sure to set the X-Forwarded-Proto header. Your application needs to know that the original request came over HTTPS, even if HTTP is used between the load balancer and backend. If you do not set this header, the application may generate HTTP links or create insecure redirects.
Optimizing SSL for better performance
To reduce handshake time and improve performance, you can enable session cache:
ssl_session_cache shared:SSL:10m;
ssl_session_timeout 10m;
This setting allows Nginx to share SSL sessions between different connections. The value 10m means about 10 megabytes of memory for session caching, which is sufficient for approximately 40,000 sessions.
Also, if all your backend servers are on a secure private network, you can use plain HTTP between the load balancer and backend. However, if your network is not trustworthy, you should also enable SSL between the load balancer and backend. In that case, use proxy_pass https:// instead of proxy_pass http:// and install certificates on each backend server.
Common errors and how to fix them
Below are several common errors you may encounter when setting up a load balancer and their solutions:
502 Bad Gateway error
This error means Nginx cannot connect to the backend server. First, check that the service is running on the backend server:
curl -I http://10.0.0.11:8080
If you do not receive a response, check the firewall on the backend server. Port 8080 must be accessible from the load balancer's address:
sudo ufw allow from 10.0.0.10 to any port 8080
504 Gateway Timeout error
This error means the backend server did not return a response within the specified time. The default Nginx timeout is 60 seconds. If your application has long processing times, increase this value:
location / {
proxy_pass http://backend_servers;
proxy_read_timeout 120s;
proxy_connect_timeout 10s;
}
Session and login issues
If your users are redirected to the home page after logging in or their sessions keep dropping, the problem is with request distribution among different servers. Use ip_hash or implement session stickiness in the application. Also, make sure the X-Forwarded-For header is correctly set, as some applications use this header to identify users.
Summary and next steps
In this article, we set up an operational load balancer with Nginx, added backend servers, configured health checks, and terminated SSL on the balancer. These basic settings are sufficient for most production applications, but if your traffic is very high or you need more advanced features such as rate limiting and caching, you can use HAProxy or managed load balancer services.
If your infrastructure runs on cloud servers, ServerNet provides the ability to set up a dedicated load balancer or use managed services, which can significantly reduce setup time. However, in any case, understanding the concepts covered in this article will help you make better decisions and troubleshoot your infrastructure issues faster.
Finally, always test load balancer settings in a staging environment and set up proper monitoring for backend servers and the load balancer itself. A load balancer without monitoring is just a new point of failure.