Scaling Without Downtime; Why Is This Critical?
When traffic to your website or service grows, the first thought that comes to mind is to increase server resources. But if you do this incorrectly, instead of improving performance, you will face service downtime. Scaling in the cloud, if not done with proper planning, can easily turn into a nightmare—especially when real users are actively using the service.
In this article, we intend to practically and without digression examine the two main scaling methods: Vertical Scaling, which means strengthening the same existing server, and Horizontal Scaling, which involves adding new nodes to the infrastructure. For both methods, we will provide specific instructions, real-world examples, and troubleshooting tips so you can upgrade your infrastructure without even a second of downtime.
Vertical Scaling; The Fastest Way But with Serious Limitations
Vertical scaling means increasing the resources of a single server: more RAM, a more powerful CPU, or a faster disk. In the cloud, this is usually done by changing the Flavor or Instance Type. For example, if you have a server with 4 GB of RAM and 2 CPU cores in a cloud panel, you can upgrade it to 8 GB of RAM and 4 cores.
Steps for Vertical Scaling Without Downtime
Many cloud management panels offer the ability to resize a server online (Live Resize). However, this feature is not always available, and in some cases, it requires a restart. To do this without downtime, follow these steps:
- Check Live Resize Support: First, review your panel's documentation. If your panel uses OpenStack, you can try the
openstack server resizecommand. But note that this command may require a restart in some distributions. - Shift Load to a Temporary Server: If Live Resize is not supported, the best approach is to create a new server with higher resources, migrate the data, and then switch traffic to it. This is done using DNS and reducing TTL.
- Use Snapshot: Before any action, take a Snapshot of the server. This allows you to revert to the previous state if any issues arise.
Limitations of Vertical Scaling
Vertical scaling has a definite ceiling. You cannot scale a server infinitely. In any cloud infrastructure, the maximum Instance size is limited. For example, if the maximum allocatable RAM is 64 GB, there is no way to scale vertically beyond that. Additionally, vertical scaling creates a Single Point of Failure; if that server encounters a hardware issue, the entire service is lost.
Common Mistake: Many users assume that increasing RAM alone will solve website slowness. However, if the issue stems from heavy database queries or disk I/O limitations, increasing RAM will have no effect. Before vertical scaling, be sure to check with top, htop, and iostat to determine which resource is actually saturated.
Horizontal Scaling; The Long-Term Solution for Sustainable Growth
Horizontal scaling means adding new nodes to the infrastructure. Instead of one large server, you place several smaller servers together and distribute the traffic load among them. This method not only overcomes the limitations of vertical scaling but also increases the reliability of the infrastructure; if one node fails, the remaining nodes continue to provide the service.
Preparing Your Application for Horizontal Scaling
The most important part of horizontal scaling is preparing your application. If your application is stateful (i.e., it stores user information on the server itself), horizontal scaling is not easily possible. To solve this problem, you need to make the following changes:
- Move Sessions to Shared Storage: Instead of storing sessions in files or local memory, use Redis or Memcached. For example, in PHP, you can move sessions to Redis by setting
session.save_handler = redisin thephp.inifile. - Upload Files to Shared Storage: Do not store user-uploaded files on the local disk. Use Object Storage such as S3 or similar services.
- Manage Configurations: Place application configurations in Environment Variables so that each node can easily receive the settings.
Adding a New Node to the Infrastructure
After preparing the application, it's time to add new nodes. Follow these steps:
- Create an Image from the Base Server: Create an Image from the main server where the application and initial settings are installed. You can store this Image in your cloud panel.
- Launch a New Node: Create a new Instance from the Image. Make sure the new node is on the same internal network (VPC) so it can communicate with the database and Redis.
- Connect to Load Balancer: Add the new node to the Load Balancer. If you are using Nginx as the Load Balancer, simply add the new node's IP address to the upstream in the configuration file:
upstream backend {
server 10.0.0.11:8080 weight=3;
server 10.0.0.12:8080 weight=3;
server 10.0.0.13:8080 weight=3; # New node
}
server {
listen 80;
location / {
proxy_pass http://backend;
proxy_set_header Host $host;
proxy_set_header X-Real-IP $remote_addr;
}
}
After applying the changes, reload Nginx with the command nginx -s reload. This is done without service downtime, and traffic is gradually distributed among the nodes.
Database Management in Horizontal Scaling
The database is usually the most challenging part of horizontal scaling. If your application uses MySQL, you can use Replication: one primary server (Master) for writes and several read servers (Replicas). In this case, the application must differentiate between these two types of servers. For example, in the database connection configuration in Laravel, you can do it like this:
'mysql' => [
'read' => [
'host' => ['10.0.0.21', '10.0.0.22'],
],
'write' => [
'host' => ['10.0.0.20'],
],
'driver' => 'mysql',
'database' => 'app_db',
'username' => 'app_user',
'password' => 'secret',
'charset' => 'utf8mb4',
]
To set up Replication in MySQL, first edit the configuration file on the primary server:
[mysqld]
server-id = 1
log_bin = /var/log/mysql/mysql-bin.log
binlog_do_db = app_db
Then apply the following settings on the Replica server:
[mysqld]
server-id = 2
relay-log = /var/log/mysql/mysql-relay-bin.log
Finally, start Replication with the following commands:
CHANGE MASTER TO MASTER_HOST='10.0.0.20', MASTER_USER='replica_user', MASTER_PASSWORD='password', MASTER_LOG_FILE='mysql-bin.000001', MASTER_LOG_POS= 107;
START SLAVE;
Common Scaling Mistakes and Their Solutions
Over the years of working with cloud infrastructures, I have encountered repeated mistakes that have led to service downtime. Here are a few important ones:
- Changing Multiple Variables Simultaneously: If you perform vertical scaling and change the application code at the same time, you won't know which change caused the issue. Always apply one change at a time.
- Forgetting Health Checks: When adding a new node to the Load Balancer, be sure to enable Health Checks. If the new node is not properly set up, the Load Balancer should automatically remove it from the rotation. In Nginx, you can use the
nginx_upstream_check_modulemodule. - Ignoring Network Capacity: Horizontal scaling may increase internal network traffic. If you are using a network with limited bandwidth, you may encounter network bottlenecks. Before adding many nodes, check the internal network bandwidth.
Conclusion; A Smart Scaling Strategy
There is no one-size-fits-all solution for scaling. The best strategy is to start with vertical scaling until you reach its ceiling, then move to horizontal scaling. But do this wisely:
- First, identify bottlenecks using monitoring tools like Prometheus and Grafana.
- Design the application for horizontal scaling from the start; even if you don't need it right now.
- Always take Snapshots and practice the Rollback process.
- Apply changes during low-traffic hours and use Load Balancer tools for gradual traffic distribution.
Finally, if you are looking for an infrastructure that enables easy and hassle-free scaling, ServerNet cloud services can be a suitable option; but the most important point is to properly plan the scaling process and know your infrastructure well before any action. By following the tips in this article, you can upgrade your infrastructure without even a second of downtime and maintain a flawless user experience.