Chat with us, powered by LiveChat
Home > Backup and Recovery Blog > Scalable Backup: How to Build Data Protection That Keeps Working as Infrastructure Grows
Updated 8th September 2026, Rob Morrison

Scalable backup is the ability to maintain backup and recovery performance as data, workloads, and retention requirements grow.

A backup platform that works well for 100 terabytes (TB) may behave very differently at 1 petabyte (PB).

As backup environments grow, scalability issues may arise. Such issues include backup windows with longer business hours, longer recovery searches, more administrative work.

Scalability has to be evaluated across the entire backup and recovery architecture, not just by looking at storage capacity. As your data environment grows, the system will need to support larger data volumes, more workloads, longer retention periods, and heavier recovery demands without creating excessive cost, complexity, or operational risk. It also needs to continue meeting the organization’s recovery point objectives (RPOs) and recovery time objectives (RTOs).

What Does Scalability Mean in a Backup Environment?

Backup scalability is the ability of a backup environment to protect and recover increasing amounts of data and workloads while continuing to meet defined performance, cost, security, RPO, and RTO requirements.

Backup environments do not grow in just one way. Data volume may increase, but so can the number of files, workloads, locations, and retention requirements. Each puts pressure on a different part of the backup infrastructure.

  • More data: More servers, databases, applications, and users increase protected capacity.
  • More files and objects: Billions of files can put considerable pressure on backup catalogs and metadata systems.
  • Higher change rates: More data changing every day means more information must be backed up within the same window.
  • More workloads: Virtual machines, databases, containers, SaaS(software as a service) applications, and cloud accounts all introduce different protection requirements.
  • Longer retention: Regulatory requirements and business policies may require backup data to be retained for months or years.

Why Backup Capacity and Backup Scalability Are Not the Same

Backup capacity is the amount of data a system can store and hold. “Backup infrastructure growth refers to the ability of the entire data protection environment to handle increasing workloads while maintaining performance, reliability, security, and recovery objectives.

As the backup environment grows, storage capacity is only one consideration. From an operational performance perspective, higher data volumes can increase backup throughput requirements, metadata and catalog processing, replication traffic, and retention costs. If these components do not scale with storage, routine backup operations can become slower and more expensive.

Growth implications for recovery. The more data you have, the longer it can take to restore, especially if you need to restore multiple systems at the same time. Even with successful backup jobs, an organization may have difficulty meeting its RTO if it does not have sufficient restore throughput and recovery capacity.

A scalable backup architecture therefore needs to accommodate growth across several interconnected areas:

  • The components that move and store backup data
  • The catalog and indexing components used to locate recovery points
  • The components that manage policies, scheduling, monitoring, and resources
  • Recovery plan

Backup performance at scale should be measured against ongoing protection needs as well as recovery needs. A scalable architecture must be able to maintain defined RPOs and RTOs as data volume, number of workloads, retention, geographic distribution and concurrent recovery demand increase.

Which Bottlenecks Appear First as Backup Environments Grow?

The first backup scalability bottlenecks typically appear in backup windows, network throughput, metadata services, replication, repository performance, or restore infrastructure. Which one appears first depends on workload characteristics and architecture.

Backup windows collide with production activity

As the amount of data changing per day increases, the backup jobs could take up a larger percentage of the backup window available. Scheduled backups can eventually fall into business hours, or run over into the next protection cycle.

Catalog and metadata growth slows recovery

As systems, recovery points, files and objects increase, the backup catalog becomes an increasingly important architecture component.

The catalog must support continuous updates, search, browse, report, and recovery operations. If you have long retention periods and many objects, a poorly designed metadata infrastructure can make it difficult to find the right recovery point quickly.

Small-file workloads create disproportionate overhead

The number and size of files can have a significant effect on backup performance. A single large database file and millions of small files may occupy the same amount of storage, but small-file workloads typically create much more processing overhead. File enumeration, metadata collection, permission checks, indexing, and file recreation can all increase backup and restore times.

Performance testing should therefore reflect the real characteristics of the workload rather than data volume alone.

Architecture choices may also need to consider small-file density. Depending on the environment, organizations can consider strategies such as file-system snapshots, image-level backups, object storage or grouping and archiving large collections of small files before protection.

Small-file workloads overwhelm metadata operations

Modern backup architectures commonly maintain multiple copies of protected data for resilience and recovery. Each additional copy affects capacity planning, and backup data that cannot be modified or deleted during a specified retention period can make that impact more significant because protected data cannot be removed before its retention period expires.

Growth planning should account for primary and secondary copies, replication, immutable retention, legal holds which can prevent backup data from expiring according to its normal lifecycle, metadata, temporary recovery space, and other overheads associated with operating the environment.

Restore throughput receives less investment than ingest

Backup ingest receives constant attention because backup jobs run regularly. Large-scale recovery is tested less frequently, which means restore limitations can remain unnoticed until they become critical.

Recovery performance can be affected by repository read speed, backup fragmentation, deduplication and reconstructing deduplicated data during recovery processes, network capacity, cloud retrieval and egress, malware scanning, and the performance of the recovery destination.

Which Metrics Reveal Scalability Problems Early?

Capacity utilization provides only part of the picture. Monitoring should therefore cover the complete protection and recovery service. Useful metrics include backup-window utilization, queue delay, backup and restore throughput, catalog response time, replication lag, and RPO coverage among others.

Scale-Up vs. Scale-Out Backup Architecture

Scale-up expands an existing backup system with more resources. Scale-out expands capacity by adding systems or nodes and distributing workloads between them.

Scale-up adds CPU, memory, storage or other resources to an existing system. It’s generally easy to deal with, but the architecture is still subject to the maximum capacity of that system.

Scale-out adds nodes, workers, storage units or other services. This method can offer more incremental growth and potentially enable workloads to be spread across multiple systems, but it also adds more coordination and operational needs.

A Step-by-Step Process for Designing Scalable Backup

A scalable backup architecture starts with the expected growth of the environment rather than the current infrastructure alone.

1. Model growth by workload instead of using one annual percentage

Track protected capacity, daily change rate, object count, job volume, retention, number of backup copies, and available backup windows for each major workload category.

Databases, virtual machines, file servers, containers, and cloud workloads often have very different protection requirements. Forecasts should also account for planned application deployments, cloud migrations, acquisitions, regulatory changes, and infrastructure retirement.

2. Convert recovery objectives into throughput requirements

Recovery objectives should determine measurable infrastructure needs. Assign recovery requirements based on application priority, not all workloads are the same. Some critical applications may need high-performance recovery resources while less time-dependent data can use lower-cost storage and recovery paths.

3. Map the complete backup and recovery data paths

Document the full data path during restore from backup storage to production and from production to backup storage. This should include source storage, backup agents/data movers, network links, firewalls, encryption and deduplication stages, repositories, replication targets, catalog services and recovery destinations. Mapping each component helps identify where throughput may be constrained, since the slowest sustained element can limit the performance of the overall backup or restore process.

Disaster scenarios introduce an additional consideration: shared-resource contention. During recovery, network, compute, identity, and storage services may already be under heavy load or partially unavailable. Recovery planning should therefore assess whether these shared resources can support backup restoration at the same time as other critical systems are being brought back online.

4. Establish separate backup resource groups

Organizations should consider separate backup domains when geography, security boundaries, workload contention, or recovery requirements make a single shared domain difficult to operate. You can reduce contention and limit the impact of a failure or unusually demanding workload by using independent backup workers, storage pools, proxies, or other resources.

NOTE: Separating also creates management overhead. Every new domain needs monitoring, patching, credentials, capacity planning and testing.

5. Automate lifecycle and placement decisions

With the increase of protected workloads, manual configuration becomes more and more difficult. Backup schedules, retention periods, storage tiers, replication requirements, and expiration policies can be assigned via policy-based automation. Automated discovery can also help discover new systems that have been deployed without appropriate protection.

6. Add security without creating an unmeasured bottleneck

While security is an important part of modern backup architecture, security controls also add additional requirements. Capacity and performance planning should take into account encryption, immutability, multifactor authentication, malware scanning and isolated recovery environments.

Organizations should also maintain offline and regularly tested backups. CISA suggests storing backups offline and encrypted, and regularly testing the availability and integrity of backups to ensure they remain usable in the event of a ransomware incident or other recovery scenario.

7. Test growth and recovery before production reaches the limit

Testing should be based on projected requirements, not just current workload levels. Load and recovery tests should measure data volume, job concurrency, catalog performance, repository utilization, replication, failure scenarios, and recovery throughput.

Backup Scalability Checklist

When designing the backup architecture, use this checklist to see if it can meet current requirements and future growth:

  • Current protected capacity: the total amount of data currently backed up.
  • Growth rate: the annual rate at which you expect protected data to grow.
  • Daily change rate: percentage or volume of data that changes between backup cycles.
  • File or Object Count: the number of individual files, objects or records that the backup system must handle.
  • Backup window: the time period available for the execution of scheduled backups.
  • Required RPO: the maximum tolerable data loss measured in time.
  • Required RTO: the maximum acceptable amount of time to restore systems or data.
  • Throughput restore: the speed at which the backup environment can return data during a recovery.
  • Number of copies in backup: how many copies are to be maintained and where?
  • Immutable retention period: the minimum period of time where backup copies should be protected from modification or deletion.
  • Replication bandwidth: the throughput required for the network to transfer backup data between environments or locations.

How to Prove That a Backup Architecture Scales

Run load and recovery tests to demonstrate scalability of backup and recovery based on projected data volumes, concurrency, retention, failures and recovery requirements.

Vendor specifications and published throughput numbers are useful for comparison, but they do not guarantee performance in a particular environment. Actual results depend on factors such as data type and file size, change rate, compression and deduplication, encryption, network latency, source performance, repository design, and recovery infrastructure.

A meaningful proof of concept should therefore reproduce the conditions the organization expects to encounter as it grows.

Testing should also include recovery from ransomware and destructive data events. NIST SP 1800-11 provides guidance on recovering data after destructive events, with particular emphasis on validating data integrity and ensuring that recovered information can be trusted and used safely.

Why Backup Costs Often Grow Faster Than Data

Backup costs can grow faster than protected data. This is because growth can increase storage copies, retention, replication, licensing, network traffic, cloud retrieval, management infrastructure, and administrative effort at the same time.

You should come up with a realistic cost model that considers at least a three-year planning period and include:

  • Data growth and daily change rates
  • Retention, replication, immutability, and legal holds
  • Software licensing and support
  • Storage and hardware refresh requirements
  • Cloud storage, operations, retrieval, and cloud data-transfer-out (egress) charges
  • Network and inter-region transfer costs
  • Catalog and management infrastructure
  • Recovery testing and temporary recovery resources
  • Administration, monitoring, patching, and incident response
  • Migration or platform-exit costs

How Bacula Enterprise Supports Scalable Backup

Bacula Enterprise stands out as an especially scalable backup and recovery solution that offers realistic answers to the scale-related issues discussed here. It uses a modular architecture, so backup management, client services, storage, and catalog components can be deployed separately based on the needs of the environment.

The platform supports a broad spectrum of workload categories including physical and virtual systems, databases, containers and cloud environments. This makes it ideal for organizations that need to protect multiple types of infrastructure under a single backup strategy.

In terms of scalability, the modular design can allow backup processing, storage and associated services to be distributed as the volume of data and workload requirements grow. Therefore, for organizations that are growing quickly, this architecture can be considered as part of a larger evaluation of how the backup infrastructure can scale over time.

Bacula’s architecture aids scale-up and scale-out in many different ways. Just some examples of this would be:

Massive data volumes: Scales from conventional enterprise environments to multi-petabyte HPC and AI infrastructures.

Billions of files: HPCAccelerator distributes file-system workloads across concurrent workers, enabling efficient protection of billions of files.

Parallel processing: Multiple backup and restore streams can run simultaneously, substantially increasing throughput.

High job concurrency: Supports thousands of simultaneous backup jobs across large and complex environments.

Distributed architecture: Backup components can be distributed across multiple servers, networks and locations to avoid reliance on a single processing bottleneck.

Parallel file-system integration: Dedicated capabilities for Lustre and IBM Storage Scale/GPFS help protect extremely large, high-performance file systems efficiently.

Efficient incremental backups: Technologies such as Lustre Changelog and BSnapDiff identify changed data without repeatedly scanning an entire massive file system.

Storage scalability: Customers can expand across disk, tape, object storage and cloud resources without being tied to a single storage vendor or technology.

Workload scalability: One platform can protect physical servers, virtual machines, containers, databases, cloud workloads and HPC environments as the infrastructure expands.

Economic scalability: Bacula does not charge according to protected data volume, so rapidly growing datasets do not automatically produce rapidly growing licence costs.

These capabilities make Bacula scalable both technically and economically, particularly for HPC, AI, complex IT estates, research and large-enterprise environments.

Frequently Asked Questions

Can cloud storage make a backup system automatically scalable?

Cloud storage can provide virtually elastic capacity, but it does not remove other limitations. Backup software, network bandwidth, API(application programming interface) limits, metadata services, retrieval performance, egress costs, and recovery infrastructure can all affect scalability. Cloud storage should therefore be evaluated as part of the complete backup and recovery architecture.

How much spare capacity should a scalable backup repository maintain?

There is no universal percentage that applies to every environment. Capacity planning should consider expected data growth, daily change rates, retention periods, immutable copies, replication, temporary workload increases, hardware or node failures, and the time required to procure and deploy additional capacity.

Does deduplication always improve backup scalability?

No. Deduplication can reduce storage consumption and network traffic, but its effectiveness depends on the type of data being protected. It also requires processing and metadata resources and may influence recovery performance depending on the implementation. Both backup efficiency and restore performance should be measured with representative workloads.

When should an organization divide one backup environment into multiple domains?

You might want to have separate backup domains when a single environment leads to resource contention, security issues, data-residency requirements or operational dependencies between workloads that are geographically or organizationally separate. The aim should be to create useful boundaries but not create unnecessary management silos. Centralized monitoring, policy management, and reporting can give you visibility across multiple environments.

Can a backup system scale if recovery remains manual?

Only to a limited extent. As the number of protected workloads increases, manual recovery tasks such as sequencing applications, managing credentials, configuring networks, and validating restored data can become significant bottlenecks.

About the author
Rob Morrison
Rob Morrison is the marketing director at Bacula Systems. He started his IT marketing career with Silicon Graphics in Switzerland, performing strongly in various marketing management roles for almost 10 years. In the next 10 years Rob also held various marketing management positions in JBoss, Red Hat and Pentaho ensuring market share growth for these well-known companies. He is a graduate of Plymouth University and holds an Honours Digital Media and Communications degree, and completed an Overseas Studies Program.
Leave a comment

Your email address will not be published. Required fields are marked *