Platform: WindowsProducts: MSP360 Backup
Article ID: s0050Last Modified: 07-Oct-2026

Client-Side Deduplication

Deduplication is an approach that involves multiple usage of the same data parts in various processes.

This functionality is not supported for the legacy backup format.

The new backup format uses client-side deduplication. This approach brings the following benefits:

  • Client-side deduplication is much faster than server-side deduplication
  • No internet connection issues
  • Reduced internet traffic
  • Ability to purge unnecessary data
  • A server-side deduplication database constantly grows, which can significantly increase expenses. Client-side deduplication uses local capacities only.

How It Works

Regardless of the backup type, the first backup is always a full backup. Subsequent backups capture data updates, so the next backup jobs are usually incremental and depend on the full backup and previous incremental backups.

The backup format ensures full backup plan independence, so each separate backup plan has its own deduplication database. Moreover, backup plan generations also have their own deduplication databases.

Once a backup plan is run, the application reads backup data in batches that are multiples of the block size. Once a block is read, it is compared with deduplication database records. If a block is not found, it is delivered to storage and assigned a block ID, which becomes a new deduplication database record. The block scanning continues, and if a block matches any of the deduplication database records, a block with such ID is excluded from a backup plan.

This approach significantly decreases the backup size, especially in virtual environments with a large number of identical blocks.

Once a deduplication database is deleted or corrupted, a full backup is always initiated.

For the image-based backup type, the approach is slightly different. Instead of cluster reading, Backup for Windows reads a Master File Table (MFT) and checks which files have been modified. This decreases source data reading exponentially.

https://git.cloudberrylab.com/egor.m/doc-help-std.git
Production