Complete Overhaul of a Production Data Center with Zero Service Interruption
The biggest challenge: the entire overhaul had to be carried out without stopping a single production service. The company operated 24/7, and any extended maintenance window was unacceptable.
Phase 1 - Audit and planning: We carried out an exhaustive inventory of all services, network dependencies, applications, and data flows. We documented every server, every VM, and every critical service to build a risk-free migration map.
Phase 2 - Physical infrastructure: We installed new redundant cooling (an N+1 system), a dual electrical supply with redundant UPS units and a backup generator, and contracted a second internet line with automatic failover. All of this was deployed in parallel with the existing infrastructure.
Phase 3 - Migration to Proxmox: We deployed a high-availability Proxmox VE cluster with Ceph shared storage. We migrated the virtual machines from vSphere and the services from the physical servers and towers progressively, service by service, validating each migration before moving on to the next.
Phase 4 - Consolidation and decommissioning: Once all services had been migrated, we retired the obsolete hardware, optimized the cluster's resources, and set up automated backups with offsite replication.
Zero downtime: The entire migration was completed without any service interruption noticeable to end users.
True high availability: The new Proxmox cluster enables live migration of VMs between nodes, guaranteeing continuity even in the event of hardware failures.
Full redundancy: Cooling, power, and internet now have redundant systems that eliminate every single point of failure.
Cost reduction: Consolidating the heterogeneous hardware into a virtualized cluster cut power consumption by 40% and maintenance costs by 60%.
Scalability: The new infrastructure makes it easy to add capacity, simply by adding nodes to the cluster.
Proxmox VE
Ceph Storage
N+1 Infrastructure