What compute infrastructure do I need for AI workloads?
AI training workloads require GPU servers with NVIDIA H100 or H200 GPUs. Each GPU has 80–141 GB of HBM3/HBM3e memory and requires 700W TDP. A single 8-GPU server (DGX H100) requires approximately 10.2 kW total power and liquid cooling. AI inference workloads can run on fewer GPUs: the requirement depends on the model size, batch size, and throughput target. Before ordering GPU hardware, verify that your facility has adequate power capacity, cooling capability, and network bandwidth.
How do I determine the right server configuration for my workloads?
Start with workload requirements: CPU core count, memory working set at peak load, storage IOPS and throughput, and network bandwidth. Size memory for peak working set plus 20% headroom: under-provisioned memory causes swap usage that degrades performance by 100x. Size CPU for the specific workload type: more cores for parallelizable workloads, higher clock speed for sequential workloads. Validate the configuration against workload requirements before procurement.
What is the difference between AMD EPYC and Intel Xeon?
AMD EPYC (Turin) offers up to 128 cores per socket, 460 GB/s memory bandwidth, and 128 PCIe 5.0 lanes: advantages for virtualization, HPC, and memory-bandwidth-bound workloads. Intel Xeon (Granite Rapids) offers up to 60 cores per socket with advantages in single-threaded performance and some ISV certifications. Platform selection should be driven by workload requirements and ISV certification requirements, not by general preference.
How long do enterprise servers last?
Enterprise servers have a typical useful life of 5–7 years. After year 5, hardware failure rates increase, manufacturer support may expire, and the performance gap between current hardware and new hardware widens. A planned 5-year refresh cadence is less expensive and less disruptive than emergency replacement after a failure. Track manufacturer support status, not just hardware age, to identify hardware that requires replacement.
What is end-of-support hardware and why does it matter?
End-of-support hardware has passed the manufacturer's end-of-support date and can no longer receive security patches or firmware updates. Vulnerabilities discovered after end-of-support cannot be remediated without hardware replacement. This is a documented attack surface that threat actors actively exploit. Organizations running end-of-support hardware in production are accepting security risk that cannot be mitigated through software controls alone.
Should I buy servers directly from the OEM or through a reseller?
Both options are valid: the right choice depends on your procurement capabilities and requirements. Direct OEM purchase provides the best pricing at volume and a direct manufacturer relationship. Authorized resellers provide deployment services, multi-OEM procurement, and financing options. The critical requirement is that the reseller is OEM-authorized, unauthorized resellers do not provide OEM warranty coverage. Verify authorization status before purchase.
What support contract level do I need?
Support contract level should match the criticality of the workloads running on the hardware. Mission-critical workloads (financial transaction processing, healthcare systems) require 2-hour on-site response with pre-positioned parts. Business-critical workloads require 4-hour on-site response. Development and test environments can typically be served by next-business-day response. Do not apply a single support tier to all hardware, differentiate by workload criticality.
How long does it take to procure GPU hardware?
NVIDIA H100 and H200 GPU hardware has lead times of 12–26 weeks from authorized OEM channels. Organizations that do not plan procurement timelines accordingly consistently miss their AI deployment targets. Plan GPU procurement at least 6 months before the required deployment date. DCS Global maintains allocation relationships with NVIDIA-authorized OEMs and can provide current lead time estimates.
What is the difference between bare metal and virtualized compute?
Bare metal servers run workloads directly on the hardware without a hypervisor layer. Virtualized servers run a hypervisor that hosts multiple virtual machines. Bare metal provides better performance for latency-sensitive and compute-intensive workloads: virtualization adds 5–15% overhead for CPU and memory, and higher overhead for I/O-intensive workloads. Virtualization provides better flexibility and utilization for general enterprise workloads. Most enterprise data centers use both: bare metal for performance-critical workloads, virtualization for general workloads.
How do I dispose of servers that have processed regulated data?
Servers that have processed regulated data (HIPAA, PCI DSS, FedRAMP) require certified data destruction before disposal. Certified data destruction involves overwriting all storage media to DoD 5220.22-M standards or physical destruction of storage media. The disposal vendor must provide a certificate of destruction that documents the serial numbers of all storage media destroyed. Retain these certificates, they may be required for compliance audits.
Questions specific to your environment?