In a major bid to tackle skyrocketing enterprise artificial intelligence (AI) costs, AI infrastructure software provider Spectro Cloud announced PaletteAI Inference Launchpad.

The turnkey, locally managed solution aims to slash token costs by up to 70% while enabling organizations to run AI inference closer to their data, applications, and end-users.

Additionally, Spectro Cloud announced expanded support for AMD-powered hardware. The platform now integrates with AMD Inc. GPUs, the AMD GPU Operator, the ROCm™ runtime, and the AMD enterprise AI reference stack, which features optimized models from the AMD Inference Microservices (AIMs) catalog.

The dual announcements position PaletteAI as a unified platform capable of managing heterogeneous AI environments. The integration offers enterprises, neoclouds, and sovereign cloud providers a flexible, governed framework to operate infrastructure across both NVIDIA Corp. and AMD silicon.

The launch comes at a critical inflection point as companies transition AI initiatives from experimental phases to full-scale production. Scaling these operations presents significant financial and technical hurdles.

According to data from Goldman Sachs Research, global AI token consumption is projected to surge 24-fold by 2030, reaching an unprecedented 120 quadrillion tokens per month.

For modern enterprises, this exponential growth turns token consumption into a primary operational bottleneck. Infrastructure teams are increasingly tasked with controlling budgets, metering usage, enforcing corporate governance, and determining whether workloads should execute locally or via external model services. Compounding these challenges is a highly fragmented AI landscape, forcing buyers to seek greater flexibility across diverse GPUs, models, and deployment environments.

PaletteAI Inference Launchpad directly addresses these operational pain points. By providing a pre-configured, locally managed solution, the platform allows organizations to establish efficient token factories without the burden of building and maintaining DIY inference stacks. The software enables intelligent routing between local and external frontier models, applies quota controls, and minimizes reliance on costly third-party services.

By combining local operational control with centralized lifecycle management, the platform delivers a vertically integrated yet horizontally open architecture. The expanded hardware compatibility has been welcomed by silicon manufacturers looking to provide clients with more deployment flexibility.

“Enterprises and cloud providers are looking for open, scalable AI infrastructure that gives them more control over cost, performance, and deployment choice,” said Kumaran Siva, corporate vice president of enterprise AI at AMD. “Spectro Cloud’s support for AMD-powered infrastructure in PaletteAI and PaletteAI Inference Launchpad helps customers accelerate production AI deployments across flexible, open AI stacks.”

Spectro Cloud has partnered with infrastructure provider NexusIgnite to expand its PaletteAI Inference Launchpad into highly regulated sectors, neoclouds, and sovereign environments. The collaboration aims to help enterprises deploy and manage AI inference infrastructure where strict data residency, governance, and operational control are critical requirements.

“Spectro Cloud’s PaletteAI Inference Launchpad fits our managed AI infrastructure strategy,” NexusIgnite CEO Greg Forrest said, noting that the collaboration gives compliance-driven customers a faster, more manageable path to production within trusted sovereign environments.