TL;DR — Key Takeaways
-
CIQ Expands Open-Weight AI Deployment: Version 4.3 of Fuzzball introduces a catalog of 11 validated AI models and agents designed to simplify enterprise AI deployment and orchestration.
-
AI Agents Gain Access to Enterprise Models and Data: Built-in LiteLLM gateways and retrieval capabilities allow AI agents to discover models and connect to enterprise documents with minimal integration.
-
Autoscaling Improves GPU Efficiency: Fuzzball enables organizations to scale AI workloads across distributed environments and scale down to zero when models are idle, freeing GPU resources for other tasks.
CIQ today made available an update to its orchestration platform to make it simpler to deploy open-weight artificial intelligence (AI) models.
Version 4.3 of Fuzz now provides access to a catalog of 11 validated models and AI agents that organizations can not only run anywhere but also connect multiple AI agents they have deployed. The first AI models in the catalog include GPT-OSS, Llama 4, Gemma 4, Qwen3-Coder, Mistral, and three Nemotron 3 model families.
The AI agents in the catalog are OpenCode and hermes-agent, two agents that can discover running models via gateways, without any manual integration because each model ships behind a built-in LiteLLM gateway that can be invoked using an OpenAI-compatible endpoint. The catalog provides a retrieval service to connect AI agents to enterprise documents.
Additionally, organizations can leverage a vLLM-based inference engine to serve any AI model that is compatible with the specifications defined by Hugging Face.
That capability makes it simpler for organizations that are finding it challenging to deploy open-weight AI models as an alternative to relying on proprietary frontier AI model services that are considerably more expensive to invoke using tokens that increase costs for every input and output, says CIQ president Bjorn Hovland. On-premises deployments of open-weight models eliminate those costs, he adds. Organizations are not going to want to pay to use frontier models for everything, says Hovland.
In general, organizations are starting to realize that less expensive open-weight models are more than capable of automating 90% of the tasks they are looking to automate, notes Hovland. The only thing holding organizations back at this point is the complexity of deploying the infrastructure needed to run those AI models, notes Hovland.
Fuzzball enables IT teams to achieve that goal by providing an ability to autoscale AI models up or out across a distributed computing environment from a single endpoint in a way that allows them to scale down to zero when those models are idle, he adds. That capability is critical because it frees up limited NVIDIA and AMD graphics processing units (GPUs) for other workloads, says Hovland. Idle pools wake only for authenticated requests that IT administrators approve.
Previously, CIQ launched an open source OpenWALDO initiative through which participants are sharing an auditable corpus of AI training data and code to foster the development of additional open-weight models. In fact, over time IT teams will be swapping in and out open-weight and open source models as providers continue to leapfrog each other in terms of capabilities, notes Hovland.
It’s not clear how widely open-weight AI models are being adopted, but it’s apparent that many more of them will eventually be deployed. The only thing left to be determined next is how many of them the average organization might need to deploy at any given time.


