TL;DR — Key Takeaways

– Arm has launched the Arm AI Portal to simplify the discovery, optimization and deployment of AI models on Arm-powered processors.

– The portal provides access to pre-optimized models, including Alibaba Qwen, Google Gemma and Ultralytics YOLO, alongside performance data, code examples and deployment workflows.

– Integration with Hugging Face and support for the Model Context Protocol (MCP) make resources accessible to both human developers and AI coding agents.

Arm has created an online portal through which it is making available tools and optimized models for teams that are building artificial intelligence (AI) applications that run on its processors.

The Arm AI Portal is designed to make AI models, performance data and workflows discoverable to both AI coding agents and human developers to streamline the development of AI applications, says Sharbani Roy, vice president of AI and developer platforms for Arm.

At launch, pre-optimized models that can be accessed include Alibaba Qwen, Google Gemma and Ultralytics YOLO using ExecuTorch, LiteRT and ONNX-RT runtimes, with support being provided by Arm partners such as Alibaba, Raspberry Pi and Ultralytics. The models themselves are hosted on the platform provided by Hugging Face and made accessible via the Model Context Protocol (MCP). The Arm AI Portal also provides access to software optimized for Arm technologies such as Scalable Vector Extension (SVE) and Scalable Matrix Extension (SME).

Arm, via an early access program, is also opening the portal to application development teams that want to optimize their own custom model for Arm processors, notes Roy. “There needs to be one place to build and customize,” she says.

The overall goal is to make it simpler for application development teams to discover the right model for the right use case by providing performance and accuracy data, latency comparisons, memory and size requirements, code examples and deployment workflows, says Roy.

The providers of the processors on which AI applications are deployed are especially keen at this moment to expand the overall size of their ecosystems. While the first wave of AI models were largely trained on graphics processing units (GPUs) provided by NVIDIA and often deployed on GPUs also provided by NVIDIA, the next wave of AI will be much more diverse, especially as organizations embrace smaller AI models that have been optimized for a specific task. As such, Arm is betting that its processors, along with rival offerings from Intel, AMD and Qualcomm, will be increasingly used to both train AI models and run the AI inference engines that wind up being embedded into an AI application.

It’s not likely that application development teams will standardize on any one class of processors, but as the number of AI applications being developed increases, the providers of the underlying processors are making a concerted effort to reduce as much friction as possible in the hopes of becoming a preferred option.

Regardless of what processor technologies are ultimately adopted, the one thing that is certain is AI applications will soon be deployed everywhere from the endpoint to the cloud and everywhere in between. Each application development team will need to decide to what degree to optimize their applications for any class of processors. As always, there might be tradeoffs that will need to be made between ensuring portability and optimizing performance by, for example, invoking a library that only runs on a specific class of processors.

The challenge, of course, is making sure that whatever tradeoff is made is worth the benefits that, hopefully, will be gained across the lifetime of the application.

Frequently Asked Questions

What is the Arm AI Portal?
The Arm AI Portal is an online resource designed to help developers discover, optimize and deploy AI models on Arm-powered processors. It provides access to optimized models, performance data, code examples and deployment workflows.
Which AI models are available through the Arm AI Portal?
At launch, the portal provides access to pre-optimized models including Alibaba Qwen, Google Gemma and Ultralytics YOLO, with support for ExecuTorch, LiteRT and ONNX Runtime.
How does the Arm AI Portal support AI coding agents?
The portal uses the Model Context Protocol (MCP) to make AI models and development resources discoverable to AI coding agents, enabling them to incorporate optimized models and workflows into application development.

TECHSTRONG AI PODCAST

SHARE THIS STORY