AMD has added the FastFlowLM team to its Artificial Intelligence Group as part of a broader effort to improve AI performance and efficiency across its hardware and software stack. Financial terms and the structure of the transaction were not disclosed.
FastFlowLM developed a lightweight inference software flow optimized for AMD-powered AI PCs and workstations. The technology is designed to run large language and multimodal models efficiently on local devices rather than relying entirely on cloud infrastructure.
AI inference is the process through which a trained model generates responses, predictions or other outputs. Faster and more efficient inference can improve application responsiveness while reducing memory, energy and computing requirements.
The FastFlowLM team will focus on strengthening AMD’s client and workstation AI software stack. AMD also expects the team to improve Day-0 enablement, meaning support for newly released AI models as soon as they become available.
FastFlowLM was developed within the open-source ecosystem using IRON, an open-source neural processing unit compiler created by AMD’s Research and Advanced Development Group. A compiler translates software instructions into forms that specialized processors can execute efficiently.
IRON is intended to support a fully open software stack for AMD’s agentic AI platforms. Agentic AI systems can perform multistep tasks, use tools and take actions with more autonomy than conventional prompt-and-response applications.
AMD said IRON was incubated internally and refined through collaboration with external developers and researchers. This approach helped establish a community around software designed to run AI workloads on AMD neural processing units.
FastFlowLM has also worked closely with Lemonade, AMD’s open-source inference initiative. Lemonade is designed to make it easier for developers to deploy AI models and applications across AMD-powered devices.
The integration supports local agentic AI, retrieval-augmented generation, coding and multimodal applications. Retrieval-augmented generation allows an AI model to use external or private information sources when producing responses, while multimodal systems can process combinations of text, images, audio and other data.
Bringing the FastFlowLM team into AMD could help the company accelerate the optimization of new AI models for its processors. The work will focus particularly on AI PCs and professional workstations, where customers increasingly expect advanced models to operate directly on the device.
On-device AI can provide lower latency, greater privacy and reduced dependence on internet connectivity. It may also lower cloud computing expenses by processing more workloads locally.
The team recently released support for Qwen3.6-35B-A3B, which AMD described as the second mixture-of-experts model released for its neural processing units. Mixture-of-experts models divide tasks among specialized parts of the neural network, activating only a subset for each request to improve efficiency.
AMD said it will continue investing in the open-source ecosystem supporting its AI processors. The FastFlowLM team’s experience is expected to contribute to the development of faster, more accessible and more efficient on-device AI applications.