Baseten joins Hugging Face Hub as an inference provider
Baseten has become a supported inference provider on the Hugging Face Hub, allowing developers to run open-weight language models directly from model pages, client SDKs, and agent harnesses. The initial integration brings support for conversational and text-generation tasks using models such as Kimi K3, DeepSeek V4 Flash, and GLM-5.2.
Drafted by our automated systems from the sources listed at the end of this article, and reviewed before publication. Every fact here should be traceable to one of those links; we publish no benchmark, price or figure that a linked source does not state. Where a company makes a claim about its own product, we report it as their claim.
According to a blog post published on 6 August 2026, Hugging Face announced that Baseten is now a supported inference provider on the Hugging Face Hub. The integration adds serverless inference options directly onto model pages on the platform. Baseten is an AI infrastructure platform providing serverless AI and training capabilities. As part of this initial launch, Baseten supports conversational and text-generation tasks. Developers can access open-weight large language models including Kimi K3, DeepSeek V4 Flash, and GLM-5.2 through the platform. Additional tasks will receive support in future updates, according to the announcement.
Integration and access methods
The integration is available through the website user interface and client SDKs. Users can set custom API keys for providers in their account settings or have requests routed through Hugging Face. When routed by Hugging Face, users authenticate with a Hugging Face token and charges apply directly to the user's Hugging Face account at standard provider rates without additional markup. For client-side access, Baseten is available via the hugging_face_hub Python package (version 1.26.1 or greater) and the @huggingface/inference JavaScript library. The platform is also integrated into agent harnesses such as Pi, OpenCode, Hermes Agents, and OpenClaw.
Is it worth it?
Whether this integration is worth using depends on how you prefer to manage your API billing and keys for open models. Developers already working within the Hugging Face ecosystem can route requests directly through their accounts or use monthly PRO plan inference credits, while teams with existing Baseten accounts can supply their own keys. You would want to check the specific model catalog and pricing rates on Baseten to decide whether routing through Hugging Face or connecting directly suits your workflow.
What this changes for tools we list
- Hugging Face — The story introduces Baseten as a new supported inference provider on the Hugging Face Hub, giving users a new backend option for running open-weight models directly from model pages and SDKs.
Sources
- Baseten on Hugging Face Inference Providers 🔥 — Hugging Face, 6 August 2026