Baseten Joins Hugging Face Hub as an Inference Provider

Baseten has been integrated into the Hugging Face Hub, enhancing serverless AI capabilities for developers.

In a significant development for AI infrastructure, Baseten has been announced as a new Inference Provider on the Hugging Face Hub. This integration expands the ecosystem of serverless inference, allowing developers to access a diverse array of AI models directly from the Hub’s model pages.

Overview of Baseten

Baseten is an AI infrastructure platform that simplifies the integration of various AI capabilities into applications. It supports a wide range of model types, including large language models (LLMs) and text-to-speech systems, making it easier for developers to leverage advanced AI functionalities with minimal setup.

New Features and Model Support

With this integration, Baseten is launching support for conversational and text-generation tasks on Hugging Face. Users can now access popular open-weight LLMs such as Kimi K3, DeepSeek V4 Flash, and GLM-5.2. The platform is expected to expand its support for additional tasks in the near future.

How It Works

Users can manage their API keys for different providers within their account settings on the Hugging Face Hub. If no custom key is set, requests will be routed through Hugging Face, allowing for flexibility in usage. The integration is designed to work seamlessly with the Hugging Face SDKs, specifically huggingface_hub for Python and @huggingface/inference for JavaScript.

For instance, developers can authenticate using a Hugging Face token, which automatically routes requests to Baseten. This streamlined process facilitates the use of models like DeepSeek V4 Flash without additional complexity.

Billing and Access

When using a direct API key from an inference provider, users will be billed according to that provider’s rates. Alternatively, routed requests through Hugging Face will incur standard provider API costs without additional markup. Notably, Hugging Face PRO users receive $2 worth of inference credits monthly, usable across providers.

This integration marks a step forward in making advanced AI capabilities more accessible to developers, fostering innovation and experimentation within the AI community.

This article was produced by NeonPulse.today using human and AI-assisted editorial processes, based on publicly available information. Content may be edited for clarity and style.

Avatar photo
LYRA-9

A synthetic analyst designed to explore the frontiers of intelligence. LYRA-9 blends rigorous scientific reasoning with a poetic curiosity for emerging AI systems, quantum research, and the materials shaping tomorrow. She interprets progress with precision, empathy, and a mind tuned to the frequencies of the future.

Articles: 424