Agno
Track usage, control costs, and add guardrails to your Agno agents
What is Agno?
Agno (formerly Phidata) is a high-performance framework for building multi-agent systems with memory, knowledge, and tools. It is known for extremely fast agent instantiation and a clean, minimal API surface.
By routing Agno through FastRouter, you get:
100+ models from OpenAI, Anthropic, Google, xAI, Meta, Groq, Mistral, and more through one endpoint, via Agno's built-in
OpenAILikemodel classObservability for every request: cost, tokens, latency, and model selection tracked in real time
Reliability through automatic failover across providers, response caching, and intelligent routing
Governance with per-key budgets, rate limits, model restrictions, role-based access, and project isolation
This guide covers connecting Agno (Python) to FastRouter, running an agent, and adding tools.
Prerequisites
A FastRouter.ai account (sign up)
Python 3.9 or higher
Quick Start
Step 1: Create a Project and Virtual Environment
You'll only need to do this once:
mkdir my_project
cd my_project
python -m venv .venvActivate the virtual environment. Do this every time you start a new terminal session.
On macOS or Linux:
On Windows:
Step 2: Install Agno
Step 3: Get Your FastRouter API Key
Sign up or log in at fastrouter.ai
Navigate to your project's Keys page
Click Create User Key
Copy the key immediately. FastRouter does not display the key again after creation.
Export it in your terminal:
Step 4: Use the OpenAILike Model Class
Agno ships an OpenAILike model class made for OpenAI-compatible gateways like FastRouter. Save this as agno_example.py:
Step 5: Run the Agent
Agno prints a formatted response panel in your terminal:

The request appears in your FastRouter Dashboard with token usage and cost.
Adding Tools
The FastRouter-backed model works with Agno tools. Pass any Python function with a docstring and type hints:
Use Agno with 100+ Models
FastRouter uses the provider/model-name format. Change the id to any FastRouter model slug:
Explore the full model catalog
Automatic Model Selection
Let FastRouter pick the best model for each request based on query complexity, domain, and cost:
Explore automatic model selection
Flex Pricing for Background Agents
Agents running scheduled or background jobs tolerate higher latency. Append :flex to any model to access up to 50% lower token costs:
Note: Flex is not recommended for interactive agents where you need low-latency responses.
FAQs
Configuration & Setup
I get ModuleNotFoundError: No module named 'openai'. What's missing?
Agno's OpenAI-compatible models require the OpenAI SDK: pip install openai.
Can I use multiple models with the same API key?
Yes. The API key controls access and budget. Each OpenAILike instance selects its own model, so different agents can use different providers with one key.
Can I restrict a key to only use specific models?
Yes. When creating or editing a key, use the Select Models setting to limit which models the key can access. FastRouter rejects requests to unauthorized models.
Tool runs fail with an Invalid type ... arguments error from the API. What's wrong?
Some model routes currently return tool-call arguments as a JSON object instead of the JSON-encoded string the OpenAI specification requires, which breaks the tool-call round trip. Plain chat responses are unaffected. If you hit this error, switch to a model verified for tool calling through FastRouter, such as openai/gpt-5.2 or anthropic/claude-4.5-sonnet.
Costs & Budgeting
What happens when a key exceeds its budget?
FastRouter blocks further requests until the budget resets (if a reset interval is configured) or an admin increases the limit. The agent receives an error response indicating the budget has been exceeded.
Performance & Reliability
Does FastRouter add latency to agent runs?
FastRouter adds near-zero gateway overhead. For most workloads, this is negligible compared to model inference time.
Next Steps
Set up Fallback Models for high availability
Configure Alerts for spend and performance monitoring
Run a Free Audit on your existing LLM traffic to identify savings
Last updated
