LlamaIndex
Track usage, control costs, and add guardrails to your LlamaIndex applications
What is LlamaIndex?
LlamaIndex is a leading data framework for building LLM applications over your own data. It provides ingestion, indexing, retrieval, and query primitives for retrieval-augmented generation (RAG) and agentic workflows.
By routing LlamaIndex through FastRouter, you get:
100+ models from OpenAI, Anthropic, Google, xAI, Meta, Groq, Mistral, and more through one endpoint
Observability for every request: cost, tokens, latency, and model selection tracked in real time
Reliability through automatic failover across providers, response caching, and intelligent routing
Governance with per-key budgets, rate limits, model restrictions, role-based access, and project isolation
This guide covers connecting LlamaIndex (Python) to FastRouter using the OpenAILike LLM class.
Prerequisites
A FastRouter.ai account (sign up)
Python 3.9 or higher
Quick Start
Step 1: Create a Project and Virtual Environment
You'll only need to do this once:
mkdir my_project
cd my_project
python -m venv .venvActivate the virtual environment. Do this every time you start a new terminal session.
On macOS or Linux:
On Windows:
Step 2: Install LlamaIndex
Step 3: Get Your FastRouter API Key
Sign up or log in at fastrouter.ai
Navigate to your project's Keys page
Click Create User Key
Copy the key immediately. FastRouter does not display the key again after creation.
Export it in your terminal:
Step 4: Point OpenAILike at FastRouter
LlamaIndex's OpenAILike class targets any OpenAI-compatible endpoint. Save this as llamaindex_example.py:
Note: Set
is_chat_model=Trueso LlamaIndex uses the chat completions endpoint.
Step 5: Run the Script

The response prints to your terminal, and the request appears in your FastRouter Dashboard with token usage and cost.
Use LlamaIndex with 100+ Models
FastRouter uses the provider/model-name format. Switch providers by changing the model slug:
The same llm object plugs into LlamaIndex query engines, chat engines, and agents—set it as the default with Settings.llm = llm.
Explore the full model catalog
Automatic Model Selection
Let FastRouter pick the best model for each request based on query complexity, domain, and cost:
Explore automatic model selection
FAQs
Configuration & Setup
Can I use multiple models with the same API key?
Yes. The API key controls access and budget. Create multiple OpenAILike instances with different models, all sharing one key.
Can I restrict a key to only use specific models?
Yes. When creating or editing a key, use the Select Models setting to limit which models the key can access. FastRouter rejects requests to unauthorized models.
Do I also need a separate embeddings model?
RAG pipelines need an embeddings model in addition to the chat model. FastRouter supports embeddings through the same endpoint—see the Embeddings API reference.
Costs & Budgeting
How do I track RAG spending?
Set budgets and rate limits on the key, and use Dynamic Tags to attribute spend per pipeline. The Dashboard breaks down costs by project, key, model, and tag.
Performance & Reliability
Does FastRouter add latency to queries?
FastRouter adds near-zero gateway overhead, negligible compared to model inference and retrieval time.
Next Steps
Set up Fallback Models for high availability
Configure Alerts for spend and performance monitoring
Run a Free Audit on your existing LLM traffic to identify savings
Last updated
