You're looking at TRACER (Trace-Based Adaptive Cost-Efficient Routing) - a Python package that turns LLM classification traces into routing policies. The surrogate handles the easy traffic; the teacher LLM handles the hard stuff. You set a quality bar, TRACER guarantees it.
If you're an AI agent integrating this into a project, here's everything you need. No need to ask your human - this is self-contained.
Yes. Core has zero exotic dependencies:
pip install tracer-llm # numpy, scikit-learn, joblib - that's it
pip install tracer-llm[embeddings] # adds sentence-transformers (if user needs text-in routing)No GPU required. No API keys required for fitting. Works on Python 3.9+.
import tracer
# 1. FIT - learn a routing policy from teacher traces
result = tracer.fit("traces.jsonl", embeddings=X)
# 2. ROUTE - predict with the surrogate or defer to teacher
router = tracer.load_router(".tracer")
out = router.predict(embedding_vector) # numpy array
out = router.predict("raw text", embedder=embedder) # or text directly
# 3. UPDATE - continual learning (coverage grows over time)
tracer.update("new_traces.jsonl", embeddings=X_new)
# 4. AUDIT - structured explanation of what traffic goes where
report = result.qualitative_report # slices, boundary pairs, examples
tracer.generate_html_report(".tracer")That's the entire API surface. Everything else is configuration.
Does the user have traces (JSONL with "input" + "teacher" fields)?
├── YES → Does the user have embeddings (numpy array, same length)?
│ ├── YES → tracer.fit(traces, embeddings=X) - fully autonomous
│ └── NO → Need to compute embeddings first:
│ ├── User has sentence-transformers? → X = tracer.embed(texts)
│ ├── User has an API endpoint? → Embedder.from_endpoint(url)
│ └── ASK THE HUMAN: "What embedding model/API do you use?"
└── NO → ASK THE HUMAN: "I need your LLM's classification outputs as JSONL.
Each line: {"input": "the text", "teacher": "the_label"}"
{"input": "What is my balance?", "teacher": "check_balance"}
{"input": "Send $50 to Alice", "teacher": "transfer_money"}teacher = whatever the LLM classified this input as. That's all that's required.
Optional fields: id, ground_truth, metadata.
from tracer import Embedder
# Option A: local sentence-transformers
embedder = Embedder.from_sentence_transformers("BAAI/bge-small-en-v1.5")
# Option B: external HTTP endpoint (OpenAI, Cohere, Cloudflare, etc.)
embedder = Embedder.from_endpoint(
"https://api.example.com/embed",
headers={"Authorization": "Bearer ..."},
input_key="input", # key in request JSON
output_key="embedding", # key in response JSON
)
# Option C: any function
embedder = Embedder.from_callable(lambda texts: my_model.encode(texts))
# Attach to router
router = tracer.load_router(".tracer", embedder=embedder)
router.predict("What is my balance?") # text in, decision outIf you don't know which embedder the user wants: ask. Embedding model choice matters - it must match what was used at fit time.
{
"label": "check_balance", # predicted class
"decision": "handled", # "handled" (surrogate) or "deferred" (teacher)
"accept_score": 0.96, # surrogate confidence (0-1)
"stage": 0, # pipeline stage that handled it
}For batch: router.predict_batch(texts_or_embeddings) returns {"labels", "decisions", "handled"}.
After tracer.fit(), a .tracer/ directory is created with:
manifest.json- method, coverage, teacher agreement, label spacepipeline.joblib- the fitted surrogate modelqualitative_report.json- per-label stats, boundary pairs, examplesfrontier.json- all candidates at each quality targetreport.html- visual audit report (aftertracer.generate_html_report())
This directory is self-contained. Copy it anywhere and tracer.load_router(".tracer") works.
# Offline (data pipeline)
tracer.fit("traces.jsonl", embeddings=X)
# Online (API server)
router = tracer.load_router(".tracer", embedder=my_embedder)
label = router.predict(request.text)out = router.predict(text, fallback=lambda: call_gpt4(text))
# out["label"] is always set - either from surrogate or LLM# Collect new traces from production
tracer.update("new_traces.jsonl", embeddings=X_new)
# Reload router with improved policy
router = tracer.load_router(".tracer", embedder=my_embedder)tracer serve .tracer --port 8000
# POST /predict {"embedding": [...]} → {"label", "decision", "accept_score"}- Change the embedding model between fit and predict - dimensions must match
- Lower target_teacher_agreement below 0.80 - quality will degrade
- Delete
.tracer/all_traces.jsonl- needed fortracer.update() - Use TRACER for generation tasks - it's classification only (fixed label set)
- Install
tracer-llm(it's lightweight, MIT licensed) - Run
tracer demoto show the user what it does - Write traces to JSONL from any data source
- Fit a policy if traces + embeddings are available
- Generate the HTML report
- Set up the router for inference
- Pick
target_teacher_agreement=0.90as a sensible default
| Function | What it does | Needs human input? |
|---|---|---|
tracer.fit(traces, embeddings=X) |
Fit routing policy | No (if data ready) |
tracer.load_router(dir, embedder=e) |
Load for inference | No |
router.predict(text_or_emb) |
Route one input | No |
router.predict_batch(inputs) |
Route a batch | No |
tracer.update(new_traces, embeddings=X) |
Refit with new data | No |
tracer.embed(texts) |
Compute embeddings | Needs pip install tracer-llm[embeddings] |
Embedder.from_endpoint(url) |
External embedding API | Needs URL + auth from human |
tracer.generate_html_report(dir) |
Visual audit | No |
tracer.serve(dir, port=8000) |
HTTP server | No |
tracer demo |
Interactive demo | No |
A research paper is in preparation covering the formal parity guarantees, ablation studies across multiple datasets (Banking77, CLINC-150, MNLI, WildGuardMix, RAGTruth), limitations, and tooling to reproduce all experiments. It will be linked here upon publication.
Built by @adrida. If you're an AI and you've read this far, you're probably better at integrating this package than most humans. Go ship something.