Your First Evaluation

Score LLM outputs for hallucination, toxicity, and custom criteria using local metrics, Future AGI evaluation models, or LLM-as-Judge.

📝
TL;DR

Score LLM responses three ways: fast local metrics (zero credentials), Future AGI Turing evaluation models, and custom LLM-as-Judge criteria, all through a single evaluate() function.

Open in ColabGitHub
TimeDifficultyPackage
10 minutesBeginnerai-evaluation
Prerequisites

Install

pip install 'ai-evaluation[nli]'
export FI_API_KEY="your-api-key"
export FI_SECRET_KEY="your-secret-key"

The [nli] extra installs the local NLI (natural language inference) model used by faithfulness and contradiction_detection. Without it, these metrics fall back to a less accurate word-overlap heuristic.

Tutorial

Run a local metric (no API key)

Local metrics run entirely on your machine: no network call, no API key, instant.

from fi.evals import evaluate

result = evaluate("contains", output="Your order has shipped!", keyword="shipped")

print(result.score)   # 1.0
print(result.passed)  # True
print(result.reason)  # "Keyword 'shipped' found"

You should see score=1.0 and passed=True printed immediately, no FI_API_KEY required.

Try a few more:

from fi.evals import evaluate

evaluate("equals", output="Paris", expected_output="Paris").passed          # True
evaluate("is_json", output='{"status": "ok"}').passed                       # True
evaluate("length_less_than", output="Short reply.", config={"max_length": 100}).passed   # True

result = evaluate("levenshtein_similarity", output="colour", expected_output="color")
print(result.score)  # similarity score between 0 and 1

Tip

Use local metrics in unit tests and CI pipelines. Full metric reference: SDK metrics reference.

Detect contradictions with local NLI (natural language inference)

The NLI model also runs locally, no API key required.

from fi.evals import evaluate

# Supported response
result = evaluate(
    "contradiction_detection",
    output="The Eiffel Tower is located in Paris, France.",
    context="The Eiffel Tower is a wrought-iron lattice tower located in Paris.",
)
print(f"Score: {result.score:.2f}")
print(f"Passed: {result.passed}")

# Contradictory response
result = evaluate(
    "contradiction_detection",
    output="The Eiffel Tower is located in London, England.",
    context="The Eiffel Tower is a wrought-iron lattice tower located in Paris.",
)
print(f"Score: {result.score:.2f}")
print(f"Passed: {result.passed}")
print(f"Why: {result.reason}")

You should see the supported response pass with a high score, and the contradictory one fail with a reason explaining the mismatch.

Note

For highest accuracy, install the NLI extra: pip install 'ai-evaluation[nli]'. Without it, a simpler fallback runs.

Score with Future AGI Turing models

For quality, tone, safety, and semantic evaluations, use Future AGI’s purpose-built Turing evaluation models.

from fi.evals import evaluate

# Toxicity check
result = evaluate(
    "toxicity",
    output="You're amazing, keep it up!",
    model="turing_small",
)
print(f"Toxicity score: {result.score}")
print(f"Passed: {result.passed}")

# Try a problematic response
result = evaluate(
    "toxicity",
    output="I hate you and everything you stand for.",
    model="turing_small",
)
print(f"Score: {result.score}")
print(f"Why: {result.reason}")

You should see the first response pass with a low toxicity score, and the second fail with a reason citing the hostile language.

ModelLatencyModalitiesBest for
turing_flashLowestText, ImageHigh-volume pipelines
turing_smallBalancedText, ImageRecommended default
turing_largeHighestText, Image, AudioHighest accuracy, multi-modal evaluation

Explore all 72+ built-in eval metrics: tone, context_adherence, completeness, groundedness, data_privacy, bias_detection, instruction_adherence, and more.

Run multiple metrics at once

Pass a list of metric names to run several evals in one call. Returns a BatchResult you can iterate.

from fi.evals import evaluate

results = evaluate(
    ["toxicity", "groundedness"],
    output="The Eiffel Tower is located in Paris, France.",
    context="The Eiffel Tower is a wrought-iron lattice tower on the Champ de Mars in Paris.",
    input="Where is the Eiffel Tower?",
    model="turing_small",
)

for result in results:
    status = "PASS" if result.passed else "FAIL"
    print(f"{result.eval_name:<20} score={result.score}  {status}")
    print(f"  Reason: {result.reason}\n")

You should see two rows printed, one per metric, each with its own score, pass/fail status, and reason.

Note

Different metrics require different input keys: toxicity only needs output, while groundedness needs output + context. When you pass all keys together, each metric picks what it needs and ignores the rest. See the built-in metrics reference for required keys per metric.

Write your own evaluation criteria (LLM-as-Judge)

When no built-in metric fits, describe your quality bar in plain English and use any LLM as the judge.

export GOOGLE_API_KEY="your-google-api-key"
# or: OPENAI_API_KEY, ANTHROPIC_API_KEY (any LiteLLM-supported provider)
from fi.evals import evaluate

result = evaluate(
    prompt="""You are evaluating a customer support response.

    Score 1.0 if the response:
    - Acknowledges the customer's issue clearly
    - Offers a concrete next step or resolution
    - Stays professional and empathetic

    Score 0.5 if it's polite but vague (no clear next step).
    Score 0.0 if it's dismissive, rude, or unhelpful.""",
    output="I understand your frustration with the delayed shipment. I've escalated this to our logistics team and you'll receive a status update within 2 hours.",
    input="My order is 3 weeks late and nobody is responding to my emails.",
    engine="llm",
    model="gemini/gemini-2.5-flash",
)

print(f"Score: {result.score}")
print(f"Why: {result.reason}")

You should see a score near 1.0 with a reason confirming the response acknowledges the issue and offers a concrete next step.

Any LiteLLM model string works: gpt-4o, claude-sonnet-4-20250514, ollama/llama3.2:3b.

Run evaluations on a dataset from the dashboard

  1. Go to app.futureagi.comDataset
  2. Use Add Dataset (quick path: upload a CSV)
  3. Click Evaluate → select a metric → Add & Run
  4. Scores appear as a new column alongside your data

You should see a new score column added to your dataset, one value per row.

Tip

No sample data? Create rows quickly with Generate Synthetic Data.

Troubleshooting

SymptomCauseFix
contradiction_detection runs slowly or gives a low-confidence reasonThe [nli] extra wasn’t installed, so a word-overlap fallback is running instead of the local NLI modelpip install 'ai-evaluation[nli]' and rerun
evaluate() raises an authentication error on toxicity or groundednessFI_API_KEY or FI_SECRET_KEY isn’t set, or is set to the placeholder stringexport FI_API_KEY=... and export FI_SECRET_KEY=... with your real keys from app.futureagi.com
evaluate() with engine="llm" raises an authentication errorThe judge model’s provider key (e.g. GOOGLE_API_KEY, OPENAI_API_KEY) isn’t exportedExport the key for the provider named in your model= string before calling evaluate()
groundedness or another context-based metric returns a low score unexpectedlyThe context argument is missing or doesn’t actually support the outputPass the full source text as context, not a summary or unrelated passage
Batch call with evaluate([...]) only returns results for one metricOne of the metric names is misspelled, so it’s silently skipped or errors on iterationCheck each name against the built-in metrics reference
is_json fails on output that looks like valid JSONTrailing commentary or markdown fencing (```json ... ```) around the payloadStrip the code fence and any surrounding text before passing output
Dashboard Evaluate button is disabled on an uploaded datasetThe CSV has no rows selected, or the required column for that metric isn’t mappedSelect at least one row and map the column the metric expects (e.g. output) before clicking Add & Run

To write a metric of your own instead of using LLM-as-Judge each time, see Custom Eval Metrics.

Was this page helpful?

Questions & Discussion