NDCG@K

Rates RAG retrieval ranking quality using Normalized Discounted Cumulative Gain, rewarding relevant results that appear earlier.

NDCG@K scores not just whether relevant chunks were retrieved, but whether they land near the top of the ranked list. Run it when ranking order matters as much as coverage.

What it does

NDCG@K is a statistical metric. It compares the retrieved chunks against the ground-truth relevant chunks and scores ranking quality, giving more credit for relevant chunks that appear earlier.

Input

Required InputTypeDescription
hypothesisstringJSON-serialized list of retrieved chunks in ranked order
referencestringJSON-serialized list of ground-truth relevant chunks

Output

FieldTypeDescription
ResultscoreA score between 0 and 1, where 1 means all relevant chunks appear at the top of the ranked list in ideal order
ReasonstringShort summary string of the score, e.g. NDCG@3: 0.469
Parameter
NameTypeDescription
eval_config (evalConfig in TypeScript)dict / Record<string, any>Optional. Pass {"k": N} to limit evaluation to the top N retrieved chunks. Defaults to using the full list

Run it from code

Call evaluate() with the template name and the eval’s required inputs. It returns the score and the reason.

Note

Before running: install the SDK and set FI_API_KEY / FI_SECRET_KEY. The model argument in the snippets is the evaluator model Future AGI uses to run the eval; turing_flash is a fast default.

import json
from fi.evals import evaluate

result = evaluate(
    "ndcg_at_k",
    hypothesis=json.dumps([
        "France is in Europe.",
        "Paris is the capital of France.",
        "Napoleon was born in Corsica.",
        "The Eiffel Tower was built in 1889.",
        "The Louvre is in Paris."
    ]),
    reference=json.dumps([
        "Paris is the capital of France.",
        "The Eiffel Tower was built in 1889.",
        "The Louvre is in Paris."
    ]),
    eval_config={"k": 5},
)

print(result.score)   # Score reflecting ranking quality
print(result.reason)
import { evaluate } from "@future-agi/ai-evaluation";

const result = await evaluate(
  "ndcg_at_k",
  {
    hypothesis: JSON.stringify([
      "France is in Europe.",
      "Paris is the capital of France.",
      "Napoleon was born in Corsica.",
      "The Eiffel Tower was built in 1889.",
      "The Louvre is in Paris."
    ]),
    reference: JSON.stringify([
      "Paris is the capital of France.",
      "The Eiffel Tower was built in 1889.",
      "The Louvre is in Paris."
    ])
  },
  { evalConfig: { k: 5 } }
);

console.log(result.score);   // Score reflecting ranking quality
console.log(result.reason);

In this example, 3 relevant chunks are scattered across positions 2, 4, and 5 instead of being at the top. NDCG penalizes this because a perfect retriever would place all 3 relevant chunks at positions 1, 2, and 3.

Batch evaluation

To evaluate multiple queries in a single call, pass a list of JSON-serialized inputs. Each element represents one retrieval evaluation:

results = evaluate(
    "ndcg_at_k",
    hypothesis=[
        json.dumps(["Paris is the capital of France.", "France is in Europe.", "Napoleon was born in Corsica."]),
        json.dumps(["The sky is blue.", "Water is wet."]),
        json.dumps(["Unrelated 1.", "Unrelated 2.", "Unrelated 3.", "The Louvre is in Paris."]),
    ],
    reference=[
        json.dumps(["Paris is the capital of France.", "The Eiffel Tower was built in 1889."]),
        json.dumps(["The sky is blue.", "Water is wet."]),
        json.dumps(["The Louvre is in Paris."]),
    ],
    eval_config={"k": 3},
)

for i, r in enumerate(results):
    print(f"Query {i+1}: {r.score}")
# Query 1: score reflects that 1 relevant chunk is at position 1 (good ranking)
# Query 2: 1.0 (both relevant chunks at top positions)
# Query 3: 0.0 (relevant chunk at position 4, outside top 3)

How it works

NDCG@K applies a logarithmic discount to lower-ranked positions, so a relevant chunk at position 1 contributes much more to the score than the same chunk at position 5.

Formula:

DCG@K  = Σ  relevance(i) / log₂(i + 1)     for i = 1 to K
NDCG@K = DCG@K / IDCG@K

Where:

  • relevance(i) is 1 if the item at position i is in the ground truth, 0 otherwise
  • IDCG@K (Ideal DCG) is the best possible DCG if all relevant items were ranked first
  • Duplicate items in the retrieved list are only credited once

A score of 1.0 means the retriever placed all relevant chunks at the very top in the best possible order. A lower score means relevant chunks are buried below irrelevant ones.

By default (without eval_config), the evaluator uses the full retrieved list. Pass eval_config={"k": N} to limit evaluation to the top N chunks. Matching is based on exact string equality.

Tip

Pass eval_config={"k": N} to evaluate only the top N retrieved chunks. For example, eval_config={"k": 3} measures ranking quality within the first 3 results only.

When to use

Run NDCG@K wherever the position of relevant chunks in the ranking matters, not just whether they were retrieved.

  • RAG and retrieval pipelines, to check whether the most relevant chunks surface near the top
  • Evaluating or tuning a re-ranking step, since NDCG@K directly rewards better ordering
  • Comparing retrieval strategies where two approaches find the same chunks but rank them differently

What to do when NDCG@K fails

If NDCG@K is low, relevant chunks are being retrieved but ranked poorly:

  • Apply a re-ranking model (cross-encoder) to reorder results by relevance after initial retrieval
  • Fine-tune the embedding model on domain-specific data to improve ranking accuracy
  • Check if your similarity metric (cosine, dot product) is appropriate for your embedding model
  • Consider using a hybrid retrieval approach where sparse (BM25) and dense scores are combined for better ranking
  • Review query preprocessing: adding context to short queries can improve ranking quality
Was this page helpful?

Questions & Discussion