Metkagram · Evaluation
Language Pattern Retrieval Benchmark
The benchmark tests one narrow job: can a system map a natural communicative goal to the intended Metkagram intent, reasoning move and an acceptable canonical pattern?
The displayed results belong to the bundled deterministic resolver and are an internal regression signal, not an independent model evaluation or evidence of learning efficacy.
What the system receives
Only the natural-language query. The system should rank an intent, identify the reasoning move and return up to three canonical pattern IDs.
Metrics
Intent top-1
The expected intent is ranked first.
Move top-1
The expected reasoning move is ranked first.
Pattern hit@3
At least one editorially acceptable pattern is present in the first three results.
How to report a run
- Record the dataset version and run date.
- Name the system/model, version and prompt/retrieval configuration.
- State whether the system had access to the Metkagram taxonomy, API or corpus.
- Report all three metrics and misses/abstentions.
Limitations
Gold labels are public, and both the benchmark and bundled resolver are maintained by the same project. That is useful for reproducibility and regression testing, but unsuitable for grand claims about outperforming LLMs. External validation needs an independently authored held-out set.