Metkagram · internal research programme
Measure the method, not just describe it.
Metkagram maintains an applied research programme at the intersection of second-language learning, retrieval practice and functional annotation. Hypotheses are recorded before results exist, so future claims can be tested rather than retrofitted.
Status · protocols prepared, no efficacy claim made
- 3,484
- reusable B2–C1 patterns
- 969
- annotated sentences
- 72
- curated document sets
- EN · DE
- learning languages
Evidence notes
What outside research tells us
These studies inform questions for Metkagram. They do not prove that the Metkagram interface works better than another method.
- 01
Keep visual cues small
A meta-analysis of visual input enhancement found a small average benefit for grammar learning, but it also found a possible cost for processing meaning. For Metkagram, this is a reason to test one useful tag against full markup instead of assuming that more highlighting is better.
Source: Lee & Huang (2008) · Visual input enhancement and grammar learning · Open paper - 02
Read for meaning before opening more help
A 2025 study on phrasal verbs found stronger learning when definitions came after reading rather than before it, and typographic enhancement supported contextual learning. This suggests a useful Metkagram test: keep the sentence readable first, then reveal the tag or explanation when the learner needs it.
Source: Contextual learning and retention of phrasal verbs (2025) · Open paper - 03
Pattern order may not work the same for everyone
Research on learning second-language constructions shows that balanced and strongly repeated input can work differently depending on the learner and the task. Metkagram should test how examples are ordered instead of treating one sequence as correct for every learner.
Source: Pulido (2024) · Optimizing input for L2-specific constructions · Open paper - 04
Spacing matters, but the schedule is still a design choice
A 2025 online replication found that spaced study sessions outperformed massed study on a delayed L2 vocabulary test. Earlier work also shows that different spaced schedules can end with similar final recall. Metkagram should test the gap between sessions instead of copying one universal SRS rule.
Source: Rogers, Nakata & Chiu (2025) · Optimizing distributed practice online · Open paper - 05
Report results by pattern family
A large study of captions and textual enhancement found clear learning effects for some targets but not for every grammar structure. If Metkagram tests the method, results should be shown by pattern family as well as in total. One average score can hide where the method helps and where it does not.
Source: Investigating textual enhancement and captions in L2 grammar and vocabulary (2020) · Open paper - 06
A cue must be noticeable enough to be processed
A 2026 artificial-language study found sensitivity to morphosyntactic violations only in the high-salience condition. This does not tell us the correct Metkagram colour or tag strength, but it gives a direct reason to test salience instead of treating visual emphasis as decoration.
Source: Fernández Santos et al. (2026) · Morphological salience in initial L2 learning · Open paper - 07
Build larger patterns from smaller chunks
A 2026 study of multi-word expressions found that L2 learners had difficulty binding relatively large chunks and argued for gradual binding of smaller units into larger ones. Metkagram can test whether a stable short chunk should appear before the full reusable frame.
Source: Wei, Wang & MacWhinney (2026) · Chunking multi-word expressions · Open paper - 08
Practice may need to change as a pattern becomes a skill
A 2025 study found evidence consistent with declarative, procedural and automatic stages in deliberate L2 skill learning. Early performance related more to declarative learning ability, while procedural ability became more important later. A learner who understands a pattern may therefore need a different task from a learner trying to use it quickly.
Source: Maie & Godfroid (2025) · Three-stage model of L2 skill acquisition · Open paper - 09
More feedback is not always better feedback
A 2025 experiment with semi-open-ended language questions found that concise immediate correct-response feedback worked better than more elaborate feedback in that task and also supported confidence judgments. Metkagram should keep detailed explanations available without forcing them after every attempt.
Source: From belief to evidence (2025) · Immediate feedback complexity · Open paper - 10
Visual notation needs a non-visual path
A 2026 study with visually impaired and sighted learners found short-term benefits from aural input enhancement, with results moderated by proficiency and not maintained at delayed testing. The direct target was vocabulary, not grammar, but the accessibility lesson is useful: important Metkagram cues should never depend on colour alone.
Source: Badri, Graham & Zhang (2026) · Aural cues for visually impaired learners · Open paper
Live pilot · H1-CUE-UTILITY-V1
The first pilot is ready to run
Compare a clean sentence with compact functional notation. The short session randomly assigns one condition and keeps results local unless the participant exports them.
Run the H1 pilot01 · Research questions
Research questions
- 01
Cue and attention
Does a minimal functional tag help learners identify structure faster than the same sentence without annotation?
Primary measure: Role-identification accuracy and time to a correct answer. - 02
Pattern and transfer
Do parallel variations support use in a new context better than a rule explanation alone?
Primary measure: Accuracy on unseen, structurally matched examples. - 03
Retrieval and retention
Does attempting the pattern before revealing it improve delayed recall?
Primary measure: Immediate and delayed cued-production scores. - 04
Annotation and data quality
Can one token-level scheme serve learning and reproducible NLP analysis at the same time?
Primary measure: Annotator agreement, offset validity and record completeness.
02 · Minimum reproducible protocol
Minimum reproducible protocol
- 01Pre-register the hypothesis, primary outcome and exclusion criteria.
- 02Vary one mechanism at a time: annotation, variation or retrieval.
- 03Measure transfer with unseen examples, not only the phrases used in practice.
- 04Publish null and negative findings alongside positive results.
- 05Protect participant privacy and collect only the data the study needs.
- 06Version the stimuli, annotation schema and analysis code.
Experiment queue
Small tests that can change the product
Each test changes one main mechanism. Transfer means using the same structure correctly in a new example, not repeating a sentence from memory.
- 01
R01 · Tag density
Compare: Compare a clean sentence, one target tag and fuller markup.
Measure: Role accuracy, sentence comprehension and response time.
Product decision: Choose the default amount of visible annotation. - 02
R02 · Progressive reveal
Compare: Compare tags shown immediately with tags revealed after the learner reads the sentence.
Measure: Meaning comprehension and delayed recall of the target form.
Product decision: Decide whether the interface should be sentence-first by default. - 03
R03 · Variation and transfer
Compare: Compare repeated examples from one context with examples that change person, tense and situation.
Measure: Production accuracy on a new sentence that was never shown in practice.
Product decision: Set a minimum diversity rule for published pattern examples. - 04
R04 · Pattern order
Compare: Compare a balanced sequence with a sequence that repeats one common form more strongly at the start.
Measure: Learning speed and transfer, reported by proficiency band.
Product decision: Decide whether pattern sets need different learning paths. - 05
R05 · Recall before feedback
Compare: Compare rereading with an attempt to produce the pattern before the answer appears.
Measure: Immediate and delayed cued production.
Product decision: Decide where active recall belongs in Pattern Practice. - 06
R06 · Annotation agreement
Compare: Give the same sentence sample to two independent annotators using the same guide.
Measure: Agreement by label, span boundary errors and adjudicated corrections.
Product decision: Revise weak labels before expanding the corpus. - 07
R07 · Salience strength
Compare: Compare no cue, a restrained cue and a strong cue for one target grammar role.
Measure: Detection of the target form, sentence comprehension and visual-load rating.
Product decision: Set evidence-based contrast and emphasis rules for tags. - 08
R08 · Progressive chunking
Compare: Compare the full pattern from the start with small chunk → larger chunk → full frame.
Measure: Immediate and delayed production, including where errors occur inside the pattern.
Product decision: Decide whether Pattern Practice should grow chunks in stages. - 09
R09 · Stage-aware practice
Compare: Compare one repeated task format with a path from explanation to recognition to timed production.
Measure: Accuracy and response speed across repeated sessions.
Product decision: Change the task when a pattern is accurate but still slow, if the data supports it. - 10
R10 · Feedback depth
Compare: Compare correct answer only, a short explanation and a detailed explanation.
Measure: Accuracy on the next unseen item, delayed accuracy, confidence and feedback time.
Product decision: Choose the default feedback depth and keep extra detail optional when possible. - 11
R11 · Cue modality
Compare: Compare the same target role through a visual tag, an accessible text label and an optional aural cue.
Measure: Role identification, later recall and usability reports.
Product decision: Make important notation understandable without colour alone. - 12
R12 · Review gap
Compare: Compare massed practice with two or more spaced schedules.
Measure: Delayed production after a fixed retention interval, practice accuracy and review completion.
Product decision: Tune review intervals from Metkagram evidence instead of copying one SRS schedule.
03 · Research-ready assets
Research-ready assets
A static corpus, canonical schema and public routes make learning stimuli inspectable and reproducible.
Measurement
How we should know if an idea works
A research page is more useful when it also defines what counts as evidence.
- 01
Measure transfer, not only memory
Use unseen production items so a learner must apply the structure in a new sentence.
- 02
Separate accuracy from speed
A learner can know a rule and still use it too slowly for fluent speech. Track response time only where the task makes timing meaningful.
- 03
Use web experiments carefully
A 2024 study found that web-based elicited imitation can provide useful morphosyntactic measurement comparable to a lab version, which makes small remote pilots realistic.
- 04
Keep comprehension as a guardrail
Whenever annotation strength changes, test whether the learner still understands the sentence. Noticing grammar at the cost of meaning is not a useful win.
- 05
Automate scoring only after calibration
Recent NLP work shows that automated L2 error analysis can agree strongly with human ratings, but context-dependent errors remain difficult. Human checks should stay in the validation loop.
Decision rules
How evidence can change Metkagram
Research is useful only when it can change the design.
- 01If extra tags improve structure detection but reduce sentence comprehension, use fewer tags by default.
- 02If sentence-first reading improves meaning and later form recall, keep explanations closed until after the first read.
- 03If varied examples improve transfer, require every published pattern to cover meaningful structural variation, not cosmetic rewrites.
- 04If different proficiency groups need different example order, build separate paths instead of one universal sequence.
- 05If stronger salience improves noticing without harming comprehension, use it only for the target role instead of highlighting the whole sentence.
- 06If progressive chunking improves full-pattern production, add chunk growth as a practice mode.
- 07If staged practice improves speed after accuracy is already high, change the task instead of simply adding more repetitions.
- 08If concise feedback works as well as or better than long explanations, keep the default feedback short and make detail optional.
- 09If an important cue cannot be understood without colour, redesign it.
- 10If annotation agreement is weak for a label, revise the label and guide before adding more data.
- 11If a result is small, mixed or negative, keep it in the research record. A useful method should be allowed to discover its own limits.
04 · Evidence boundary
Evidence boundary
Metkagram does not yet claim that its interface outperforms another learning method. Research on attention, focus on form, retrieval and spacing informs the design; the effect of this specific implementation must be measured separately.
05 · Propose a study
Propose a study
We welcome preregistered pilots, corpus studies, annotation-quality reviews and supervised student research.
Discuss a research partnership