Measuring source fidelity in AI-generated flashcards

This internal study tests quote matching in Quanta's Source-First pipeline: 1,997 of 2,042 generated candidates passed. It measures text binding, not factual accuracy.

100% source-backed · 97.8% pass the quote-match · 2,042 cards

Abstract

Across 100 generations (50 exam topics from upper-secondary and university level, 2 runs each) Quanta's source-first pipeline produced 2,042 flashcards. 45 did not pass the deterministic quote-match against their source and were dropped server-side; 1,997 were delivered, a hit rate of 97.8%. All 1,997 delivered cards carry a source evidence quote (100%, by construction). The study measures source fidelity (where a statement comes from), not factual correctness. Every number is recomputable from the open raw dataset. Measured on a German corpus; the method is language-neutral.

From 2,042 generated to 1,997 backed cards

2,042
generated & checked
100 generations, target 20 cards/run
−45
dropped
evidence not found, never shown
1,997
delivered
100% with source evidence
97.8%
Hit rate (quote-match)
delivered / generated = 1,997 / 2,042
100%
of delivered cards source-backed
by construction: the unbacked is dropped

As of June 28, 2026. Full computation under "Results" below.

The evidence stays where you need it: on the card.

The study measures the same gate used before delivery in the Source-First path. Quanta retrieves real source text, generates from it and keeps the passage attached to the study content.

  1. 01You start with a topic or your own material.
  2. 02Quanta links every delivered Source-First card to its source and passage.
  3. 03If the evidence cannot be confirmed in that text, the candidate is dropped.

Confirmed evidence makes a statement traceable. It is not an automatic fact check.

Live search · Alpha-Zerfall

Serlo · „Radioaktivität"

Picked

Wikipedia · „Alphastrahlung"

Wikibooks · Physik

Quanta searches vetted, openly licensed knowledge sources for your topic and picks the best passage.

Source fidelity, not correctness

Definition first, so the number stays honest.

The quote-match hit rate is the share of model-generated, well-formed cards whose evidence quote a deterministic algorithm finds in the source text (exact, normalized or contiguous token match). It is not an accuracy or correctness rate; it measures provenance.

The decisive part is source selection: Quanta generates exclusively from established, vetted, openly licensed educational sources (see below) and makes every card traceable back to one. That reverses the burden of proof: instead of claiming "correct", we show you the vetted source and the exact location for every card. You can verify correctness yourself rather than taking our word for it.

hit rate = backed / (backed + dropped) = 97.8%
backed = 1,997, dropped = 45, generated (well-formed) = 2,042.

Source-first: the source before the card

Sample

German upper-secondary (Abitur) and bachelor level. 50 typical exam/test topics across the subject spectrum, no adversarial inputs.

Retrieve the source

Per topic, real full text from verified, openly licensed sources (Serlo, Wikibooks, Wikipedia, Wikiversity). Substance and license gate per source.

Generate extractively

Cards are produced ONLY from that text (temperature 0, no model knowledge). Quote-first: the verbatim anchor first, then the question and answer around it.

Verify deterministically

Every card is checked by quote-match against the source (no second AI model): exact substring, punctuation/formula-normalized (>=0.9), token containment (>=0.85). The unbacked is dropped.

Parameters & assumptions

Generation model (28 June 2026)
Google Gemini (card generation), temperature 0, thinkingBudget 0
AI provider today
Since 22 August 2026, Mistral AI SAS, based in Paris; the measurement documented here ran on the generation model of the time and has not been repeated since
Verification
Deterministic quote-match (no second AI model)
Match tiers
exact (score 1.0) · normalized (0.95–0.9) · token containment (>=0.85)
Minimum quote length
16 characters (shortest delivered: 17)
Sources
Serlo, Wikibooks, Wikipedia, Wikiversity (all CC-BY-SA 4.0)
Sample
50 exam topics (25 Abitur + 25 university) x 2 runs = 100 generations
Target per generation
20 cards (reached 97x, 19 cards 3x after dropping)
Corpus
German exam-prep topics and German open educational sources (June 2026 snapshot)

Source-backed rate by level and subject

Absolute numbers per group. Rate = delivered / generated cards.

By level

Leveln (gen.)generateddelivereddroppedrate
Abitur (upper-secondary)50101110001198.9%
University (bachelor)5010319973496.7%

By subject

Subjectn (gen.)generateddelivereddroppedrate
Economics81601600100%
Medicine61201200100%
Philosophy480800100%
Geography240400100%
Chemistry12239238199.6%
History8162160298.8%
Physics14284280498.6%
Biology14284280498.6%
Law48280297.6%
Mathematics163363191794.9%
Computer Science122552401594.1%

Weakest fields: mathematics (94.9%) and computer science (94.1%), formula- and algorithm-heavy material with thinner open German source coverage. No subject below 94%.

How it is computed

hit rate = delivered / generated

Total   = 1,997 / 2,042 = 0.9780 = 97.8%

Abitur  = 1,000 / 1,011 = 0.9891 = 98.9%

Univ.   = 0,997 / 1,031 = 0.9670 = 96.7%

generated = delivered + dropped (well-formed cards). Non-parseable model outputs in the denominator: measured 0.

Hit rate by subject

Share of backed cards per subject, descending. Honest 0–100% scale (not truncated); amber marks fields under 96%.

96% and above under 96% (mathematics, computer science)

What happened to the 45 dropped cards

The score distribution shows: almost all drops were clearly unbacked, barely any borderline cases.

44
score 0
evidence quote not found (contiguously) in the source; clearly unbacked or stitched from several places
1
score 0.73
the only borderline case, just under the 0.85 threshold
Why this supports the study: the 0.85 threshold is not chosen after the fact to fit: 44 of 45 drops sit at score 0, far beyond any threshold debate. There are practically no borderline decisions. The system drops stricter rather than laxer: the one 0.73 case was factually correct but failed on formatting; better one card too few than one unbacked too many. Cards per generation: 97x a full 20, 3x 19 after dropping.

Real cards from the run

Straight from the dataset, including a dropped card. Fronts/backs are the model's German output (English gloss added); the evidence quote is verbatim from the German source.

History · Abitur Verified

Was ist der technologische Kern der dritten industriellen Revolution?

What is the technological core of the third industrial revolution?

Die mikroelektronische Revolution seit Mitte der 1970er Jahre wird als technologischer Kern einer neuen, dritten industriellen Revolution angesehen.

Evidence quote, found verbatim in the German source

"Die mikroelektronische Revolution seit Mitte der 1970er Jahre wird als technologischer Kern einer neuen, dritten industriellen Revolution angesehen"

Source: Wikipedia · "Industrielle Revolution" (CC-BY-SA)

Biology · Abitur Verified

Was versteht man unter Selektion?

What is meant by selection?

Unter Selektion versteht man Prozesse, die den Genpool verändern und bewirken, dass sich Arten umbilden.

Evidence quote, found verbatim in the German source

"Unter Selektion versteht man Prozesse, die den Genpool verändern und bewirken, dass sich Arten umbilden."

Source: Serlo · "Selektion" (CC-BY-SA)

Mathematics · University Dropped, never shown

Was ist die Form einer linearen Funktion?

What is the form of a linear function?

Claimed evidence, not found contiguously in the source

"Eine lineare Funktion hat die Form $f(x)=m\cdot x+b$ . Ihr Graph ist eine Gerade ."

Dropped (score 0.73 < 0.85): the multi-sentence quote fell just below the threshold because of the spaces before the periods. The system drops when in doubt rather than deliver something uncertain (stricter, not laxer).

Inspect all 2,042 measured cards

Search the German measurement corpus by topic, subject, level and verdict. Every delivered card includes its evidence quote and source; dropped candidates include their rejection reason.

The report is translated; the measured cards remain German because this release used a German-language corpus.

Fully traceable

Generation and verification use the same retrieved full text as both the basis and the match corpus (no circular AI-grades-AI setup). The quote-match is pure, deterministic string work; the full parameters are listed above under "Parameters & assumptions".

How to recompute it yourself: download the raw dataset, and for each entry count the cards in ausgeliefert (delivered) and verworfen (dropped). delivered / (delivered + dropped) gives the rate, per topic, summed per subject, per level and overall. Every delivered card carries its evidence quote and sourceUrl, every dropped card its score and reason.

Which sources, and why reliable

Quanta only uses sources that are editorially or community-reviewed, versioned/citation-required and openly licensed. That is exactly what makes the reverse argument hold: a vetted source plus per-card traceability.

SourceLicenseWhy reliable
SerloCC-BY-SA 4.0Non-profit learning platform (Serlo Education e.V.), editorially curated by teachers and subject authors, focused on German STEM didactics.
Wikibooks (DE)CC-BY-SA 4.0Open Wikimedia textbooks, community-reviewed with full version history and a sourcing requirement.
Wikipedia (DE)CC-BY-SA 4.0Citation-required encyclopedia with a flagged-revisions system and complete public version control.
Wikiversity (DE)CC-BY-SA 4.0Wikimedia teaching and course material, open and versioned.

Distribution of the 1,997 citations across sources

Source (host)Citations
de.wikipedia.org800
de.serlo.org605
de.wikibooks.org545
de.wikiversity.org47

What this number does NOT say

  • It checks source backing, not factual correctness. The source itself can contain errors.
  • Only the evidence quote comes verbatim from the source. The front and back are written by the model and can introduce their own inaccuracies, even when the quote is correctly backed.
  • "Backed" covers exact, normalized and token matches (>=85% of words contiguous). A normalized or token match is not character-exact, but identical up to case, punctuation and formula notation.
  • It is not an external test against a curriculum or answer key, and not a statement about learning outcome.
  • Sample: 50 curated, non-adversarial German topics with good open source coverage, not transferable to arbitrary or niche queries (primary n = 50 topics / 100 generations).
  • Sources are vetted for provenance and license, not subject-matter didactics (verified = provenance, not didactics).
  • Scope: the topic-based source-first path. For your own PDF/scan upload, the evidence is checked against your document, not against open sources.
  • A human-labeled validation subsample (matcher precision) is still pending and will follow.

Checkable, not asserted

Across a realistic cross-section of upper-secondary and university material (50 topics, 2 runs, 2,042 cards), Quanta's source-first architecture delivers checkable study material: 97.8% of generated cards pass the deterministic source check, and 100% of delivered cards carry a source citation. The result is stable across both runs and all eleven subjects, none below 94%.

The burden of proof is reversed

Not "trust us, it is correct", but "here is the vetted source and the location". Correctness is made checkable rather than asserted.

The weakness is structural

Lower rates (mathematics, computer science) come from formula- and algorithm-heavy material with thin open German sources. They show up as fewer cards, never as unbacked ones. The gate holds.

The check is not flattered

44 of 45 dropped cards sit at score 0, not just under the threshold. The system drops stricter rather than laxer.

Outlook: next are a human-labeled validation subsample (matcher precision), an English-corpus run, broader and deliberately adversarial topics, better sources for formula-heavy STEM, and a per-card match score in the dataset. The number reported here measures source fidelity, not factual correctness. It is the most honest figure we can prove against the open dataset.

Frequently asked questions about the study

Does "97.8% source-backed" mean 97.8% of the cards are correct?
No. The study measures source fidelity: whether every card is backed by a quote found verbatim in the source. What matters is which sources, and Quanta draws exclusively from established, editorially or community-reviewed, openly licensed educational sources (Serlo, Wikibooks, Wikipedia, Wikiversity). Every delivered card is traceable to one of them (source and evidence quote per card). We deliberately do not claim a correctness rate; we make correctness checkable.
Why are 100% of delivered cards source-backed but only 97.8% pass the check?
The two numbers measure different things. 97.8% is the share of model-generated, well-formed cards (1,997 of 2,042) whose evidence passes the check. The other 45 are dropped server-side and never shown, which is why 100% of the delivered cards carry a source citation.
What does "source-backed" mean, and how verbatim is the evidence?
Every delivered card carries an evidence quote that a deterministic algorithm (not a second AI model) finds in the source text, in three tiers: exact (character for character), normalized (case, punctuation and formula notation unified) or a contiguous token match (at least 85% of the words in the same source window). The front and back are written by the model; the evidence points to the exact location. Every quote is inspectable in the open dataset.
Which corpus was this measured on?
A German exam-prep corpus: 50 typical exam/test topics from German upper-secondary (Abitur) and bachelor study, 2 runs each, drawing on German CC-BY-SA educational platforms. The method (retrieve real source text, generate only from it, verify each quote) is language-neutral; the measured numbers are for this German corpus. An English-corpus run is planned.
What happens when Quanta finds no source for a topic?
It generates no cards from model knowledge, and instead returns an honest message ("no citable source found, upload a PDF or URL"). Across the 50 real topics in this study that never happened: all 100 generations produced source-backed cards.
What are the limits of this measurement?
It checks source backing, not factual correctness, interpretation or learning outcome. The front and back are written by the model and can introduce their own inaccuracies; only the evidence quote comes from the source. Sources are vetted for provenance and license, not subject-matter didactics. The sample is German and curated. A human-labeled validation subsample (matcher precision) is still pending. Scope: the topic-based source-first path; for your own PDF/scan upload the evidence is checked against your document.
AM
Amos Matzke·Founder & Managing Director, Full-Stack Architect · former MINT-EC student·June 28, 2026

Study material you can back up

Source-First cards show their evidence status; without a suitable source no flashcards are created. Eligible users can try Essential for 7 days.