Skip to content
NakodaAI

AI FIELD MANUALRESEARCHERS, STUDENTS

A grad student needs forty real sources for a literature review, not forty plausible-sounding ones

Why asking an LLM to "write a literature review with citations" produces fabricated references, and the retrieve-then-summarize workflow that avoids it.

Last reviewed 1 September 2026

THE PROBLEM

A master's student writing the literature review section of a thesis on urban heat mitigation asks ChatGPT to "summarize the key literature on cool-roof interventions with citations." The response reads like a competent literature review: named authors, plausible journal titles, specific years. Roughly a third of the citations, when checked, do not exist - the model generated statistically plausible author names and titles in the shape of a citation, not references to real papers it retrieved.

This is not a rare glitch. It's a predictable consequence of how a general-purpose chat model generates text: it produces the next most probable token, and a citation-shaped string of tokens is exactly as available to it as a real one, whether or not the underlying paper exists. Submitting this list without checking every entry is the single fastest way to damage a thesis defense.

The problem the student actually has is a workflow problem: they asked a generation tool to do a retrieval task.

THE APPROACH

Retrieve real papers first, from a tool that searches an actual indexed corpus of academic literature, before any AI-generated summarization happens. Only summarize text the tool actually retrieved and can show you the source for - never let a model produce a citation from memory.

This is the retrieve-then-generate pattern again, but for citations specifically: the retrieval step (find real papers) and the generation step (summarize what they say) must stay separated, with a verification check between them.

WHY IT WORKS

Tools built for literature search (Elicit, Consensus, Semantic Scholar) query an actual database of indexed papers and return real records - the paper you see is the paper that exists, because the tool is doing lookup, not free generation. A general chat model has no such database behind every claim; it is producing text that resembles a citation, which is a different operation entirely.

Verifying at the retrieval stage is far cheaper than verifying at the end: checking that a search tool's ten results are real papers (click through, confirm they exist) takes minutes; discovering that a third of a finished forty-source bibliography is fabricated after the fact means redoing the whole review.

STEP BY STEP

  1. 01

    Retrieve

    Search an actual indexed corpus (Elicit, Consensus, Semantic Scholar) - real records only.

  2. 02

    Verify

    Click through and confirm each paper exists and matches what the tool claims.

  3. 03

    Extract

    Pull only findings actually present in the verified paper's text.

  4. 04

    Synthesize

    LLM drafts narrative from verified material only - explicitly no new citations.

  5. 05

    Re-verify

    Every citation in the final draft checked against the verified list before submission.

  1. Search with a literature-specific tool, not a chat model

    Use Elicit, Consensus, or Semantic Scholar to search the actual research question. These return real, indexed papers with links, not generated text.

  2. Verify each result resolves

    Click through a sample of the results (or all of them, for a review under ~50 sources) and confirm the paper exists, matches the title and authors shown, and is actually about what the tool says it's about.

  3. Extract only what the paper actually says

    For each verified paper, pull the specific finding relevant to your review - either read the abstract/full text yourself, or use the tool's own extraction feature (Elicit's structured data extraction reads the actual paper text, it doesn't generate from memory).

  4. Only now bring in a general LLM, for synthesis of verified material

    Paste the verified list of papers and their extracted findings into Claude or ChatGPT and ask it to draft the narrative synthesis - explicitly instructed to reference only the papers provided and add no additional citation.

  5. Re-verify the draft's citation list against your verified list

    Before submitting, confirm every citation that appears in the drafted narrative is one of the papers you verified in step 2 - not a new one the model reintroduced during synthesis, which does happen even with explicit instructions not to.

  6. Format citations from the real records, not from the draft

    Pull the final citation details (exact title, authors, year, DOI) from the original tool's record or the publisher page, not from whatever the LLM typed - transcription drift in an AI-generated bibliography is common and easy to miss.

TOOLS

LIMITATIONS

  • Literature-search tools have their own coverage gaps - non-English literature, very recent preprints, and paywalled or non-indexed journals may be underrepresented. A field with a lot of non-English or grey literature needs manual search to supplement this workflow, not replace it.

  • "Retrieved and real" is not the same as "correctly interpreted." A tool's auto-summary of a paper can still misstate a finding's strength or scope; reading the actual abstract (not just the tool's summary of it) for any source doing real work in your argument is worth the extra few minutes.

  • This workflow prevents fabricated citations. It does not by itself prevent cherry-picking - a search tool will return what matches your query, and a literature review that only searches for confirming evidence will find it. Deliberately search for the strongest counter-arguments too.

EXAMPLE

A literature review section needs ~15 sources on cool-roof urban heat mitigation interventions.

  1. Elicit search: "cool roof reflective surface urban heat island mitigation effectiveness" returns 20 candidate papers with abstracts and links.

  2. Each of the 20 checked by clicking through; 18 resolve to real, matching papers, 2 are dropped (one had a mismatched abstract, one link 404'd).

  3. Elicit's extraction pulls the reported temperature-reduction finding from each of the 18 verified papers into a table.

  4. Claude drafts a three-paragraph synthesis from the 18-paper table, instructed to cite only those papers.

  5. Final check: all citations in the draft cross-referenced against the 18 verified records - one citation the model had subtly reworded (wrong year) was caught and corrected from the original record.

An 18-source literature review section where every citation is a real, verified paper, built in under two hours instead of the half-day a fully manual search would have taken - without the fabrication risk of asking a chat model to write it directly.

RELATED

Nakoda editorial · last reviewed

This entry describes a workflow Nakoda recommends - it is not a claim about how any named tool behaves in every case, and it is not paid placement. Spotted something out of date? Tell us.

MORE FOR THIS AUDIENCE