> Canonical: https://drkishanrees.com/dispatches/the-first-prompt/

# The First Prompt: My earliest ChatGPT conversation and the problem of fabricated references

> **Machine-facing title.** This Markdown companion uses the expanded browser and structured-data title. The canonical HTML page preserves the human-facing headline: “Dispatches from the Archive 001: The First Prompt”.

*Dr Kishan Rees returns to his first prompt to ChatGPT, sent on 20 December 2022, and to the academic-reference test that still explains why he checks the sources.*

By Dr Kishan Rees

With contributions from OpenAI GPT-5.6 Sol (Codex) and Anthropic Opus 5 (Claude Code).

## Archive record

- Historical record date: 20 December 2022
- Original platform: ChatGPT

> Source note: The historical record is the ChatGPT conversation “Post-COVID Healthcare Trends”, preserved in the original export. The article distinguishes that record from present-day interpretation and verification.

## Evidence profile

- **Historical record.** The earliest ChatGPT conversation in the preserved export, dated 20 December 2022.
- **Verbatim evidence.** The five-part first prompt and the later instruction to insert academic references.
- **Present-day verification.** A re-check of the final submitted thesis found no evidence that the questionable references entered its bibliography.
- **Human interpretation.** The author distinguishes what the historical record proves from the context he now recalls.
- **Related concept or outcome.** The account connects the early test to doctoral research practice, including EndNote and human critical review.

I have recently started building an Obsidian vault from several years of
conversations with large language models. The original exports are being
preserved separately, while working copies are converted into a searchable and
interconnected archive. At the moment, that archive contains 1,669 conversations:
962 with ChatGPT, spanning 20 December 2022 to 4 August 2026, and 707 with
Claude, spanning 19 July 2023 to 1 August 2026. Even that figure is incomplete.
My conversations with Claude Code and Codex have not yet been exported and
incorporated into the archive.

I began doing this largely because I wanted to take ownership of a body of
material that had accumulated almost without me noticing. Until now, most of
these conversations had existed wherever the platforms happened to put them:
accessible through search or an endlessly scrolling sidebar, but hardly arranged
in a way that encouraged systematic reflection. Putting them into an archive
changes their character. Conversations separated by months or years can sit
alongside one another. Ideas can be traced backwards. Repeated questions become
visible. What previously felt ephemeral begins to resemble a longitudinal record,
not just of what the models could do at different points in time, but of what I
was asking them to do and how my own expectations of them were changing.

One of the first things I wanted to know was where that record began.

My ChatGPT export contains 962 conversations, and the earliest dates from
20 December 2022, less than three weeks after ChatGPT was released publicly. The
conversation was titled Post-COVID Healthcare Trends. At 13:59 that afternoon, I
typed my first prompt into ChatGPT:

> ⦿ What do HCPs and Patients want after COVID-19?
> ⦿ Disruption has become the new normal; are we ready?
> ⦿ How do Medical Affairs embrace disruption and augment human-only traits in the organization?
> ⦿ How does disruption help differentiate strategies?
> ⦿ Emerging trends and disruption opportunities.

There is something reassuringly consistent about those questions when I read them
again nearly four years later. Healthcare, communication, technology and
organisational change remain subjects I spend a great deal of time thinking
about. What has changed dramatically is the sophistication of the machines I use
to think alongside, and perhaps equally importantly, the sophistication of the
relationship I have developed with them.

It was what happened less than two hours later that really caught my attention.

ChatGPT was new, but the murmurings about its limitations had already begun.
Among them was a particularly troubling claim for anyone working in academia:
apparently, if you asked it for references, it could simply make them up. Not get
a reference slightly wrong or confuse two papers, but produce something that had
all the surface characteristics of an academic citation while referring to a
source that did not exist.

I wanted to see that for myself.

So I took some of the material ChatGPT had generated and gave it a deliberately
simple instruction:

> “insert academic references into the following text”

It did.

The answer looked academic. Claims were accompanied by names and dates: Doshi et
al., Zhou et al., Kuo et al., Fujimoto et al., Teece et al., Wang et al. and Hood
et al., alongside organisations including the World Health Organization and the
Centers for Disease Control and Prevention. A bibliography began underneath.

The significance of that exchange is therefore slightly different from how it
might look if the prompt were read without its original context. I wasn’t
building a reference list for a piece of academic work and blindly trusting
ChatGPT to do the scholarship for me. I was testing something I had heard about
this strange new system. Could it really manufacture academic authority
convincingly enough that the result looked real?

Nearly four years later, with the original exchange sitting in front of me again,
I decided to re-check what it had actually produced.

Some of the sources were genuine. Some pointed towards real bodies of evidence
but were inadequately specified or incorrectly attributed. Others could not be
substantiated in the form in which ChatGPT had presented them. One was
particularly interesting. The model began constructing a reference attributed to
Peter Doshi, Tom Jefferson and Daniela Rivetti. These are real researchers, with
genuine scholarly connections between their work, but I could not verify the
particular reference ChatGPT appeared to be constructing.

In some ways, that is more revealing than simply inventing three fictional
academics. The names were credible because they were credible. The surrounding
subject matter was credible. Many of the underlying propositions in the text were
also reasonable: COVID-19 had disrupted healthcare access, telemedicine had
expanded rapidly and healthcare systems were adopting digital technologies at
pace. The model appeared capable of assembling pieces of a credible scholarly
landscape without reliably preserving the provenance required to demonstrate that
a particular citation actually supported a particular claim.

So the murmurings I had heard in December 2022 had substance. ChatGPT could
produce something that occupied the form of academic evidence without necessarily
having the evidential foundations that form implied. Plausibility and provenance
were two different things.

What happened to me in 2026, however, is probably the part I find most
interesting.

I completed my medical doctorate two and a half years after that first conversation.
References were not something I simply asked an LLM to generate and pasted into a
thesis. Academic writing involved the much less glamorous machinery of
conventional scholarship: building and maintaining a reference library in
EndNote, returning to papers, checking bibliographic information, matching
references to claims, reviewing the bibliography, and then checking it again. My
thesis went through supervision, examination and the usual processes surrounding
doctoral research. The final submitted document also explicitly describes my use
of generative AI within parts of the research process and the importance of human
critical review of its outputs.

I knew all of that. More importantly, I remembered doing it. I had checked,
checked and checked again.

And yet, when I rediscovered those fabricated or questionable citations from
December 2022, a thought appeared almost immediately: Oh my God. What if one of
them is in the thesis?

Rationally, I was about 99 per cent certain they weren’t. The original
interaction had been an experiment. I knew how I had subsequently managed the
literature for my doctorate. I had used EndNote rather than treating an
LLM-generated bibliography as a source of truth. I had returned to original
sources and checked that references supported the claims for which I was using
them. There were multiple layers of review between the first exploratory ChatGPT
conversation in 2022 and a submitted doctoral thesis.

But suddenly 99 per cent did not feel sufficient.

So I checked.

I went back to the final submitted thesis and specifically searched for the
questionable references from that first conversation. They had not propagated
into it. There are incidental overlaps involving common author names and genuine
papers, but I found no evidence that the fabricated or misattributed references
generated during that first afternoon with ChatGPT had entered the final
bibliography.

The relief was less interesting than the fact that I had felt compelled to check
at all.

There is something here about the peculiar persuasiveness of generative AI. A
fluent model does not necessarily persuade us by making outrageous claims. Often
it does something subtler: it produces something close enough to the shape of
knowledge that even years later, and even when I knew the verification processes
I had followed, seeing those references again was sufficient to introduce a
nagging doubt. I knew the provenance of my doctoral work. I knew the processes I
had followed. I knew that I had repeatedly checked the references. Yet the output
of a model from four years earlier was persuasive enough to make me reopen the
finished thesis and establish the answer for myself.

That doubt is probably healthy. It is also an argument for some distinctly
unfashionable tools and habits.

EndNote is not exciting. Manually checking whether a paper exists is not
exciting. Opening the source and making sure it actually supports the sentence
beside which it is cited is certainly not exciting. Neither are stable
bibliographic records, versioned documents, supervisors’ comments, or the slow
accumulation of a reference library over the course of a research project.
Compared with asking a language model to generate a perfectly formatted
bibliography in seconds, these methods can feel almost analogue.

Yet that friction is part of their value.

A reference manager is not there to generate the appearance of scholarship. Used
properly, it maintains a relationship between the document being written and a
source that the researcher has actually encountered and recorded. Checking a
reference requires an action outside the generative system. Following a DOI,
opening a paper, reading the relevant section and deciding whether it supports a
claim creates a chain of provenance that fluency alone cannot provide.

There is also something reassuring about the fact that those older methods leave
traces. Nearly four years after the first ChatGPT experiment, I did not have to
ask another language model whether I had probably been careful. I could inspect
the finished thesis. I could search its bibliography. I could return to the
sources. The infrastructure of conventional academic practice provided an
independent record against which the machine-generated material could be tested.

The lesson I take from that first conversation therefore isn’t that academics
should avoid generative AI. My own subsequent history would make that a difficult
argument to sustain. Generative AI eventually became an explicitly documented
part of my doctoral research process, and I now use considerably more capable
systems for forms of research, analysis, writing, coding and knowledge work that
would have seemed extraordinary to me in December 2022.

The lesson is almost the opposite. The more capable and persuasive these systems
become, the more valuable some of our older scholarly disciplines may become.
Verification does not become obsolete because generation gets better, and
provenance does not matter less because a model is usually right. A reference does
not become trustworthy because the title, authors, journal and DOI all look
exactly as we expect them to look. If anything, increasing fluency raises the
premium on having methods that exist independently of the model and allow us to
establish what is actually true.

This is also why the Obsidian project is becoming more interesting to me than I
initially expected. I thought I was organising old conversations. Instead, I seem
to be assembling evidence of two things developing simultaneously: the
capabilities of the models and my own literacy in using them.

The first conversation provides a useful baseline. In December 2022, I was
sufficiently intrigued by reports of ChatGPT hallucinating academic references
that, within hours of first using it, I deliberately tested the claim. By the
time I completed my doctorate, generative AI could be incorporated explicitly
into a research methodology, but within a system that retained human critical
review and conventional mechanisms of academic verification. By August 2026, I am
using another generation of these systems to help preserve and interrogate the
history of those interactions themselves.

That evolution is difficult to reconstruct from memory. The archive makes at
least part of it observable. It allows me to ask when I first trusted a model
with a particular kind of task, when I first tested something I did not trust,
how my prompts became more sophisticated, which ideas recurred over time, and
which failure modes disappeared or persisted as the technology changed.
Eventually, it may be possible to trace ideas from an initial conversation
through subsequent iterations and into projects, publications or decisions in the
physical world.

There are limits to what such an archive can tell me. A conversation history is
not a transcript of thought, and retrospective interpretation brings its own
biases. The prompt itself, for example, proves that I asked ChatGPT to insert
references; it does not prove why I asked. The context that I had heard
murmurings about fabricated references and wanted to see the behaviour for myself
comes from my recollection now. That distinction matters, particularly in a
project concerned with provenance.

It is one reason I am preserving the original exports separately from the working
Obsidian vault. I don’t want to turn several years of ChatGPT and Claude
conversations into an infallible autobiography. I want to preserve the original
record, distinguish it from subsequent interpretation, and then use the working
archive to interrogate it.

There is a pleasing symmetry to that. The weakness I was testing in my first
ChatGPT conversation was a failure of provenance: an answer could look convincing
even when the evidential chain underneath it did not withstand scrutiny. Nearly
four years later, provenance is precisely what makes the archive useful. I don’t
have to rely solely on my recollection of how I used ChatGPT in 2022. I can
return to the record, see what I actually asked, see what the model actually
answered, and distinguish that contemporary evidence from the context I now
remember surrounding it.

And the scale of that record is becoming significant. I am not casually browsing
through some old chats. I am beginning a longitudinal analysis of a corpus of at
least 1,669 LLM conversations spanning 1,324 days: 962 with ChatGPT and 707 with
Claude. Even that corpus is incomplete. My conversations with Claude Code and
Codex have not yet been exported and incorporated into the archive.

I began building the vault because I wanted a better way to preserve my
conversations with artificial intelligence. I am starting to think it may contain
something more valuable: a longitudinal record of how I learnt to work with it.

And, occasionally, a reminder of why I still check the references.

Cite as: Rees, Kishan. 'Dispatches from the Archive 001: The First Prompt'. drkishanrees.com, first published 8 August 2026. https://drkishanrees.com/dispatches/the-first-prompt/

## Production note

- **First published:** 8 August 2026.
- **Last amended:** 13 August 2026.
- **OpenAI GPT-5.6 Sol (Codex):** Archive-series concept, Dispatches index prototype and implementation handoff.
- **Anthropic Opus 5 (Claude Code):** Independent second-desk review.
- **Final editorial responsibility:** Dr Kishan Rees.
