Dispatches from the Archive 001: The First Prompt

Dr Kishan Rees returns to his first prompt to ChatGPT, sent on 20 December 2022, and to the academic-reference test that still explains why he checks the sources.

With contributions from OpenAI GPT-5.6 Sol (Codex) and Anthropic Opus 5 (Claude Code).

Read as: HTML Markdown

Dispatches from the Archive 001 Archive record

Archive record

Historical record date
Original platform
ChatGPT

Source note. The historical record is the ChatGPT conversation “Post-COVID Healthcare Trends”, preserved in the original export. The article distinguishes that record from present-day interpretation and verification.

Evidence profile

  1. Historical record The earliest ChatGPT conversation in the preserved export, dated 20 December 2022.
  2. Verbatim evidence The five-part first prompt and the later instruction to insert academic references.
  3. Present-day verification A re-check of the final submitted thesis found no evidence that the questionable references entered its bibliography.
  4. Human interpretation The author distinguishes what the historical record proves from the context he now recalls.
  5. Related concept or outcome The account connects the early test to doctoral research practice, including EndNote and human critical review.

I have recently started building an Obsidian vault from several years of conversations with large language models. The original exports are being preserved separately, while working copies are converted into a searchable and interconnected archive. At the moment, that archive contains 1,669 conversations: 962 with ChatGPT, spanning 20 December 2022 to 4 August 2026, and 707 with Claude, spanning 19 July 2023 to 1 August 2026. Even that figure is incomplete. My conversations with Claude Code and Codex have not yet been exported and incorporated into the archive.

I began doing this largely because I wanted to take ownership of a body of material that had accumulated almost without me noticing. Until now, most of these conversations had existed wherever the platforms happened to put them: accessible through search or an endlessly scrolling sidebar, but hardly arranged in a way that encouraged systematic reflection. Putting them into an archive changes their character. Conversations separated by months or years can sit alongside one another. Ideas can be traced backwards. Repeated questions become visible. What previously felt ephemeral begins to resemble a longitudinal record, not just of what the models could do at different points in time, but of what I was asking them to do and how my own expectations of them were changing.

One of the first things I wanted to know was where that record began.

My ChatGPT export contains 962 conversations, and the earliest dates from 20 December 2022, less than three weeks after ChatGPT was released publicly. The conversation was titled Post-COVID Healthcare Trends. At 13:59 that afternoon, I typed my first prompt into ChatGPT:

  • ⦿ What do HCPs and Patients want after COVID-19?
  • ⦿ Disruption has become the new normal; are we ready?
  • ⦿ How do Medical Affairs embrace disruption and augment human-only traits in the organization?
  • ⦿ How does disruption help differentiate strategies?
  • ⦿ Emerging trends and disruption opportunities.

There is something reassuringly consistent about those questions when I read them again nearly four years later. Healthcare, communication, technology and organisational change remain subjects I spend a great deal of time thinking about. What has changed dramatically is the sophistication of the machines I use to think alongside, and perhaps equally importantly, the sophistication of the relationship I have developed with them.

It was what happened less than two hours later that really caught my attention.

ChatGPT was new, but the murmurings about its limitations had already begun. Among them was a particularly troubling claim for anyone working in academia: apparently, if you asked it for references, it could simply make them up. Not get a reference slightly wrong or confuse two papers, but produce something that had all the surface characteristics of an academic citation while referring to a source that did not exist.

I wanted to see that for myself.

So I took some of the material ChatGPT had generated and gave it a deliberately simple instruction:

“insert academic references into the following text”

It did.

The answer looked academic. Claims were accompanied by names and dates: Doshi et al., Zhou et al., Kuo et al., Fujimoto et al., Teece et al., Wang et al. and Hood et al., alongside organisations including the World Health Organization and the Centers for Disease Control and Prevention. A bibliography began underneath.

The significance of that exchange is therefore slightly different from how it might look if the prompt were read without its original context. I wasn’t building a reference list for a piece of academic work and blindly trusting ChatGPT to do the scholarship for me. I was testing something I had heard about this strange new system. Could it really manufacture academic authority convincingly enough that the result looked real?

Nearly four years later, with the original exchange sitting in front of me again, I decided to re-check what it had actually produced.

Some of the sources were genuine. Some pointed towards real bodies of evidence but were inadequately specified or incorrectly attributed. Others could not be substantiated in the form in which ChatGPT had presented them. One was particularly interesting. The model began constructing a reference attributed to Peter Doshi, Tom Jefferson and Daniela Rivetti. These are real researchers, with genuine scholarly connections between their work, but I could not verify the particular reference ChatGPT appeared to be constructing.

In some ways, that is more revealing than simply inventing three fictional academics. The names were credible because they were credible. The surrounding subject matter was credible. Many of the underlying propositions in the text were also reasonable: COVID-19 had disrupted healthcare access, telemedicine had expanded rapidly and healthcare systems were adopting digital technologies at pace. The model appeared capable of assembling pieces of a credible scholarly landscape without reliably preserving the provenance required to demonstrate that a particular citation actually supported a particular claim.

So the murmurings I had heard in December 2022 had substance. ChatGPT could produce something that occupied the form of academic evidence without necessarily having the evidential foundations that form implied. Plausibility and provenance were two different things.

What happened to me in 2026, however, is probably the part I find most interesting.

I completed my medical doctorate two and a half years after that first conversation. References were not something I simply asked an LLM to generate and pasted into a thesis. Academic writing involved the much less glamorous machinery of conventional scholarship: building and maintaining a reference library in EndNote, returning to papers, checking bibliographic information, matching references to claims, reviewing the bibliography, and then checking it again. My thesis went through supervision, examination and the usual processes surrounding doctoral research. The final submitted document also explicitly describes my use of generative AI within parts of the research process and the importance of human critical review of its outputs.

I knew all of that. More importantly, I remembered doing it. I had checked, checked and checked again.

And yet, when I rediscovered those fabricated or questionable citations from December 2022, a thought appeared almost immediately: Oh my God. What if one of them is in the thesis?

Rationally, I was about 99 per cent certain they weren’t. The original interaction had been an experiment. I knew how I had subsequently managed the literature for my doctorate. I had used EndNote rather than treating an LLM-generated bibliography as a source of truth. I had returned to original sources and checked that references supported the claims for which I was using them. There were multiple layers of review between the first exploratory ChatGPT conversation in 2022 and a submitted doctoral thesis.

But suddenly 99 per cent did not feel sufficient.

So I checked.

I went back to the final submitted thesis and specifically searched for the questionable references from that first conversation. They had not propagated into it. There are incidental overlaps involving common author names and genuine papers, but I found no evidence that the fabricated or misattributed references generated during that first afternoon with ChatGPT had entered the final bibliography.

The relief was less interesting than the fact that I had felt compelled to check at all.

There is something here about the peculiar persuasiveness of generative AI. A fluent model does not necessarily persuade us by making outrageous claims. Often it does something subtler: it produces something close enough to the shape of knowledge that even years later, and even when I knew the verification processes I had followed, seeing those references again was sufficient to introduce a nagging doubt. I knew the provenance of my doctoral work. I knew the processes I had followed. I knew that I had repeatedly checked the references. Yet the output of a model from four years earlier was persuasive enough to make me reopen the finished thesis and establish the answer for myself.

That doubt is probably healthy. It is also an argument for some distinctly unfashionable tools and habits.

EndNote is not exciting. Manually checking whether a paper exists is not exciting. Opening the source and making sure it actually supports the sentence beside which it is cited is certainly not exciting. Neither are stable bibliographic records, versioned documents, supervisors’ comments, or the slow accumulation of a reference library over the course of a research project. Compared with asking a language model to generate a perfectly formatted bibliography in seconds, these methods can feel almost analogue.

Yet that friction is part of their value.

A reference manager is not there to generate the appearance of scholarship. Used properly, it maintains a relationship between the document being written and a source that the researcher has actually encountered and recorded. Checking a reference requires an action outside the generative system. Following a DOI, opening a paper, reading the relevant section and deciding whether it supports a claim creates a chain of provenance that fluency alone cannot provide.

There is also something reassuring about the fact that those older methods leave traces. Nearly four years after the first ChatGPT experiment, I did not have to ask another language model whether I had probably been careful. I could inspect the finished thesis. I could search its bibliography. I could return to the sources. The infrastructure of conventional academic practice provided an independent record against which the machine-generated material could be tested.

The lesson I take from that first conversation therefore isn’t that academics should avoid generative AI. My own subsequent history would make that a difficult argument to sustain. Generative AI eventually became an explicitly documented part of my doctoral research process, and I now use considerably more capable systems for forms of research, analysis, writing, coding and knowledge work that would have seemed extraordinary to me in December 2022.

The lesson is almost the opposite. The more capable and persuasive these systems become, the more valuable some of our older scholarly disciplines may become. Verification does not become obsolete because generation gets better, and provenance does not matter less because a model is usually right. A reference does not become trustworthy because the title, authors, journal and DOI all look exactly as we expect them to look. If anything, increasing fluency raises the premium on having methods that exist independently of the model and allow us to establish what is actually true.

This is also why the Obsidian project is becoming more interesting to me than I initially expected. I thought I was organising old conversations. Instead, I seem to be assembling evidence of two things developing simultaneously: the capabilities of the models and my own literacy in using them.

The first conversation provides a useful baseline. In December 2022, I was sufficiently intrigued by reports of ChatGPT hallucinating academic references that, within hours of first using it, I deliberately tested the claim. By the time I completed my doctorate, generative AI could be incorporated explicitly into a research methodology, but within a system that retained human critical review and conventional mechanisms of academic verification. By August 2026, I am using another generation of these systems to help preserve and interrogate the history of those interactions themselves.

That evolution is difficult to reconstruct from memory. The archive makes at least part of it observable. It allows me to ask when I first trusted a model with a particular kind of task, when I first tested something I did not trust, how my prompts became more sophisticated, which ideas recurred over time, and which failure modes disappeared or persisted as the technology changed. Eventually, it may be possible to trace ideas from an initial conversation through subsequent iterations and into projects, publications or decisions in the physical world.

There are limits to what such an archive can tell me. A conversation history is not a transcript of thought, and retrospective interpretation brings its own biases. The prompt itself, for example, proves that I asked ChatGPT to insert references; it does not prove why I asked. The context that I had heard murmurings about fabricated references and wanted to see the behaviour for myself comes from my recollection now. That distinction matters, particularly in a project concerned with provenance.

It is one reason I am preserving the original exports separately from the working Obsidian vault. I don’t want to turn several years of ChatGPT and Claude conversations into an infallible autobiography. I want to preserve the original record, distinguish it from subsequent interpretation, and then use the working archive to interrogate it.

There is a pleasing symmetry to that. The weakness I was testing in my first ChatGPT conversation was a failure of provenance: an answer could look convincing even when the evidential chain underneath it did not withstand scrutiny. Nearly four years later, provenance is precisely what makes the archive useful. I don’t have to rely solely on my recollection of how I used ChatGPT in 2022. I can return to the record, see what I actually asked, see what the model actually answered, and distinguish that contemporary evidence from the context I now remember surrounding it.

And the scale of that record is becoming significant. I am not casually browsing through some old chats. I am beginning a longitudinal analysis of a corpus of at least 1,669 LLM conversations spanning 1,324 days: 962 with ChatGPT and 707 with Claude. Even that corpus is incomplete. My conversations with Claude Code and Codex have not yet been exported and incorporated into the archive.

I began building the vault because I wanted a better way to preserve my conversations with artificial intelligence. I am starting to think it may contain something more valuable: a longitudinal record of how I learnt to work with it.

And, occasionally, a reminder of why I still check the references.

Something sparked your interest? Let's talk!