Large language models (LLMs) have notably advanced beyond interpretation and generation of text. Critical for our purposes here, they can also now "reason", interpret images, and responsively self-augment their knowledge base by retrieving new information (i.e., agentic). Given these new capabilities, in particular, as well as my own desire to understand how artificial intelligence may complement (or perhaps, eventually, supplant) my expertise, it seemed timely to benchmark how well they can currently make sense of the temporal attributes of open web content. To that end, I conducted experiments with some of my go-to example websites and pages:
- A SCOTUSBlog post
- The U.S. Supreme Court website
- A webpage on the U.S. Supreme Court website featuring a speech by Chief Justice Rehnquist
- The website for the 1996 film Space Jam
- A post from my own blog
I chose these specifically because their ostensible publication dates range from being obviously self-declared to entirely opaque or dependent on consulting third-party sources. A number of them additionally offer multiple, contradictory date signals, and so afford the opportunity to assess which signals are either legible or relatively more salient to LLMs.
I ran three types of tests:
- Zero-shot: asking when a particular website or webpage was first published
- Iterative: if the zero-shot prompt yielded a poor answer, I followed up by asking the model either how it came up with the date (i.e., if wildly off in its first response) or to venture a guess (i.e., if it demurred on being specific in its first response)
- Screenshot interpretation: asking the model to estimate the publication date of a webpage based on a screenshot
For this exercise, I used historical versions of the webpages in question to ensure better alignment between the actual (i.e., ground-truth) publication dates and the (closer-to-) contemporaneous website designs. As I've emphasized previously, the Internet Archive Wayback Machine (IAWM) was indispensable for establishing ground-truth publication or, at least, earliest-available dates.
I tested four freely available frontier models (in late May):
with memorization of previous chats disabled, to the extent possible, and previous chats manually deleted before each new chat session.
With that preamble, I'll cover the individual tests in sequence and then summarize my observations of the models' capabilities.
SCOTUSBlog post
The SCOTUSBlog post has the most straight-forward published date. The body text timestamp says that it was published 2 June 2021. This same date is corroborated by a dateCreated metadata attribute in the source code, a contemporaneous IAWM capture, and the Google bylineDate.
Not all models passed all tests, however; Gemini provided only a gross range for the screenshot interpretation test, and Grok was off by several weeks for both zero-shot and iterative prompting. All of the other model-tests yielded the correct answer.
Here is a summary view of the models' performance:
| Model | Version | Query type | Date in response |
|---|---|---|---|
| ChatGPT | GPT-4o | Zero-shot | 2021-06-02 |
| ChatGPT | GPT-4o | Screenshot interpretation | 2021-06-02 |
| Gemini | 2.5 Flash | Zero-shot | 2021-06-02 |
| Gemini | 2.5 Flash | Screenshot interpretation | Late 2020 / early 2021 |
| Grok | 3 | Zero-shot | 2021-06-29 |
| Grok | 3 | Iterative | 2021-06-29 |
| Grok | 3 | Screenshot interpretation | 2021-06-02 |
| Sonnet | 4 | Zero-shot | 2021-06-02 |
| Sonnet | 4 | Screenshot interpretation | 2021-06-02 |
U.S. Supreme Court website
The initial publication date of the U.S. Supreme Court website was progressively more ambiguous than that of the SCOTUSBlog post. From previous investigations
, I determined that the original domain of the U.S. Supreme Court website had been supremecourtus.gov. Performing a WHOIS lookup for this domain, it appears to have been registered on 1 December 1997, establishing the absolute earliest date that a website could have gone live. The earliest capture in IAWM dates to 20 May 2000. Examining the x-archive-orig-last-modified HTTP header pushed the date back to 27 April 2000, so I'd say that this is the earliest date that I can substantiate that the website had been published.
The models' guesses varied widely and were generally much less precise than for the SCOTUSBlog post. Sonnet performed the worst, on both the zero-shot and iterative tests; it suggested that the website had been published as early as the mid-1990s, implying that the U.S. Supreme Court was implausibly part of the early web's technological vanguard.
However, two different models (ChatGPT and Gemini) each for one of the tests (zero-shot and iterative, respectively) were able to come up with a plausible date — 17 April 2000 — that I hadn't otherwise discovered. It turns out that the go-live date for the U.S. Supreme Court website had been publicized and recorded contemporaneously on other websites (e.g., Researching Constitutional Law on the Internet: World Constitutions/Comparative Constitutional Law).
On some casual follow-up searching on DuckDuckGo and Google with reasonable search terms, I wasn't able to readily turn up this or related references. Of course, such statements alone would only provide circumstantial evidence that the website had been published earlier, but the date is plausible based on the earliest capture date in IAWM.