Secret dates in system prompts undermine language model evaluation – Unite.AI
Without the ability to compare Large Language Models (LLM), it is difficult for consumers and businesses to understand how much progress a model has made compared to recent releases and how it compares to its competitors:
The influential LLM ranking on arena.ai. Source
Since LLMs are not deterministic (i.e. they will Not always produce coherent outputs starting from the same inputs), evaluating them is complicated. Even when researchers use identical prompts, model settings, and benchmark datasets, seemingly minor differences in the execution environment can produce different responses and significantly alter performance scores.
Factors such as hardware configuration, numerical precision, inference batch size, and even the ordering of multiple-choice answers can affect the results. This can make it difficult to see whether a reported improvement reflects real progress or simply a semi-random variation in the conditions under which the model was tested.
To some extent, some of these variables can be taken into account, or at least established what margin of error they generate, so that the comparative evaluation becomes meaningful. It’s important to try, since a lot of money and a lot of reputation depend on the ability to compare AI systems of this kind with some degree of accuracy.
However, according to new research, a particular variable can not only be destructive to benchmarks, but is also very difficult to eliminate from the stimuli that define them: today’s date.
Times change
The new paper*, an academic collaboration between Germany, Mexico and the United States, states that the fact that the current date is automatically and secretly included in the system of all borders and in many open-weight model implementations means that reproducibility may be almost impossible:
“We identify a critical, often overlooked, source of non-determinism in LLM assessment: the hidden insertion of the current date into system prompts.
“Across 9 models, 6 datasets, and four tasks, this dynamic metadata alters model performance and shuffles leaderboard rankings, overcoming variance introduced by other system-level factors such as batch size or numerical precision.”
The researchers also note that the identified effect is largest for tasks requiring generated responses, with performance varying by up to 6% on multiple-choice questions, 14% on mathematical reasoning, and 7% on code generation – and with machine translation scores varying significantly.
Providing sample answers and encouraging step-by-step reasoning did not solve the problem – in fact, this one did worse:
Accuracy fluctuations between different dates in 2024 for Llama 3.1 (8B) on the MMLU benchmark. Step-by-step reasoning (red) produces substantially greater variability than direct responses (blue), demonstrating how the suggested chain of thoughts amplifies date sensitivity. Source
The effect was also identified in proprietary models, with GPT-5.1 showing accuracy fluctuations of up to 4% on three multiple-choice benchmarks during a week of testing in December 2025. The researchers used a blank system prompt, disabled reasoning, and set randomness to zero, but the date was still filled in automatically by the vendor:
GPT-5.1’s accuracy fluctuated for seven consecutive days in December 2025, despite identical user instructions and model settings. The largest change occurred on GPQA (orange), reaching four percentage points between the best and worst days. The dotted line represents the average accuracy of each benchmark over the course of the week.
The researchers recommend removing the date from system prompts where possible, or correcting and documenting it, to ensure fair comparisons. However, they note that the date-centric influence may be more deeply rooted:
“One possible reason for the date sensitivity is that the system prompt may be corrected during supervised fine-tuning (SFT) and human feedback reinforcement learning (RLHF), making the model fragile to any slight modification.”
To investigate the problem, the researchers tested six different wordings for the system prompt and found that these changes affected accuracy as much as changing the date (0.78%). Therefore the date seems to have the same influence as deliberate timely engineering, except that it changes automatically, without the user even knowing it
Other approaches have been attempted to mitigate the “current date” problem, including few-shot learning (where the model obtained five sample responses before responding). This helped a little, lowering the average change from 2.52% to 2.27%, but it didn’t solve the problem.
Changes to GPU hardware, batch size, numerical precision, response order, and system prompt wording were also tested. The largest effects were observed with response order and timely wording, which most closely matched the impact of the date change.
Last days
It is reasonable to expect an LLM/VLM to include the date from the start of a chat; however, there appears to be no explicit reason to put it in the system prompt (the invisible rubric that imposes guardrails and conditions LLM behaviors in ways inaccessible to the user) when LLM could routinely make a sub-Kb RAG call to acquire the last date, as a minor cleanup routine before engaging with the user.
No doubt there are other possibilities for resolving the issue; however, since entering the current date in the system prompt has not been seen as a problem thus far, it is likely that there has been little or no investigation into this matter.
Online dating
For testing, identical prompts were used for each model, only the date in the system prompt was changed. Every day of 2024 was tested, from January 1st to December 31st, keeping all other settings unchanged.
Six benchmarks were used: MMLU; GPQA; and ARC-Challenge, for multiple choice questions (scored based on the probability of the response token). GSM8K for step-by-step mathematics (final controlled answer); HumanEval for generating Python code (unit tested); and WMT for English-German; From English to Finnish; and translation from English to Czech (the entire output will be evaluated). Time-dependent questions were excluded.
Nine models were tested: Llama 3.1 Instruct (8B and 70B); Gemma 3 Instruct (4B and 27B); Qwen3 (4B); Qwen3-Forward (80B); Phi-4 (14B); and GPT-OSS (20B and 120B).
Accuracy (the percentage of questions answered correctly) was used to score the multiple-choice and math tests, while expected calibration error (ECE) measured how well the models’ confidence matched their actual performance.
The code was checked using pass@1 (the percentage of generated code solutions that pass all tests on the first try); translations were evaluated using BLEU and chrF.
The results confirmed the researchers’ hypothesis: simply changing the date in the system prompt changed the precision with which the models answered the same questions.
Test results show the performance of five leading models on MMLU when the system-required date was changed during 2024. Accuracy is shown at the top, with expected calibration error (ECE) at the bottom. All five models showed fluctuations in both measures, despite no other changes to the testing conditions. Lower ECE scores indicate better alignment between confidence and accuracy.
Together with other results shown earlier in the article, this is potentially bad news for LLM benchmarking, because a model could score better or worse depending on the day it was tested, possibly changing its ranking, without any actual improvement or decline in its capabilities.
Conclusion
This issue highlights the gap between the deterministic computing systems we were used to before around 2023, and the very different nature of diffusion-based and similar AI systems that have evolved and continue to evolve since then.
The date resolution was essentially resolved on January 1, 1970, but has apparently come back to haunt the computing world in the form of cutoff dates, among other timing concerns.
* Titled “Dating the Model: Hidden Dates in System Prompts Affect LLM Grading”
First published Friday 9 October 2026



Post Comment