The April 2023 technical report Scaling Transformer to 1M tokens and beyond with RMT explored how recurrent memory could extend the sequence length processed by a BERT-based model. Its widely discussed result was approximately two million tokens on synthetic memory tasks. Understanding the task is essential to interpreting that number.
The work is relevant to research on long documents, but processing a long test sequence does not establish general comprehension of an equally long document. It also does not demonstrate the ability to generate coherent novels or retain a person’s complete history.
What’s BERT?
BERT is an encoder model introduced by Google researchers for learning language representations. It can be adapted to tasks such as classification and extractive question answering. The BERT backbone in this experiment should not be treated as an autoregressive text generator.
BERT uses subword tokens. A token can represent a word, part of a word or another text unit, so token counts depend on the tokenizer and the material being processed.
Tokenization and long inputs
Comparisons with GPT-4 and CoLT5 in the original report belong to the April 2023 research context. They are not a current inventory of model context limits. Token counts also cannot be converted into a fixed number of pages or conversation hours without specifying the text and tokenization method.
A long input capacity describes how much material a system can process. Evaluation must separately establish what information it retrieves, how it reasons about that information and which errors it makes.
The 2304.11062 research paper
The paper record includes later revisions. The 19 April 2023 version evaluated a BERT-based Recurrent Memory Transformer on constructed tasks involving fact memorization, fact detection and simple reasoning. Answers were represented as six classification options, with irrelevant text inserted between relevant facts and the question.
The model processed segments sequentially and passed memory vectors from one segment to the next. This limits the attention computation within each segment while allowing selected information to persist. The report describes extrapolation to roughly two million tokens; its headline length and the effective content count differ because segment space is reserved for special and memory tokens.
These experiments support a specific claim about memory retrieval and length extrapolation under the tested conditions. They do not show that all information in a large technical archive will be retained, or that results transfer unchanged to free-form generation. The original report identified broader language tasks as further work.
Conclusion
For an energy-industry application, the next evaluation would need representative reports, maintenance records and questions whose answers can be checked. Missing information, conflicting records and retrieval errors would matter as much as maximum input length. The result offers a research direction; deployment benefits still require application-specific evidence.
Implementation material linked by the authors is available in the research repository.




