Ai2 Open Source AstaBrief 8B for rapid generation of scientific reports – Unite.AI
The Allen Institute for AI (Ai2) said on Oct. 2, 2026, it is open-sourcing AstaBrief 8B, a model that turns a search question and retrieved literature excerpts into a cited report, releasing model weights and training data as the system goes live as Fast mode in Asta, its agentic platform for scientific work.
Fast mode in generating Asta reports
AstaBrief is available in Asta’s Generate a Report feature as a Quick mode, which works alongside the existing Claude-based Thinking mode, according to Ai2’s announcement. Ai2 says its goal was to see whether a small, open model trained specifically for scientific report generation could match the report quality of proprietary models it had used, while reducing generation time and running costs. Ai2 reports that across the entire Asta pipeline, Fast mode averages 51.1 seconds per report compared to 178.5 seconds for Thinking mode, a difference it describes as approximately 3.5 times faster and as nearly an order of magnitude reduction in build time compared to tracked proprietary models.
The AstaBrief 8B model sheet states that the model is licensed under Apache 2.0, is based on Qwen3-8B, and is intended for research and educational use under the Ai2 Responsible Use Guidelines. Because the weights are open, Ai2 says institutions can run the model on their own hardware, even behind their own firewall, which it describes as necessary when research questions touch on sensitive or unpublished work. In addition to the weights, Ai2 has released a sample workflow in its GitHub repository ai2-scholarqa-lib that researchers can adapt to generate reports from their own PDFs.
Training and filtering data
Ai2 started from Qwen3-8B and created AstaBrief with supervised tuning followed by direct preference optimization (DPO). The announcement states that the team considered reinforcement learning-based training, which its previous DR Tulu work had shown can improve long-running report generation for open-weighted models, but chose the simpler recipe because RL training can be unstable and expensive, and the team wanted a setup that was cheaper and easier to debug and iterate on.
For the sake of speed, AstaBrief was trained to write the complete report in a single pass from the user’s query and retrieved snippets, skipping the snippet summarization and clustering steps used by the Claude-based thinking mode and skipping section-by-section writing. Ai2 says it found this was possible without sacrificing performance.
The training pipeline began with real user questions submitted via the system behind Ai2’s ScholarQA framework, which powers Asta’s report generation. The team filtered the logs for quality, relevance, and privacy, removing traffic from beta testers and bots, eliminating queries that were too short to be meaningful, and using an LLM-based pass to detect non-English queries, non-scientific requests, and prompts containing personal information. This left a pool of 90,000 search-focused queries.
For the supervised phase, comprehensive reporting targets were generated with the multi-stage ScholarQA pipeline supported by a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. Quality filtering left 47,000 usable examples. For DPO, the pairs came from a separate subset of queries not used when generating SFT data: a report from the ScholarQA pipeline, typically supported by Claude 3.5 or 3.7 Sonnet, versus a report generated by o3, o4-mini, DeepSeek-V3, or DeepSeek-R1. Two judging models, GPT-4.1 and DeepSeek-R1, chose a winner for each pair. Ai2 says the judges agreed with human preferences 95% of the time, and only the pairs that both judges agreed on were retained, producing around 6,000 final examples. According to the model sheet, the released model was initialized by the AstaBrief-8B-SFT checkpoint and fine-tuned on the AstaBrief_DPO_Mix dataset, with DPO training conducted in the Ai2 open-instruct framework on 8xH100 GPU.
Evaluation results
Ai2’s main development focus was SQABench-CS2, which the announcement describes as a set of 200 user-written computer science research questions. The team tracked four metrics: rubric score, which measures how much necessary content a report covered; response accuracy, which measures whether each paragraph is relevant to the question; citation accuracy, which measures whether each citation supports the claim attached to it; and citation recall, which measures whether a report’s claims are fully supported by the citations provided. Secondary evaluations used DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent arXiv papers, as well as pairwise comparisons with Claude-based pipeline reports.
The announcement reports that the team tested four filters based on summary training report statistics: output-to-input token ratio, citation relevance, citation density, and citation diversity. Ai2 says the biggest gains came from filtering reports with low citation density, while more aggressive filters, filter combinations, and learning rate controls brought no significant benefits. He describes the broader lesson as evidence that scientific specialization is not necessarily a matter of adding more scientific text to pre-training, and that the composition and quality of post-training data can materially change the performance of the resulting model.
Ai2 says that early SFT checkpoints improved overall content quality, but fell behind the Claude-based pipeline in terms of response accuracy and citation quality, and that the DPO phase brought AstaBrief into range of the Claude pipeline and DR Tulu on report generation. The model sheet reports that on the ScholarQA-CS2 test set, AstaBrief-8B scored an average of 87 across tracked metrics, compared to 83.7 for SFT checkpoint and 77.3 for base Qwen3-8B, with AstaBrief-8B scoring 90.2 on ingredient recall, 89 on answer precision, 90.5 on citation precision, and 78.2 on citation recall. The sheet also reports LLM-assessed win rates for AstaBrief-8B versus the Asta ScholarQA pipeline of 55% on the development split and 72% on the testing split, and lists DeepScholarBench scores of 53.50 for AstaBrief-8B, 60.25 for Asta ScholarQA, and 56.26 for DR-Tulu-8B.
In a separate human study described in the announcement, three scientific researchers each contributed four to five questions out of a set of 14 questions and ranked the reports from the three systems based on overall preference, completeness, relevance, organization and accuracy of citations, with links allowed. Ai2 reports that DR Tulu won in terms of overall preference, while two of the three researchers preferred AstaBrief over the other systems in terms of citation accuracy.
Ai2 cautions that the majority of the training and evaluation was completed in 2025, that the proprietary models used to generate training data and serve as points of comparison reflect the frontier at that time, and that it has not re-run the full evaluation against the current frontier models.
Initial usage and next steps indicated
Ai2 reports that among 374 Asta users who tried Quick mode, 29.1% used it for two or more days, users generated an average of 3.67 report threads, 23% never returned to Thinking mode for future threads, and another 18% alternated modes depending on their goals, using Quick mode for about 40% of threads. Positive feedback was 84.2% for Fast mode compared to 85.2% for Thinking mode, which Ai2 characterizes as a similar rate while noting that feedback is generally too poor to support strong conclusions.
Ai2 says it is exploring more detailed preference learning, stronger RAG-plus-RL approaches, multi-turn and multi-tool capabilities, additional scientific data sources, and query decomposition, along with evaluations that test whether a model preserves the evidentiary scope of its sources rather than expanding what underlying studies have established. The announcement places the work within Ai2’s broader science modeling efforts, including NSF OMAI, a national U.S. initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, and describes AstaBrief as one experiment in a longer line of work stretching from ScholarQA and DR Tulu to future releases of Olmo.



Post Comment