OCTOBER 11, 2026
Subscribe
Global Press Media · World Report
Technology

Anthropic Removes Public Model Evaluation Datasets, Calling It a Temporary Precaution

Anthropic Removes Public Model Evaluation Datasets, Calling It a Temporary Precaution

On Tuesday, Anthropic—the company that develops the Claude line of language models—said it will pull its publicly available model evaluation datasets from the web, characterizing the action as a temporary precaution rather than an irreversible closure.

These evaluation collections, containing benchmark results, prompts and performance logs, have been reachable to scholars and developers for a few months, offering an uncommon view into Anthropic's model abilities and safety traits. The open release was intended to promote transparency and allow external parties to verify the firm’s assertions.

Anthropic explained in a short statement that the pull‑back stems from rising worries about possible abuse of the evaluation material and the competitive dynamics of large‑language‑model progress. The company warned that leaving the data accessible might unintentionally assist parties attempting to reverse‑engineer or misuse model behavior, particularly as the sector races toward ever stronger systems.

The move has ignited debate among AI researchers regarding the trade‑off between openness and security. Academics depend on common benchmarks to gauge advancement, replicate findings, and spot safety vulnerabilities. With Anthropic’s datasets now offline for the time being, certain scholars caution that the action could hamper collaborative work and restrict independent evaluation of the firm’s safety assertions.

Anthropic gave no date for when the datasets might return, yet stressed that the pull‑back is temporary. It said it is looking into other methods of publishing evaluation outcomes that reduce risk while maintaining scientific standards. Analysts observe that comparable limits have emerged at other major AI companies, pointing to a wider shift toward more cautious sharing of model performance data as the field balances innovation with accountability.

Source: Gizmodo
Editorial Desk — Editorial desk.

Comments (0)

Be the first to comment.

Join the discussion

Protected by reCAPTCHA v3

Related