On September 10, 2026, Anthropic's Frontier Red Team and Threat Intelligence team published evaluations of two capabilities that sit outside the usual cyber and biology risk lists: intelligence targeting (identifying and locating people) and conventional weapons tasks built around drones. They tested Claude Mythos Preview, Mythos 5, Opus 5 and Sonnet 5 alongside the open-weights models Kimi K3 and GLM 5.2, with models working alone in a sandbox without internet access.
On targeting, the results are strong. In photo geolocation on 6,000 Flickr images from the YFCC100M dataset, Mythos Preview had a median error of 37.0 km (23.7 percent within 1 km), Mythos 5 47.2 km and Opus 5 181 km, against 151 km for players in the GeoGuessr Champion Division. Placing 1,697 anonymized Twitter users from their text alone, the best models reached median errors around 20 km, and 135 users (8 percent) were reliably placed within 1 km by at least one model. On 200 identity-linkage tasks set in fictional protest scenarios, models processed a median 37,000-word sample in about 11 minutes, work Anthropic estimates would take a human analyst about 2.5 hours just to read.
The drone results are much weaker. In simulated terminal guidance, Opus 5 hit a parked, high-visibility vehicle 80 percent of the time but fell to 47 percent against a vehicle moving at road speed, and camouflaged or evading targets were essentially unsolved; across nine settings it hit the target on 20 percent of 540 launches. Payload drops on a static target landed almost every time, but against a moving target in wind Opus 5 succeeded on 28 percent of sorties, and under subtle GPS spoofing all models performed poorly. Anthropic's conclusion is that frontier models already clear the capability floor for low-resource groups, and that the open-weights ecosystem is close enough behind "that the gap should not be mistaken for safety."
Anthropic lists the limits itself: the social media data is synthetic and not fully realistic, the drone work is simulation-only with camera rendering far simpler than reality, and the evaluations speak mainly to enabling low-resource actors rather than measuring real-world uplift. The paper also argues that the most dangerous system may be one trained deliberately for surveillance or weapons work, which no evaluation of general-purpose models can measure.