AI benchmark finds GPT-5.6 Sol completes Fast16 malware probe
SentinelLabs’ Fast16 benchmark tested frontier AI models; only OpenAI’s GPT-5.6 Sol finished all eight investigative stages, while others stalled or stopped early.
SentinelLabs, the research arm of cybersecurity firm SentinelOne, built a long-horizon reverse-engineering benchmark using its Fast16 malware investigation. The Fast16 analysis was published in April and served as a test case to see if large AI models can carry a full, multi-stage malware inquiry when new evidence contradicts earlier conclusions.
Fast16 is a 2005 Windows malware that targets LS-DYNA, an engineering simulation package. SentinelLabs links Fast16 to efforts to disrupt Iran’s nuclear-weapon development and describes it as resembling earlier sabotage malware. The benchmark tracks whether a model can sustain an investigation through eight escalating stages rather than completing isolated tasks.
SentinelLabs tested four frontier model families: OpenAI’s GPT-5.5 and GPT-5.6 Sol, Z.ai’s GLM-5.2, and Anthropic’s Opus 4.7 and 4.8. GPT-5.6 Sol completed all eight stages in three separate runs using different reasoning-effort settings. GPT-5.5 did not advance past the first stage. GLM-5.2 stalled partway through the process, and the Opus models often reported tasks as finished before core defects were resolved.
The report frames the gap as a failure of “project-scale recovery.” SentinelLabs defines that term as the ability to withdraw a disproven conclusion, identify and update downstream results that depended on it, fix the root cause, and carry those corrections through the remainder of an investigation. Researchers found many models could perform useful local analyses but struggled to revise earlier reasoning and update downstream work when new evidence overturned prior assumptions.
SentinelLabs noted even the strongest runs produced errors. The report observes GPT-5.6 Sol made semantic mistakes, accepted weak quality checks and claimed readiness prematurely. “Senior reverse engineers remain essential,” the report adds. It recommends the “best current use” as supervised investigative agency, with human analysts defining objectives, exposing blind spots and retaining final publication authority.
SentinelLabs describes the benchmark as the first long-horizon reverse-engineering assessment for frontier AI systems. The firm released the results to document model behavior across a full investigative workflow and to inform how organizations combine AI assistance with human oversight.








