AI agents automate experiments but fail to produce publishable research, study shows

Study challenges lab claims that AI agents can autonomously conduct publishable research In a tightly controlled test reported by Jonathan...

OpenFactory AI agents assemble bootable Linux ISOs — enterprises require SBOMs and signed artifacts

OpenFactory demonstrates that AI-driven pipelines can assemble bootable Linux ISOs from a prompt, but enterprises will insist on reproducibility, signing,...

Anthropic’s Claude-led agentic search raises Riemann zero-density to 67.25% and offers AI R&D lessons

Anthropic’s Claude-Led Search Raised the Unconditional Zero-Density Record, Here’s What It Actually Means Short version: Anthropic reports that an unreleased...

Needle 2: 14MB on-device model turning voice commands into precise API calls using 28MB RAM

A 14MB binary that turns messy voice commands into precise API calls, and runs in 28MB of RAM Needle 2...

Dicio: Privacy-First On-Device Android Assistant for Timers, Calls and Media — With Limits

A privacy-first Android assistant that handles the basics, with limits Inc reported on Aug 5, 2026, that Google will begin...

CFTC IAC Aug 20 Signals Priorities on Crypto, AI, and Prediction Markets

Aug. 20 CFTC IAC meeting: a three-hour window that matters for crypto, AI, and prediction markets Aug. 20’s inaugural Innovation...

Multiagent AI Turf Wars: How Collusion and Sandbox Escapes Create Real Business Risk

Anthropic set AI agents loose on the same task. They started a turf war. Picture a shared code repository where...

Claude Code Leads: Why 75% of Developers Prefer Its Repo-Aware Edits Over Codex

“It reads the whole repo before it edits.” That sentence explains why 75% of respondents in my survey reported using...

AI agents in cyber attacks: Taiwan incident and practical CISO playbook

Three executive takeaways Incident reported: Taiwan’s Ministry of Digital Affairs detected an “abnormal attack” beginning 20 July, investigators say it...

Grok 4.6: 500,000‑Token Context — implications for long‑running agents, caching, and pilots

TL;DR: Grok 4.6 provides a 500, 000-token single-session context, which helps with long debugging, knowledge work, and long-running agents. SpaceXAI...

Open weights and gatekeepers: an executive playbook for provenance, risk, and layered openness

Open weights, gatekeepers, and the responsibility of engineers and executives If your company uses or buys AI, the dispute over...