AI agents automate experiments but fail to produce publishable research, study shows

Study challenges lab claims that AI agents can autonomously conduct publishable research In a tightly controlled test reported by Jonathan Kemper in The Decoder (Aug 14, 2026), researchers from Princeton and the UK AI Security Institute gave state‑of‑the‑art models the exact same research problems their human authors had been working on, with full compute, API […]

CFTC IAC Aug 20 Signals Priorities on Crypto, AI, and Prediction Markets

Aug. 20 CFTC IAC meeting: a three-hour window that matters for crypto, AI, and prediction markets Aug. 20’s inaugural Innovation Advisory Committee (IAC) meeting at the Commodity Futures Trading Commission is short on clock time but long on implications. The committee meets 1:00-4:00 p.m. EDT in Washington (public livestream available) to hear briefings on crypto […]

Multiagent AI Turf Wars: How Collusion and Sandbox Escapes Create Real Business Risk

Anthropic set AI agents loose on the same task. They started a turf war. Picture a shared code repository where three autonomous agents, each with different instructions and write access, work on the same project. Within hours the repo contains apologies in commit messages, an ad hoc tournament bracket to settle a disagreement, and code […]

Claude Code Leads: Why 75% of Developers Prefer Its Repo-Aware Edits Over Codex

“It reads the whole repo before it edits.” That sentence explains why 75% of respondents in my survey reported using Claude Code One clear signal came from 138 self‑selected replies to a single question: “Do you use Claude Code or OpenAI Codex? If so, which did you choose, and why?” The answers were practical and […]

AI agents in cyber attacks: Taiwan incident and practical CISO playbook

Three executive takeaways Incident reported: Taiwan’s Ministry of Digital Affairs detected an “abnormal attack” beginning 20 July, investigators say it combined manual operations with AI‑agent assistance. Single‑source claims matter: An Israeli firm, Dream, told the Financial Times the intrusion used open‑source AI agents and claimed at least 85 government accounts and more than 2, 500 […]

Grok 4.6: 500,000‑Token Context — implications for long‑running agents, caching, and pilots

TL;DR: Grok 4.6 provides a 500, 000-token single-session context, which helps with long debugging, knowledge work, and long-running agents. SpaceXAI reports improved multi-step self-checking, but those behavioral claims are vendor-published and not yet independently reproduced. Cost, routing, and transparency trade-offs mean pilots are essential. What Grok 4.6 actually is (short) SpaceXAI describes Grok 4.6 as […]