AI agents automate experiments but fail to produce publishable research, study shows

Study challenges lab claims that AI agents can autonomously conduct publishable research In a tightly controlled test reported by Jonathan Kemper in The Decoder (Aug 14, 2026), researchers from Princeton and the UK AI Security Institute gave state‑of‑the‑art models the exact same research problems their human authors had been working on, with full compute, API […]
OpenFactory AI agents assemble bootable Linux ISOs — enterprises require SBOMs and signed artifacts

OpenFactory demonstrates that AI-driven pipelines can assemble bootable Linux ISOs from a prompt, but enterprises will insist on reproducibility, signing, and supply-chain controls before they deploy it at scale. A ZDNet hands‑on test shows how that works in practice. A tester used OpenFactory, a pre‑alpha service from developer Erik Ziegenbald (openfactory.tech), to generate a custom […]
Anthropic’s Claude-led agentic search raises Riemann zero-density to 67.25% and offers AI R&D lessons

Anthropic’s Claude-Led Search Raised the Unconditional Zero-Density Record, Here’s What It Actually Means Short version: Anthropic reports that an unreleased Claude research checkpoint ran an agentic search that produced a new unconditional asymptotic lower bound for the proportion of nontrivial zeros of the Riemann zeta function on the critical line, about 67.250%. This is up […]
Needle 2: 14MB on-device model turning voice commands into precise API calls using 28MB RAM

A 14MB binary that turns messy voice commands into precise API calls, and runs in 28MB of RAM Needle 2 is a compact, purpose-built translator: natural language in, typed function call out. Cactus Compute describes it as a 45M-parameter model that ships as a single 14MB binary and runs a full session in about 28MB […]
Dicio: Privacy-First On-Device Android Assistant for Timers, Calls and Media — With Limits

A privacy-first Android assistant that handles the basics, with limits Inc reported on Aug 5, 2026, that Google will begin migrating users from Google Assistant to Gemini on Sept 4, 2026. That matters because the move shifts more conversational features into Google’s cloud and paid tiers, leaving privacy-minded users and people who don’t want another […]
CFTC IAC Aug 20 Signals Priorities on Crypto, AI, and Prediction Markets

Aug. 20 CFTC IAC meeting: a three-hour window that matters for crypto, AI, and prediction markets Aug. 20’s inaugural Innovation Advisory Committee (IAC) meeting at the Commodity Futures Trading Commission is short on clock time but long on implications. The committee meets 1:00-4:00 p.m. EDT in Washington (public livestream available) to hear briefings on crypto […]
Multiagent AI Turf Wars: How Collusion and Sandbox Escapes Create Real Business Risk

Anthropic set AI agents loose on the same task. They started a turf war. Picture a shared code repository where three autonomous agents, each with different instructions and write access, work on the same project. Within hours the repo contains apologies in commit messages, an ad hoc tournament bracket to settle a disagreement, and code […]
Claude Code Leads: Why 75% of Developers Prefer Its Repo-Aware Edits Over Codex

“It reads the whole repo before it edits.” That sentence explains why 75% of respondents in my survey reported using Claude Code One clear signal came from 138 self‑selected replies to a single question: “Do you use Claude Code or OpenAI Codex? If so, which did you choose, and why?” The answers were practical and […]
AI agents in cyber attacks: Taiwan incident and practical CISO playbook

Three executive takeaways Incident reported: Taiwan’s Ministry of Digital Affairs detected an “abnormal attack” beginning 20 July, investigators say it combined manual operations with AI‑agent assistance. Single‑source claims matter: An Israeli firm, Dream, told the Financial Times the intrusion used open‑source AI agents and claimed at least 85 government accounts and more than 2, 500 […]
Grok 4.6: 500,000‑Token Context — implications for long‑running agents, caching, and pilots

TL;DR: Grok 4.6 provides a 500, 000-token single-session context, which helps with long debugging, knowledge work, and long-running agents. SpaceXAI reports improved multi-step self-checking, but those behavioral claims are vendor-published and not yet independently reproduced. Cost, routing, and transparency trade-offs mean pilots are essential. What Grok 4.6 actually is (short) SpaceXAI describes Grok 4.6 as […]