Ambient Agents on AWS: From S3 Events to Scalable, Human-in-the-Loop Automation

A claims PDF lands in S3. The agent extracts policy numbers, spots a low‑risk match, auto‑approves it, and only surfaces the handful of high‑risk claims to a human reviewer. That single example shows what an ambient agent does. It watches event streams (S3 uploads, scheduled jobs, webhooks, DB changes), turns matching events into durable work […]

Chains-of-thought leakage on cloud-hosted models: require per-session crypto and adversarial tests

TL;DR Researchers showed attackers could replay hosted session artifacts, called “reasoning blobs, ” to recover models’ intermediate chains-of-thought on some cloud-hosted endpoints. OpenAI says it detected and shut down a coordinated “adversarial distillation” campaign (banning more than 15, 000 accounts and applying fixes), but the researchers reported the same extraction techniques still worked on certain […]

pplx-embed-v2-context-9b-preview: Perplexity’s evidence-aware embeddings for provenance in RAG

Perplexity Releases pplx-embed-v2-context-9b-preview: a contextual embedding that retrieves answers and their evidence Perplexity Research (with turbopuffer) published pplx-embed-v2-context-9b-preview, an embedding model trained to surface both relevant answer chunks and the supporting sentences inside documents. The weights are available on Hugging Face (perplexity-ai/pplx-embed-v2-context-9b-preview) under the MIT license and the release is a self-hosted preview. Perplexity notes […]

Gemini 4 Argon: Google narrows the frontier gap — pilot and measure token usage first

Gemini 4 Argon, Google narrows the frontier gap but don’t swap vendors yet In short: Argon meaningfully closes capability gaps with top OpenAI and Anthropic models, and its 1, 000, 000-token context/output limits plus Long-Decode Continuation make new long-form workflows possible. But Argon’s higher average output-token use and several unresolved runtime and billing details mean […]

Agentic Retrieval with Amazon Bedrock Knowledge Bases for Cited Insurance Claim Answers

You shouldn’t have to hunt through a dozen PDFs, emails, and scanned notes to answer “Has the estimate for claim CLM‑100482 been approved?” That demo question teaches a simple architectural lesson: use a managed Retrieval‑Augmented Generation (RAG) pattern to ground answers in source documents, show citations for traceability, and keep humans in the loop for […]

AI-augmented workers: Build a stitched mega-platform to verify, upskill and place talent

Five great solutions. One missing link. The HP Future of Work Accelerator Pitch Fest in New York (September 30, 2026) put five sharp ideas on stage: Skillionaire Games, SkillUp Coalition, PHIND, Extern and the Millennium Campus Network. I judged the event alongside Antara Lahiri, Catherina Gioino, and Michele Malejki. It was produced by Conspiracy of […]

Zhipu GLM‑5.3 nears Claude Mythos on exploit benchmarks; Anthropic’s claims need verification

Anthropic reports Zhipu’s open‑weight GLM‑5.3 approaches Claude Mythos Preview on exploit benchmarks, claims await independent verification Anthropic’s Frontier Red Team published an analysis (Sep 30, 2026) reporting that Zhipu AI’s open‑weight GLM‑5.3 can autonomously produce working exploits at rates close to Anthropic’s restricted Claude Mythos Preview on public benchmarks. The headline numbers: GLM‑5.3 “built a […]

AI Voice for Email Approvals: Use as Narrator, Not Sign-Off

I Let an AI Voice Approve My Emails. 3 Were Wrong. Siraj Raval ran a tight, uncomfortable experiment: he layered a text-to-speech voice over an existing AI agent that manages his inbox, closed the visual log, and approved every action by listening only. After the session he opened the log and announced the count: “I […]