AI-assisted targeting: accountability and governance lessons for executives

Investigative reports: algorithms shaping life‑and‑death targeting decisions

Investigative reporting has documented cases where algorithmic tools reportedly flagged phones, homes, and individuals for rapid human review, a workflow critics say effectively lets an algorithm shape lethal decisions. The most detailed accounts come from +972 Magazine and were summarized by outlets including TIME and The Guardian. Those reports, and the legal and humanitarian responses they provoked, matter for anyone building, buying, or governing AI for high‑risk uses.

What the reporting says about Gaza

Journalists interviewed current and former Israeli intelligence officers who described two systems reported under the names Lavender and “Where’s Daddy?”:

  • Lavender, reported to identify phone‑use patterns that analysts treated as indicators of membership in Hamas or other armed groups.
  • Where’s Daddy?, reported to aggregate location and home‑presence signals, pointing analysts to suspected operatives at particular residences.

The reporting quotes intelligence officers saying analysts sometimes spent about “20 seconds” reviewing each AI‑flagged target before approving it. Sources told journalists that Lavender’s suggestions were understood to be incorrect in roughly 10% of cases. The Israel Defense Forces publicly told reporters, via a statement reproduced in TIME, that the tools “help intelligence analysts review and analyze existing information” and that they “do not constitute the sole basis for determining targets eligible to attack.”

“I would invest 20 seconds for each target at this stage, and do dozens of them every day. I had zero added-value as a human, apart from being a stamp of approval. It saved a lot of time.”, quoted intelligence officer (reported by +972 Magazine and summarized in TIME)

“I have much more trust in a statistical mechanism than a soldier who lost a friend two days ago. Everyone there, including me, lost people on October 7. The machine did it coldly. And that made it easier.”, quoted Israeli soldier (reported by +972 Magazine and summarized in TIME)

Why this matters for law, ethics, and risk

Humanitarian law still governs warfare. The IHL principle of proportionality forbids attacks expected to cause incidental civilian loss that would be excessive in relation to the concrete and direct military advantage anticipated. Legal critics including Kenneth Roth argue that the reported tolerance of high civilian harm in some targeting calculations violates that test. Roth summarized the situation bluntly: “An algorithm has been deciding who lives and who dies.” He further judged the conduct described in the reporting as “war crimes.” Those are forceful legal positions based on the reporting. They are not judicial findings.

The International Committee of the Red Cross has warned about human deference to machine outputs, a phenomenon commonly called “automation bias.” Automation bias looks like analysts favoring a machine’s probability score over ambiguous human evidence, or treating a model’s output as a default. That turns human reviewers into ceremonial approvers rather than genuine decision‑makers, and it undermines the evidentiary record needed to show that commanders made lawful proportionality assessments.

What meaningful human control should actually require

“Human in the loop” cannot be a ritual stamp. Make it operational and auditable:

  • Provenance and model context: reviewers must see timestamped sensor logs, data source identifiers, pre‑processing steps, model version, and a plain‑language description of the training data and known limitations.
  • Calibrated uncertainty metrics: display calibrated confidence scores, prediction thresholds, and expected error rates (precision/recall) so reviewers can judge reliability for the operational context.
  • Minimum review time and escalation: adopt tiered controls, e.g., a short triage window for low‑risk leads, but a minimum review period (and multi‑person escalation) for any target where the predicted civilian cost could be high.
  • Immutable audit trails: log every recommendation, who viewed it, what data they consulted, why they approved or rejected it, and preserve logs in tamper‑evident storage for independent review (multi‑year retention recommended for accountability).
  • Independent validation and red‑teaming: require third‑party operational testing and adversarial stress tests under realistic conditions before deployment, and periodic re‑validation post‑deployment.
  • Clear command accountability: define who signs strike authorizations. Empower legal counsel with stop‑work authority when proportionality or distinction cannot be confidently assessed.

Example workflow (realistic and enforceable): an AI flags a high‑risk target and attaches data provenance and a calibrated confidence score. A trained analyst then has a documented minimum of 10-30 minutes to review raw logs and alternative hypotheses for high‑risk cases, after which a supervising officer must sign a structured authorization form that is logged and archived. Shorter cycles are acceptable for low‑risk, time‑sensitive tactical decisions, but they must be defined in advance with explicit legal and operational criteria.

Practical reforms for governments, vendors, and executives

Prioritize governance over velocity. Five practical steps that materially reduce risk:

  1. Procure with transparency: require model cards and dataset datasheets from vendors, documented chain of custody for training data, and evidence of pre‑deployment operational testing.
  2. Design human‑centred interfaces: surface uncertainty, provenance, and alternative hypotheses. Avoid single‑click approval UIs that encourage rubber‑stamping.
  3. Institutionalize external validation: mandate independent labs, academic partners, or accredited national test centers to run adversarial and operational tests before fielding.
  4. Bake in auditability: require immutable logs, multi‑year retention, and accessible records to enable ex post legal and technical review.
  5. Empower legal and ethical stopgaps: give counsel and designated ethics officers demonstrable authority to pause operations pending rapid review of proportionality or distinction questions.

Case study and disputed claims, what is verified and what remains contested

The core operational details about Lavender and Where’s Daddy? and the 20‑second and ~10% claims come from investigative interviews published by +972 Magazine and summarized by TIME and The Guardian. The IDF has acknowledged developing analytic tools but maintains they do not autonomously pick targets and that analysts verify information. That creates a documented tension, with on‑the‑record anonymous sources describing perfunctory checks and significant error while official statements emphasize human verification.

Other assertions deserve caution. Kenneth Roth and others describe weakened Geneva diplomacy and attribute specific textual rollbacks to legal delegations sent by powerful states; those diplomatic attributions appear in commentary but require primary UN/CCW documents for independent verification. Similarly, claims about the operational use of “fully autonomous” weapons in other conflicts should be judged against detailed technical and OSINT evidence before being treated as established fact.

Diplomacy, norms, and the governance gap

Talks under the Convention on Certain Conventional Weapons in Geneva have debated lethal autonomous systems for years, and “meaningful human control” is a recurring demand from humanitarian actors. Yet most negotiating text and political debate still focus on fully autonomous weapons that fire without any human decision. That narrow focus leaves a large grey zone, because algorithmic systems that recommend, filter, or prioritize targets upstream of human action are not covered by a simple ban on “killer robots, ” and they can produce the same perilous outcomes in practice.

Beyond Gaza: trajectory and open questions

Reports from other theatres suggest militaries are increasingly using autonomy, sensors, and algorithmic decision aids. What remains open is how states will measure and publish error rates, how they will standardize meaningful human oversight, and who will be held accountable when algorithm‑assisted decisions produce unlawful civilian harm. Those are policy and procurement questions now, not hypotheticals for the future.

Accountability is not optional

Faster decision tempo should not mean thinner judgment. If institutions want the benefits of automation, they must invest in governance that preserves human responsibility: better UIs, clearer rules of engagement, robust audits, independent testing, and empowered legal oversight. Otherwise, organizations will discover that outsourcing judgment to opaque models brings moral and legal liabilities measured in human lives and institutional legitimacy.


  • Who reported on Lavender and “Where’s Daddy?” and the 20‑second review claim?

    Investigative reporting by +972 Magazine, reported and summarized in outlets including TIME and The Guardian, documented those programs and quoted intelligence officers describing brief human checks; the IDF issued a statement saying the tools are analytic aids and not sole determinants of targets.

  • Is the “10% error” figure an official IDF concession?

    The roughly 10% figure is reported by sources interviewed in the +972/TIME reporting. The IDF’s public statement emphasized analyst review and did not formally adopt or rebut that specific percentage in a policy document.

  • Are the actions described legally “war crimes”?

    Kenneth Roth and other legal commentators argue the reported practices could meet the tests for unlawful conduct under IHL, particularly on proportionality. That is a reasoned legal judgment based on the reporting, not a judicial determination recorded in the public record.

  • Have Geneva negotiations already banned these systems?

    Geneva/CCW talks have long addressed lethal autonomous weapon systems, but treaty language focused only on fully autonomous weapons would leave AI‑assisted targeting methods in a legal grey zone. Specific claims about text changes or the number of governments involved need primary UN/CCW documents for confirmation.

Key takeaways for executives and procurement leaders

  • Treat targeting AI as mission‑critical.

    Require transparency (model cards, datasheets), independent operational testing, and legal signoff before any deployment that could cause loss of life.

  • Make human control meaningful and auditable.

    Design interfaces and processes that surface provenance and uncertainty, enforce minimum review and escalation rules for high‑risk targets, and preserve tamper‑evident logs for review.

  • Require vendor accountability and continuous monitoring.

    Procure only from vendors willing to disclose provenance, accept third‑party validation, and support post‑deployment monitoring and rapid rollback if real‑world performance diverges from tested expectations.

  • Empower legal stopgaps.

    Give counsel and appointed ethics officers the authority to pause operations pending quick legal review when proportionality or distinction is uncertain.

  • Document governance publicly.

    Publishing governance practices preserves institutional legitimacy and helps demonstrate good faith if incidents occur.

Reality check: algorithmic assistance in targeting is not a futuristic issue. The reporting from Gaza shows how easily AI can migrate from analytic aid to decisive influence. Businesses in the AI supply chain, defense contractors, and procurement leaders must choose now: build governance that preserves human responsibility, or accept legal, moral, and reputational risks when machines shape who lives and who dies.