Shortlisting ML vendors in 2026: platforms, services and contract SLAs for production

How to shortlist ML vendors in 2026, platforms, services and the contract items that separate pilots from products

Vendor roundups are plentiful. Procurement-ready guidance is not. If your goal is models that survive contact with the real world, including retraining cadence, monitoring, on-call ownership and integration into a system of record, you need a different checklist than marketers publish. The list below adapts SmartDataCollective’s procurement-focused framework and adds concrete, measurable asks you can include in an RFP.

Basis for this assessment: documented capabilities, published service descriptions, stated pricing models, deployment options and category positioning. Not hands-on trials, not procurement records, and not vendor briefings. We have not verified headcounts, hourly rates or client relationships, so this page prints none of them, unlike most of the pages competing for the same search. Where a specific claim couldn’t be checked, the sentence was written without it.

Quick shortlist: platforms vs services (and why the split matters)

These are different purchases. Platforms sell infrastructure, managed models or tooling (model registries, feature stores, serving). Services firms sell people and processes, including design, integration, change management and operations. Shortlist one of each matched to the problem you need fixed, not to the brand you like.

Platforms (who to consider and why)

  1. Databricks: Noted for a unified lakehouse and close MLflow/feature-store integrations. Good for engineering-led teams that want a single workspace for data engineering, training and serving.
  2. Amazon Web Services (SageMaker, Bedrock): Strong when your estate already runs on AWS. Broad tooling for training, hosting and managed foundation models tied into the rest of your cloud services.
  3. Google Cloud (Vertex AI, BigQuery): Useful where analytics are warehouse-centric. Vertex AI and BigQuery adjacency make it straightforward to combine managed models with custom training on analytics data.
  4. Microsoft Azure: The obvious choice for organizations already committed to Azure and Microsoft enterprise stacks. Strong identity, governance and hybrid deployment options for regulated enterprises.
  5. DataRobot: Oriented to business-analyst-owned forecasting and tabular prediction workflows. Packaged tooling for model lifecycle and explainability is geared toward analyst teams.
  6. Hugging Face: Central repository for open-weight models, datasets and evaluation results. Useful when you need full model ownership, fine-tuning or want to avoid per-token hosting costs.
  7. OpenAI: Fastest route to ship language features via hosted APIs. Best where time-to-market beats model ownership and you accept per-token economics and managed updates.

Services firms (who to consider and why)

  1. Accenture: Deep experience in multi-country rollouts and the change management needed to drive adoption across global operations.
  2. Deloitte: Strong for regulated environments. Model risk management, audit readiness and regulator engagement are core capabilities.
  3. QuantumBlack (McKinsey): Suited to upstream prioritization and deciding what’s worth building before committing engineering effort.
  4. Infosys: Good for sustained delivery capacity against legacy systems and long-running operational maintenance.
  5. Fractal Analytics: Decision-science depth for CPG, retail and financial services. Focused on improving commercial decisions, not just model accuracy.
  6. Quantiphi: Cloud-native ML engineering with strong hyperscaler alignment. Well-suited for vision workloads and production engineering.
  7. Tredence: Practical for retail and supply-chain models that must survive daily operations and integrate with planning systems.
  8. Faculty: Emphasizes explainability and defensibility. Useful where public or regulatory scrutiny demands transparent decisioning.
  9. ScienceSoft: Combines scoped ML work with the software delivery and support needed to embed models into products.
  10. InData Labs: Positioned for mid-market teams taking on a first ML project and needing rapid, pragmatic delivery.
  11. Addepto: Focused on ML built on top of warehouses and BI stacks. Useful when analytics teams own the data and want to convert insights into models.
  12. Tooploox: Research-leaning firm strong on computer vision and product-shaped R&D. Good where prototype-to-product transitions matter.

Rapid vendor sieve, a compact checklist to use in vendor interviews

Use these questions to separate vendors who can ship production systems from vendors who hand over pilots.

  • Does the firm’s public record show production systems, or pilots? Look for named, contactable production references and artifacts (dashboards, model registry entries, live endpoints).
  • Is ML the core practice or a line item? A firm that treats ML as a service line will typically staff differently and price differently than one where ML is mission central.
  • Is the positioning specific enough to be falsifiable? If a vendor’s claims are vague marketing, you won’t be able to test acceptance criteria.
  • Would you know who to call at 3am in month nine? Get named escalation contacts, expected on-call response times and a documented incident path.

“The gap between a firm that ships and a firm that hands over is the whole ballgame, and it usually shows up in how they describe support, not how they describe modelling.”, SmartDataCollective

MLOps and operational deliverables you must insist on (make them measurable)

Don’t accept a ZIP file and a slide deck. Require specific artifacts and SLAs:

  • Feature management: a feature store (e.g., Feast or vendor equivalent) and proof that training and inference use identical features.
  • Model registry & serving: registry entries, versioned weights and a hosted or exportable serving endpoint with an availability SLO. Example: 99.9% inference availability.
  • Monitoring: data and concept-drift alerts (daily checks) with thresholds that trigger retraining. Require sample alert definitions and an on-call runbook.
  • Retraining cadence: either automated retraining triggered by drift thresholds or a scheduled cadence (e.g., every 4-12 weeks). Specify acceptable model-performance degradation before retrain.
  • Rollback & incident response: a rollback playbook with RTO/RPO targets and an incident response SLO. Example: under 4-hour initial response for Sev-1 incidents.
  • Ownership & exportability: clarity on who owns fine-tuned weights, derived datasets and code. Include a migration test that exports artifacts within 30 days on termination.

Procurement caution: supplier ownership and conflict risk

Vendors change. Ownership or customer shifts can create procurement risk even if technical capability remains. There have been public reports of major ownership and customer relationship shifts at prominent ML vendors in 2025. Buyers should verify current supplier relationships, data-usage terms and commercial commitments before signing contracts that assume continuity.

Pricing and engagement models, pros, cons and red flags

  • Fixed-scope discovery / POC
    When it fits: de-risk a single hypothesis. Red flag: vague success criteria and scope creep. Mitigation: require acceptance tests and a measurable KPI to trigger next-phase funding.
  • Time-and-materials
    When it fits: exploratory or changing scope. Red flag: opaque seniority mix and runaway hours. Mitigation: require a rate card, a weekly burn report and senior-time caps.
  • Dedicated squad (per-person-per-month)
    When it fits: sustained delivery and ops. Red flag: low seniority on critical roles. Mitigation: require named roles, minimum senior FTE allocation and a retention guarantee.
  • Outcome-based deals
    When it fits: measurable downstream business impact you can reliably attribute to model performance. Red flag: mis-specified outcomes and measurement disputes. Mitigation: define attribution, measurement windows and dispute resolution in the contract.

Due-diligence checklist, require these items in writing

  • Named delivery team members with minimum retention (e.g., named leads committed for X months) and replacement standards.
  • Contactable production references with verification that the vendor delivered ongoing operational support (not a handover).
  • Monitoring & retraining deliverables: dashboard access, alert definitions, retraining triggers and automated retrain proof-of-concept where applicable.
  • SLOs and incident response targets (example: 99.9% serving availability; under 4-hour Sev-1 response) and an on-call escalation path.
  • Clear IP and data ownership clauses: who owns fine-tuned models, derived datasets and any accelerators used.
  • Exportability and migration plan: a tested export of models, weights and metadata within a contractual window (e.g., 30 days).
  • Regulatory and explainability deliverables where required: model cards, audit logs, decision-level explainability artifacts.

Common board questions, short, honest answers

  • Which vendors should we shortlist now?
    Shortlist a platform aligned with your cloud and data architecture (Databricks, AWS, Google, Azure) and a services partner with industry and regulatory experience (Accenture, Deloitte, Fractal, QuantumBlack). Match vendor capabilities to the outcome you need: platform for reuse and scale, services for embedment and organizational change.
  • How do we avoid “pilot‑itis”?
    Treat production as a separate phase with deliverables: monitoring, retraining, ownership and SLAs. Require these as contract deliverables and validate with a live production reference.
  • When is a hosted LLM API (OpenAI/Bedrock) appropriate?
    When speed matters and you accept per-token economics and limited model ownership. For high-volume, IP-sensitive or long-running workloads, model licensing or open-weight hosting (Hugging Face or self-hosted) usually lowers long-term TCO and increases control.
  • Are hyperscalers always the right pick?
    Hyperscalers make sense when your estate already lives there. Otherwise evaluate lock-in, migration costs and whether warehouse-adjacent ML (BigQuery ML, Snowflake integrations) better fits your analytics flows.

Five hiring questions to bring to vendor interviews

  1. Is the bottleneck the model or the pipeline?
  2. What decision changes when the prediction exists, and who makes it?
  3. Who owns it in month nine?
  4. Who signs off when the model is wrong?
  5. Are you buying capability (expertise, IP, architecture) or capacity (hands to execute)?

Follow these with a documented request: named delivery staff, a live production reference, and contract clauses requiring monitoring, retraining and an incident escalation path.

What we didn’t evaluate here (and what you must verify)

  • No hands-on benchmarking between vendors, performance varies by dataset and integration.
  • No review of vendor headcounts, hourly rates or confidential client contracts, verify these directly during procurement.
  • No legal review of supplier ownership changes or disputed market reporting, confirm current ownership and customer relationships before signing.

Practical next steps for procurement

  • Build an RFP that separates platform requirements from services requirements and includes the measurable MLOps artifacts above.
  • Score vendors on production evidence (live deployments, monitoring artifacts, SLOs) as much as on capability or price.
  • Require a migration test before go-live and include exit mechanics in the contract so you can move models and data if relationships change.

Key takeaways, questions and honest, short answers

  • How do I tell a production partner from a pilot vendor?
    Ask for contactable production references, monitoring dashboards, retraining triggers and named on-call contacts, if they can’t provide these, you’re buying a pilot.
  • What are the non-negotiables in the contract?
    Ownership of artifacts, exportability, retraining and monitoring deliverables, incident SLOs, and a tested migration plan.
  • When should we use hosted LLM APIs vs open or self-hosted models?
    Hosted APIs for speed and lower operational burden. Open or self-hosted for cost predictability at scale, IP control and regulatory needs.
  • How should we price vendor engagements to avoid surprises?
    Use acceptance tests in fixed-scope POCs, transparent rate cards for T&M, named staffing and retention guarantees for squads, and clear measurement windows for outcome deals.
  • What single change reduces the most risk?
    Treat production as a distinct phase and contract for operational deliverables up front, monitoring, retraining, named on-call and a migration clause.

Buying ML is three decisions, not one: pick the platform that fits your data and cloud estate, choose the services partner that brings the operational muscle and industry knowledge you lack, and lock the contract on operational deliverables before you pay for models. If you want a plug-and-play RFP template that maps these asks to evaluation weights, that can be converted into a vendor scorecard and shared with procurement teams.