← All articlesML Engineering

    Machine Learning in Production: From Prototype to Contractual SLA with MLOps

    A technical guide to transforming machine learning prototypes into monitored, reproducible services covered by an SLA.

    October 07, 2026 · 8 min read

    Putting machine learning into production requires transforming an experimental model into a reproducible, observable, secure service that can operate under contractual targets. MLOps enables this transition by controlling data, code, artifacts, infrastructure, and metrics, while the SLA objectively defines availability, latency, quality, support, and responsibilities.

    Why a Good Prototype Is Still Not a Product

    A notebook may demonstrate that there is a predictive signal in the data, but it usually does not answer the questions required to operate the model continuously:

    • Do production data have the same format and distribution as the training data?
    • Is it possible to reproduce a previous version exactly?
    • What happens when a variable arrives empty or late?
    • What is the 95th-percentile latency, rather than just the average?
    • How can a decline in quality be detected when the ground-truth label takes days to arrive?
    • Who approves, publishes, monitors, and rolls back a version?
    • Which component is covered by the SLA?

    In conventional projects, code implements explicit rules. In machine learning, behavior also depends on data, features, hyperparameters, and the runtime environment. A seemingly harmless change to a SQL query can alter the input distribution and degrade the model without causing a technical error.

    Therefore, the deployable unit should not be just a weights file. At a minimum, it must include inference code, transformations, input schema, dependencies, configuration, training metadata, and acceptance criteria.

    The Stages Between Experiment and Production

    1. Defining the Business Decision

    Before choosing an algorithm, define which decision will be supported and the cost of each error. Accuracy alone is rarely sufficient.

    In clinical risk detection, for example, false negatives and false positives have different consequences. In demand forecasting, the statistical metric must be related to stockouts, inventory, and tied-up capital. In service request classification, it is also necessary to measure time saved and incorrect routing.

    The technical contract must document:

    • the population and situations in which the model may be used;
    • the prediction unit and time horizon;
    • the target variable and label source;
    • minimum metrics for each relevant segment;
    • known limitations and prohibited uses;
    • the expected action for each result range.

    2. Reproducible Prototype

    Initial validation must separate training, validation, and testing without temporal or identity leakage. If records from the same patient, machine, or customer appear on both sides of the split, the evaluation may become artificially optimistic.

    To make the experiment reproducible, record:

    • the version of the code and dependencies;
    • the data period and version;
    • the definition of each feature;
    • the random seed and hyperparameters;
    • global and subgroup metrics;
    • the generated artifact and runtime environment.

    The outcome of this phase is not merely “a model with 92%.” It is a traceable package accompanied by a report explaining the evaluated population, time range, baseline, decision threshold, and limitations.

    3. Inference Engineering

    The next step is choosing how the prediction will be consumed:

    • Batch: suitable for periodic predictions, such as daily propensity or weekly demand.
    • Synchronous API: appropriate when another system needs the response during an operation.
    • Streaming: required when continuous events demand low latency.
    • Embedded execution: useful when connectivity, privacy, or response time prevents external calls.

    The choice changes costs and the SLA. A real-time API requires capacity, load balancing, timeouts, queues, a circuit breaker, and scalability. Batch processing tolerates higher latency but requires controls for completeness, processing windows, and idempotent reruns.

    Data contracts must validate types, required fields, ranges, allowed categories, and semantics. Invalid inputs must not be silently converted into plausible predictions.

    Minimum MLOps Architecture

    A production architecture does not need to start out complex, but it must cover the entire lifecycle:

    1. Versioning: code, data or immutable references, configuration, and artifacts.
    2. Data pipeline: ingestion, validation, transformation, and feature generation.
    3. Training pipeline: reproducible execution, evaluation, and conditional publishing.
    4. Model registry: versions, metrics, stage, and the person responsible for approval.
    5. CI/CD/CT: code, data, and model tests; deployment; and controlled training when necessary.
    6. Serving: batch service, API, stream, or embedded component.
    7. Observability: logs, metrics, traces, drift, and predictive quality.
    8. Governance: permissions, auditing, documentation, retention, and rollback.

    CI/CD does not mean automatically promoting every new model. In critical contexts, automation should run tests while promotion remains subject to human approval and documented evidence. Continuous training is not mandatory either: retraining without criteria may replace a stable version with a worse one.

    What to Monitor After Deployment

    Operational Metrics

    Technical metrics show whether the service is working:

    • monthly availability;
    • request or record volume;
    • error and timeout rates;
    • p50, p95, and p99 latency;
    • CPU, memory, and accelerator consumption;
    • queue size and processing delay;
    • cost per thousand inferences.

    Averages hide tails. An average latency of 100 ms does not guarantee a good experience if p99 exceeds several seconds.

    Data and Model Metrics

    It is also necessary to monitor:

    • missing fields and schema violations;
    • feature distributions;
    • new categories;
    • input and output drift;
    • probability calibration;
    • precision, recall, specificity, F1, AUC, or error, depending on the use case;
    • metrics separated by relevant segments.

    Drift does not prove degradation, but it indicates that the population has changed and requires investigation. Likewise, the absence of drift does not guarantee quality: the relationship between inputs and outcomes may have changed.

    When labels arrive with a delay, use two levels. Immediate monitoring tracks schema, distribution, and prediction behavior; delayed monitoring calculates actual performance when outcomes become available.

    How to Convert Metrics into a Contractual SLA

    An SLA is a measurable commitment between the provider and the customer. It must be distinguished from an SLO, the operational target, and an SLI, the indicator used for measurement.

    Structural example:

    • SLI: proportion of valid requests answered successfully.
    • SLO: monthly availability of 99.9%.
    • SLA: contractual obligation based on this indicator, including exclusions, calculation, and consequences of noncompliance.

    To avoid ambiguity, each item must define the formula, data source, window, percentile, time zone, exclusions, and the party responsible for calculation.

    Availability and Latency

    The SLA must clarify whether availability covers only the inference API or also ingestion, database, authentication, and external integrations. A 99.9% target allows approximately 43 minutes of downtime in a 30-day month; 99.99% reduces that margin to approximately 4 minutes and 19 seconds, significantly increasing architectural costs.

    For latency, prefer percentiles: for example, p95 below the agreed limit for valid requests, measured at the service entry point. Also define the maximum payload and load conditions.

    Predictive Quality

    Model quality should not be promised in the same way as availability. It depends on the data received, event prevalence, label delay, and use within the validated population.

    A more defensible commitment specifies:

    • the metric and decision threshold;
    • the evaluation dataset or period;
    • the minimum sample size;
    • the source and quality of labels;
    • the evaluated segments;
    • the procedure to follow when the metric declines;
    • responsibilities regarding data changes.

    In many cases, the SLA should guarantee monitoring and a response to degradation, while minimum quality remains conditional on the documented assumptions.

    Incidents, Support, and Recovery

    Classify incidents by severity and associate them with different response times. The contract must also define RTO, the maximum time to restore operations, and RPO, the acceptable amount of data loss measured in time.

    Include procedures for rollback, degraded operation, use of fallback rules, communication, post-incident reporting, and escalation. If the model influences critical decisions, there must be a safe way to continue operating without it.

    Production Readiness Checklist

    Before release, confirm:

    • [ ] the objective and population of use are documented;
    • [ ] the baseline and acceptance criteria have been defined;
    • [ ] there is no leakage between training and testing;
    • [ ] data, code, and model can be reproduced;
    • [ ] the input schema is validated;
    • [ ] unit, integration, load, and security tests have been performed;
    • [ ] operational and predictive metrics have alerts;
    • [ ] version records and approval exist;
    • [ ] gradual, canary, or shadow deployment has been considered;
    • [ ] rollback has been tested;
    • [ ] permissions, secrets, and sensitive data are protected;
    • [ ] SLIs, SLOs, and the SLA use verifiable formulas;
    • [ ] responsibility for incidents and retraining has been defined;
    • [ ] external dependencies and contractual exclusions are documented.

    Trade-Offs That Must Be Explicit

    Performance versus explainability: more complex models may improve a metric but make auditing and diagnosis more difficult. The choice depends on the risk of the decision, not on technological preference.

    Latency versus cost: maintaining idle capacity reduces response time but increases spending. Batch or asynchronous processing may be better when the decision is not immediate.

    Automation versus control: automatic promotion accelerates cycles but may be unsuitable in healthcare, finance, or critical operations. Manual gates increase governance and delivery time.

    Retraining frequency versus stability: frequent cycles capture changes but increase operational risk and validation costs. The trigger should combine schedule, drift, quality, and business context.

    Aggressive SLA versus complexity: high availability requires redundancy, recovery testing, on-call support, and less dependence on single points of failure. The contracted level must reflect the actual impact of downtime.

    How Predictor Solutions Solves This

    Predictor Solutions treats machine learning as a production system, not as an isolated artifact. Its approach combines data engineering, applied artificial intelligence, cloud/DevOps, security, and custom development to create reproducible pipelines, observable APIs, version controls, monitoring, and measurable contractual criteria.

    In healthcare, the company works with HL7 v2 and FHIR integrations and maintains the products Predictor Health, focused on dashboards and wearables, and Predictor AI Hospitals, aimed at predicting sepsis, heart attacks, and pneumonia in the ICU. These scenarios require data contracts, traceability, population-based validation, resilient integration, and safe operating mechanisms when predictions are unavailable.

    Predictor Solutions has served 9 medium-sized and large companies. Across completed projects, the reported consolidated results include average savings of R$ 1.32 million per customer per year, an average productivity increase of 70%, and profit growth of 43% in six months; these indicators are portfolio results and do not replace the definition of specific metrics for each ML implementation.

    Contact: contato@predictorsolutions.com / WhatsApp +55 31 98835-3246

    Frequently asked questions

    How do you take a machine learning model from a notebook to production?

    First, make the data, code, dependencies, and artifacts reproducible. Then implement input validation, a deployment pipeline, version records, monitoring, rollback, and objective acceptance criteria before exposing the model as an API, batch process, or stream.

    What should be included in a machine learning SLA?

    The SLA should define availability, percentile-based latency, component scope, support, incident severity, RTO, RPO, and the calculation method. Predictive quality requires additional conditions, such as the valid population, label source, minimum sample size, metric, threshold, and responsibilities for data changes.

    What is the difference between MLOps, CI/CD, and model monitoring?

    CI/CD automates code and artifact testing and deployment. MLOps covers a broader lifecycle, including data, training, registration, approval, serving, drift monitoring, delayed evaluation, and governance; monitoring is only one part of this lifecycle.

    Does drift mean that the model needs to be retrained?

    Not necessarily. Drift indicates changes in the data or predictions, but retraining should only occur after verifying the impact on quality, the cause of the change, label availability, and approval criteria. Automatically retraining without validation can make the system worse.

    How much does it cost to require 99.99% availability for an AI API?

    There is no universal figure because the cost depends on traffic, infrastructure, redundancy, dependencies, and support. However, 99.99% allows approximately 4 minutes and 19 seconds of downtime in a 30-day month, which typically requires redundant architecture, observability, tested recovery, and more expensive operations.

    Keep reading