Putting machine learning into production requires transforming an experimental model into a reproducible, observable, secure service that can operate under contractual targets. MLOps enables this transition by controlling data, code, artifacts, infrastructure, and metrics, while the SLA objectively defines availability, latency, quality, support, and responsibilities.
Why a Good Prototype Is Still Not a Product
A notebook may demonstrate that there is a predictive signal in the data, but it usually does not answer the questions required to operate the model continuously:
- Do production data have the same format and distribution as the training data?
- Is it possible to reproduce a previous version exactly?
- What happens when a variable arrives empty or late?
- What is the 95th-percentile latency, rather than just the average?
- How can a decline in quality be detected when the ground-truth label takes days to arrive?
- Who approves, publishes, monitors, and rolls back a version?
- Which component is covered by the SLA?
In conventional projects, code implements explicit rules. In machine learning, behavior also depends on data, features, hyperparameters, and the runtime environment. A seemingly harmless change to a SQL query can alter the input distribution and degrade the model without causing a technical error.
Therefore, the deployable unit should not be just a weights file. At a minimum, it must include inference code, transformations, input schema, dependencies, configuration, training metadata, and acceptance criteria.
The Stages Between Experiment and Production
1. Defining the Business Decision
Before choosing an algorithm, define which decision will be supported and the cost of each error. Accuracy alone is rarely sufficient.
In clinical risk detection, for example, false negatives and false positives have different consequences. In demand forecasting, the statistical metric must be related to stockouts, inventory, and tied-up capital. In service request classification, it is also necessary to measure time saved and incorrect routing.
The technical contract must document:
- the population and situations in which the model may be used;
- the prediction unit and time horizon;
- the target variable and label source;
- minimum metrics for each relevant segment;
- known limitations and prohibited uses;
- the expected action for each result range.
2. Reproducible Prototype
Initial validation must separate training, validation, and testing without temporal or identity leakage. If records from the same patient, machine, or customer appear on both sides of the split, the evaluation may become artificially optimistic.
To make the experiment reproducible, record:
- the version of the code and dependencies;
- the data period and version;
- the definition of each feature;
- the random seed and hyperparameters;
- global and subgroup metrics;
- the generated artifact and runtime environment.
The outcome of this phase is not merely “a model with 92%.” It is a traceable package accompanied by a report explaining the evaluated population, time range, baseline, decision threshold, and limitations.
3. Inference Engineering
The next step is choosing how the prediction will be consumed:
- Batch: suitable for periodic predictions, such as daily propensity or weekly demand.
- Synchronous API: appropriate when another system needs the response during an operation.
- Streaming: required when continuous events demand low latency.
- Embedded execution: useful when connectivity, privacy, or response time prevents external calls.
The choice changes costs and the SLA. A real-time API requires capacity, load balancing, timeouts, queues, a circuit breaker, and scalability. Batch processing tolerates higher latency but requires controls for completeness, processing windows, and idempotent reruns.
Data contracts must validate types, required fields, ranges, allowed categories, and semantics. Invalid inputs must not be silently converted into plausible predictions.
Minimum MLOps Architecture
A production architecture does not need to start out complex, but it must cover the entire lifecycle:
- Versioning: code, data or immutable references, configuration, and artifacts.
- Data pipeline: ingestion, validation, transformation, and feature generation.
- Training pipeline: reproducible execution, evaluation, and conditional publishing.
- Model registry: versions, metrics, stage, and the person responsible for approval.
- CI/CD/CT: code, data, and model tests; deployment; and controlled training when necessary.
- Serving: batch service, API, stream, or embedded component.
- Observability: logs, metrics, traces, drift, and predictive quality.
- Governance: permissions, auditing, documentation, retention, and rollback.
CI/CD does not mean automatically promoting every new model. In critical contexts, automation should run tests while promotion remains subject to human approval and documented evidence. Continuous training is not mandatory either: retraining without criteria may replace a stable version with a worse one.
What to Monitor After Deployment
Operational Metrics
Technical metrics show whether the service is working:
- monthly availability;
- request or record volume;
- error and timeout rates;
- p50, p95, and p99 latency;
- CPU, memory, and accelerator consumption;
- queue size and processing delay;
- cost per thousand inferences.
Averages hide tails. An average latency of 100 ms does not guarantee a good experience if p99 exceeds several seconds.
Data and Model Metrics
It is also necessary to monitor:
- missing fields and schema violations;
- feature distributions;
- new categories;
- input and output drift;
- probability calibration;
- precision, recall, specificity, F1, AUC, or error, depending on the use case;
- metrics separated by relevant segments.
Drift does not prove degradation, but it indicates that the population has changed and requires investigation. Likewise, the absence of drift does not guarantee quality: the relationship between inputs and outcomes may have changed.
When labels arrive with a delay, use two levels. Immediate monitoring tracks schema, distribution, and prediction behavior; delayed monitoring calculates actual performance when outcomes become available.
How to Convert Metrics into a Contractual SLA
An SLA is a measurable commitment between the provider and the customer. It must be distinguished from an SLO, the operational target, and an SLI, the indicator used for measurement.
Structural example:
- SLI: proportion of valid requests answered successfully.
- SLO: monthly availability of 99.9%.
- SLA: contractual obligation based on this indicator, including exclusions, calculation, and consequences of noncompliance.
To avoid ambiguity, each item must define the formula, data source, window, percentile, time zone, exclusions, and the party responsible for calculation.
Availability and Latency
The SLA must clarify whether availability covers only the inference API or also ingestion, database, authentication, and external integrations. A 99.9% target allows approximately 43 minutes of downtime in a 30-day month; 99.99% reduces that margin to approximately 4 minutes and 19 seconds, significantly increasing architectural costs.
For latency, prefer percentiles: for example, p95 below the agreed limit for valid requests, measured at the service entry point. Also define the maximum payload and load conditions.
Predictive Quality
Model quality should not be promised in the same way as availability. It depends on the data received, event prevalence, label delay, and use within the validated population.
A more defensible commitment specifies:
- the metric and decision threshold;
- the evaluation dataset or period;
- the minimum sample size;
- the source and quality of labels;
- the evaluated segments;
- the procedure to follow when the metric declines;
- responsibilities regarding data changes.
In many cases, the SLA should guarantee monitoring and a response to degradation, while minimum quality remains conditional on the documented assumptions.
Incidents, Support, and Recovery
Classify incidents by severity and associate them with different response times. The contract must also define RTO, the maximum time to restore operations, and RPO, the acceptable amount of data loss measured in time.
Include procedures for rollback, degraded operation, use of fallback rules, communication, post-incident reporting, and escalation. If the model influences critical decisions, there must be a safe way to continue operating without it.
Production Readiness Checklist
Before release, confirm:
- [ ] the objective and population of use are documented;
- [ ] the baseline and acceptance criteria have been defined;
- [ ] there is no leakage between training and testing;
- [ ] data, code, and model can be reproduced;
- [ ] the input schema is validated;
- [ ] unit, integration, load, and security tests have been performed;
- [ ] operational and predictive metrics have alerts;
- [ ] version records and approval exist;
- [ ] gradual, canary, or shadow deployment has been considered;
- [ ] rollback has been tested;
- [ ] permissions, secrets, and sensitive data are protected;
- [ ] SLIs, SLOs, and the SLA use verifiable formulas;
- [ ] responsibility for incidents and retraining has been defined;
- [ ] external dependencies and contractual exclusions are documented.
Trade-Offs That Must Be Explicit
Performance versus explainability: more complex models may improve a metric but make auditing and diagnosis more difficult. The choice depends on the risk of the decision, not on technological preference.
Latency versus cost: maintaining idle capacity reduces response time but increases spending. Batch or asynchronous processing may be better when the decision is not immediate.
Automation versus control: automatic promotion accelerates cycles but may be unsuitable in healthcare, finance, or critical operations. Manual gates increase governance and delivery time.
Retraining frequency versus stability: frequent cycles capture changes but increase operational risk and validation costs. The trigger should combine schedule, drift, quality, and business context.
Aggressive SLA versus complexity: high availability requires redundancy, recovery testing, on-call support, and less dependence on single points of failure. The contracted level must reflect the actual impact of downtime.
How Predictor Solutions Solves This
Predictor Solutions treats machine learning as a production system, not as an isolated artifact. Its approach combines data engineering, applied artificial intelligence, cloud/DevOps, security, and custom development to create reproducible pipelines, observable APIs, version controls, monitoring, and measurable contractual criteria.
In healthcare, the company works with HL7 v2 and FHIR integrations and maintains the products Predictor Health, focused on dashboards and wearables, and Predictor AI Hospitals, aimed at predicting sepsis, heart attacks, and pneumonia in the ICU. These scenarios require data contracts, traceability, population-based validation, resilient integration, and safe operating mechanisms when predictions are unavailable.
Predictor Solutions has served 9 medium-sized and large companies. Across completed projects, the reported consolidated results include average savings of R$ 1.32 million per customer per year, an average productivity increase of 70%, and profit growth of 43% in six months; these indicators are portfolio results and do not replace the definition of specific metrics for each ML implementation.
Contact: contato@predictorsolutions.com / WhatsApp +55 31 98835-3246