Putting machine learning into production requires more than publishing an API: data, versions, infrastructure, predictive quality, security, and incident response must all be controlled. A sustainable contractual SLA should cover only measurable and controllable indicators, such as availability, latency, and restoration time, while model metrics need their own SLOs, revalidation rules, and explicit limits of responsibility.
Why a prototype is not ready for production
A notebook proves that a given approach can work on a known sample. It does not usually demonstrate that the solution will remain available, fast, and reliable when it receives real-world data, distribution shifts, concurrent access, or invalid inputs.
The difference between demonstration and operation appears across five dimensions:
- Reproducibility: code, data, parameters, and environment must be versioned.
- Reliability: the service must respond correctly even under partial failures.
- Observability: the team and the client need to know when data, the model, or infrastructure degrades.
- Governance: each prediction must be traceable to the version that produced it.
- Operations: incidents require designated owners, rollback procedures, and defined deadlines.
It is also necessary to distinguish software failure from predictive error. An API can have 99.9% availability while delivering unsuitable predictions because the population being served has changed. Conversely, a model can retain good offline accuracy but become unavailable because of infrastructure issues.
The minimum MLOps architecture
MLOps applies software engineering, data engineering, and operations principles to the model lifecycle. It does not necessarily require an expensive platform, but it does require clear components and responsibilities.
Versioning and traceability
Each release must record, at a minimum:
- the commit of the training and inference code;
- the version or immutable reference of the data;
- hyperparameters and random seeds;
- dependencies and the runtime image;
- validation metrics;
- technical or business approval;
- the version of the input and output schema.
Saving only the model file is not enough. Without its lineage, the team cannot reproduce a result, investigate a prediction, or safely restore a previous version.
Training pipeline
Training must move out of the manual notebook and into a repeatable pipeline. A typical sequence includes data validation, transformation, training, evaluation, bias testing when applicable, packaging, and artifact registration.
The pipeline should fail automatically when there are, for example:
- missing columns or incompatible types;
- an abnormal increase in null values;
- leakage between training and validation;
- a metric below the approved threshold;
- a vulnerable dependency or an unverifiable image;
- incompatibility between training and production features.
Serving and integration
Inference can be synchronous, asynchronous, batch-based, or embedded. The choice affects cost, latency, and complexity:
- Synchronous API: suitable when the system needs the prediction immediately; it requires strict latency and availability controls.
- Asynchronous queue: tolerates later processing and absorbs spikes, but increases total response time.
- Batch: reduces costs for large volumes without urgency, but produces stale data.
- Edge or embedded: reduces network dependency but makes updates and telemetry more difficult.
In critical systems, model unavailability should not necessarily interrupt the entire workflow. A fallback can be defined using a business rule, the last valid prediction, or referral for human review, provided that the behavior is documented.
CI/CD/CT: three different types of automation
In MLOps, continuous integration and delivery do not cover the entire lifecycle. There are three complementary workflows:
- CI: tests code, data contracts, transformations, and integration.
- CD: deploys approved services, configurations, and models.
- CT: reruns training when sufficient data or signs of degradation emerge.
Continuous training does not mean automatic promotion. In higher-risk sectors, a new model can be trained automatically but should remain a candidate until it passes validation and approval. This separation reduces the chance that a data change will promote an inferior model directly to production.
A safe deployment strategy uses shadow, canary, or champion-challenger. In shadow mode, the candidate receives copies of production traffic without influencing decisions. In a canary deployment, it serves a limited fraction of requests. In champion-challenger, its performance is compared with the current model over a defined period.
What to monitor after deployment
Monitoring must cover four layers. Monitoring only CPU and memory leaves the main source of risk invisible.
Infrastructure and service
The most common indicators are availability, latency at the p50, p95, and p99 percentiles, error rate, saturation, resource consumption, and queue size. Percentiles are preferable to averages because they reveal the experience of the slowest requests.
Data quality
Schema, range, cardinality, null values, unknown categories, and update delays should be monitored. A valid change in the source system can silently break a feature even if the API continues responding with HTTP 200.
Model behavior
Data drift indicates a change in the input distribution. Concept drift occurs when the relationship between input and outcome changes. Neither one alone proves that the model has lost quality, but both justify investigation.
When the ground-truth label arrives late, metrics such as precision, recall, F1, MAE, or AUC will also be delayed. During this interval, the team can monitor prediction distribution, confidence, decision rate by class, and feature stability without treating them as definitive substitutes for actual performance.
Business impact
The technical metric must be connected to the use case. A reduction in statistical error does not guarantee less fraud, shorter service times, or higher conversion. Monitoring should record what action occurred after the prediction and, when possible, compare the operational outcome with a baseline.
How to transform metrics into SLIs, SLOs, and SLAs
The three concepts are not equivalent:
- SLI: an observed measurement, such as the percentage of valid requests answered within 300 ms.
- SLO: an internal target, such as keeping that percentage above 99.5% per month.
- SLA: a contractual commitment with scope, exclusions, and consequences for noncompliance.
The internal SLO should be stricter than the SLA. This margin creates an error budget for maintenance, deployments, and incidents without immediately violating the contract.
An inference SLA should clarify:
- the covered endpoint, region, schedule, and environment;
- the availability formula and measurement window;
- the latency percentile and limit;
- the contracted volume and behavior above that volume;
- the definition of incident and severity;
- acknowledgment and restoration times;
- scheduled maintenance and exclusions;
- dependencies under the client’s responsibility;
- the telemetry source accepted for auditing;
- applicable compensation or service credit.
It is not advisable to contractually promise fixed accuracy without defining the population, time window, ground-truth label availability, and minimum data quality. The metric may depend on external phenomena that the provider does not control. An alternative is to contract the process: monitoring, evaluation frequency, revalidation triggers, and mitigation deadline.
Example of an operational matrix
| Indicator | Illustrative internal target | Possible contractual commitment |
|---|---:|---:|
| Monthly availability | 99.95% | 99.9% |
| p95 latency | up to 250 ms | up to 300 ms |
| Critical incident acknowledgment | 15 min | 30 min |
| Restoration or fallback | 2 h | 4 h |
| Relevant drift | automatic alert | analysis within the agreed deadline |
| Predictive quality | threshold by use case | periodic review, if measurable |
The values are illustrative. They must be calculated according to architecture, criticality, budget, volume, and external dependencies. The difference between 99.9% and 99.99% availability may require regional redundancy, on-call support, and substantially higher costs.
Checklist for moving from prototype to SLA
Before making a contractual commitment, verify that:
- [ ] the problem, population, and supported decision are defined;
- [ ] the business baseline and statistical baseline have been recorded;
- [ ] data, code, model, and environment are reproducible;
- [ ] input and output contracts have automated validation;
- [ ] unit, integration, load, and model regression tests are in place;
- [ ] secrets and personal data do not appear in code or logs;
- [ ] access follows the principle of least privilege and leaves an audit trail;
- [ ] deployment supports canary, shadow, or rollback;
- [ ] dashboards cover the service, data, model, and business;
- [ ] alerts have designated owners and executable procedures;
- [ ] the fallback has been tested, not merely documented;
- [ ] the SLI has a formula, source, and measurement window;
- [ ] SLA exclusions are explicit;
- [ ] cost per prediction and maximum capacity have been measured;
- [ ] model revalidation and decommissioning have objective criteria.
If several items still depend on unrecorded manual actions, the service remains at the pilot stage. The SLA should reflect the existing operational capability, not the architecture desired for the future.
How Predictor Solutions addresses this
Predictor Solutions structures machine learning projects by combining data engineering, custom software, cloud/DevOps, security, and observability. Its work ranges from prototype validation to the creation of reproducible pipelines, inference APIs, drift monitoring, access controls, rollback strategies, and the technical definition of SLIs and SLOs before SLA negotiations.
The company also applies AI in healthcare products: Predictor Health integrates dashboards and wearables, while Predictor AI Hospitals works with predictions of sepsis, heart attacks, and pneumonia in ICUs. In these contexts, integration through HL7 v2 and FHIR, traceability, and explicit failure handling are parts of the architecture, not later additions.
Based in Lavras, Minas Gerais, Predictor Solutions has already served 9 medium-sized and large companies. Across its project portfolio, it reports average savings of R$ 1.32 million per client per year, an average productivity increase of 70%, and profit growth of 43% in six months; these results are historical references and do not replace specific goals and metrics for each deployment.
Contact: contato@predictorsolutions.com / WhatsApp +55 31 98835-3246