Why DevSecOps for Machine Learning Doesn't Work Out of the Box
In the usual DevSecOps we protect code, dependencies and infrastructure. SAST searches for injections in source, SCA checks CVE in packets, DAST tests HTTP-endpoints. For the usual application - enough. ML-pipeline adds artifacts that are simply not in the traditional conveyor. More details - in our review security llm applications.
Data - datasets for learning and validation. They cannot be scanned by Bandit or Trivy. The poisoned dataset does not generate CVE and does not appear in the advisory.
The model is a binary artifact (weight file) that determines the behavior of the system. Model substitution is equivalent to a business logic substitution, but no tool from the standard DevSecOps stack sees it.
Pipeline Learning - DAG in Kubeflow, Airflow or analogue. Compromising one step (feature engineering, augmentation, validation) castading breaks all the security of AI systems below the stream.
According to the Stanford AI Index 2025, about 78% of the organizations surveyed use AI in at least one business function. At the same time, a large part of the organizations implementing LLM recognizes the lack of maturity to protect against AI-specific threats. Here is the problem: the rate of implementation is growing, and MLSecOps-practices are not. This is a specific entry point for the attackers. According to the CrowdStrike Global Threat Report, the average time of lateral movement within the network continues to decrease (62 to 18 minutes in different reports depends on the year). For the ML infrastructure, where security control is weaker than the main circuit, the response window is still.
MLSecOps (Machine Learning Security Operations) - applying security practices to the ML security lifecycle, from design and data preparation to inference in production. In fact, DevSecOps, adapted to the specifics of secure ML development: data as code, model as a vulnerable artifact, drift as an anomaly.
Machine Learning Safety: TTPs Card Through MITRE ATT&CK
In order for the SOC to work with ML threats, they must be described in TTPs. Below is the mapping attacks on ML-pipeline to MITRE ATT&CK techniques.
ML-Piping stage Attack vector MITRE ATT&CK Technique Tactics
Data collection Dataset substitution through compromised source T1195.002 - Compromise Software Supply Chain Initial Access
Dependencies Malware package in requirements.txt T1195.001 - Compromise Software Dependencies and Development Tools Initial Access
Data storage Access S3/GCS-boat with training data T1530 - Data from Cloud Storage Collection
Model Repository Weight theft from Git LFS / MLflow T1213.003 - Code Repositories Collection
Training Deployment of a malicious container in a cluster T1610 - Deploy Container Execution
Configuration API keys in config.yaml / .env files T1552.001 - Credentials In Files Credential Access
Inference Server Exfiltration through model inversion T1005 - Data from Local System Collection
Preparation of the attack Using ML tools to generate adversarial examples T1588.007 - Artificial Intelligence Resource Development
This mapping is not an academic exercise. Part of the ML threats (T1530, T1552.001, T1610) already covered with standard Sigma packs. It is worth separately T1588.007 (Artificial Intelligence, Resource Development) - using attacking ML tools to prepare adversarial attacks. The SOC command remains to specify the rules for ML-specifics: data poisoning, model substitution, anomalies in the inference API. Half of the work has already been done - it is necessary to finish the remaining.
Classification OWASP LLM Top 10 (2025) is also coming here. LLM04 - Data and Model Poisoning describes manipulation of training data and embeddings. LLM01 - Prompt Injection covers attacks on inference through crafted inputs, bypassing safety controls. LLM10 - Unbounded Consumption - DoS and model extraction through resource exhaustion. Binding to OWASP allows you to use these identifiers when auditing the security of AI models and generating reports for the regulator.
AI system data protection: poisoning and insider threat
Data poisoning is the most insidious attack on the ML system. The consequences are delayed in days and weeks. The model is retrained, passes validation (metrics can even improve on a clean sample) and goes to the prod. The detective happens when the business has already suffered losses. In fact, a time bomb that no scanner can see.
Business logic of attack
Why would an attacker poison data? Three typical motives:
Financial benefit - the scoring model begins to approve fraudulent applications. Classic fraud through compromised pipeline.
Insider sabotage - data scientist with access to the DVC repository and MLflow discreetly replaces the dataset or adjusts the hyperparameters so that the model degrades on a certain segment of the input data in the prod.
The second motive - insider threat - is practically not covered in Russian-language materials on MLSecOps. In vain. Data scientist with write rights in the data-repository and model registry - a privileged user with access to critical artifacts. For SOC, it is an analogue of DBA with rights to the prod database: a person whose actions require separate monitoring and baseline. Understanding the motives of the attacker (external attacker or insider) is the starting point for detection.
What to monitor in SIEM
To detect ML data poisoning and insider attacks, baseline normal data pipeline security behavior is needed:
The volume of commits in the data-repository is an abnormal increase in records or a change in the distribution of tags
Learning Pailine Start Time - Reassembly off schedule or from a non-standard branch
Hyperparameter changes - diff between current and previous experiment in MLflow
Data storage access - access to the S3-boucket with datasets from uncharacteristic IP or service accounts (T1530)
Example of a Sigma rule for detecting an abnormal commit in a data repository:
YAML:
title: Anomalous Data Commit to ML Dataset Repository
status: experimental
logsource:
category: vcs
product: git
detection:
selection:
EventType: push
Repository|contains:
- 'datasets'
- 'training-data'
- 'dvc'
filter_schedule:
User: 'ci-bot'
condition: selection and not filter_schedule
level: medium
The rule will work on a push in a repository with datasets from any user except the CI-bot. Next, the analyst checks the diff: whether the distribution of labels, file size, the format of the data have changed. If the baseline is fixed (the average commit is 500 records, and 12 000) - this is a reason for incident response.
Regulatory context: FZ-152 and negotiable fines
If the training data contain personal data (FZ-152, art. 3), the operator is obliged to ensure their confidentiality (St. 7) and obtain the consent of the subjects for processing with specific purposes (St. 9). The use of PD for model training is a self-directed processing goal that should be explicitly stated in the consent. Leakage or substitution of the dataset with the PD is not only the degradation of the model, but also the basis for negotiable fines. Re-leakage threatens with a negotiable fine on the scale established by law - specific thresholds depend on the volume of the leak and the current version of the Administrative Code. The fine can be millions of rubles, and this is without taking into account the damage from the compromised model.
Protecting ML models from attacks on chain supply
The model is an artifact with the same supply-chain risks as the Docker image. The difference in maturity: for containers already there is Notary/cosign and a culture of signing, and for ML models, the signing of artifacts is not yet everywhere.
According to LegitSecurity, the supply chain ML includes: package managers (PyPI, conda), pre-trained weights (Hugging Face), public datasets, open-source notesbooks, CI/CD for AI models. Each component is a vector for T1195.001 (Compromise Software Dependencies) and T1195.002 (Compromise Software Supply Chain).
Practical control measures by steps:
Step 1. Signing models via cosign (Sigstore). Sign the weight files before placing the registry in the model. If - MLflow, screw the signing as post-hook in CI/CD. Team cosign sign --key cosign.key model-v1.2.onnx takes seconds, but blocks the substitution of the artifact.
Step 2. AI Bill of Materials. Fix for each model: DVC-hash dataset, dependencies with lockfile, git commit pipeline, hash final artifact. It is an analogue of SBOM, but for ML. Without AI BOM, the investigation of the incident turns into archaeology - it is impossible to say on which data the model of semi-annual limitation was trained.
Step 3. Scanning of basic images. ML-papelines use heavy Docker images (nvidia/cuda, tensorflow/serving). Scan them in the same way as any other: trivy image --severity HIGH,CRITICAL <image> will show CVE in the dependencies of the basic image.
Step 4. OPA policies for model registry. Limit the publication of models through policy-as-code. Do not rely on conventions and verbal agreements - they do not experience the rotation of the team.
Example of OPA policy (Rego) for the control of the model deploy:
Code:
package mlsecops.model_deploy
deny[msg] {
input.model.signed != true
msg := "Model artifact must be signed"
}
deny[msg] {
input.pipeline.branch != "main"
msg := "Deploy only from main branch"
}
Two controls - signature and branch - cut off most of the random and part of the purposeful substitution of the artifact. The policy is embedded through OPA/Gatekeeper in the Kubernetes cluster of training and deploy. MLOps security at the policy-as-code level is reproduced, audited and does not depend on whether the ML-engineer remembered about security at the next push.
In the usual DevSecOps we protect code, dependencies and infrastructure. SAST searches for injections in source, SCA checks CVE in packets, DAST tests HTTP-endpoints. For the usual application - enough. ML-pipeline adds artifacts that are simply not in the traditional conveyor. More details - in our review security llm applications.
Data - datasets for learning and validation. They cannot be scanned by Bandit or Trivy. The poisoned dataset does not generate CVE and does not appear in the advisory.
The model is a binary artifact (weight file) that determines the behavior of the system. Model substitution is equivalent to a business logic substitution, but no tool from the standard DevSecOps stack sees it.
Pipeline Learning - DAG in Kubeflow, Airflow or analogue. Compromising one step (feature engineering, augmentation, validation) castading breaks all the security of AI systems below the stream.
According to the Stanford AI Index 2025, about 78% of the organizations surveyed use AI in at least one business function. At the same time, a large part of the organizations implementing LLM recognizes the lack of maturity to protect against AI-specific threats. Here is the problem: the rate of implementation is growing, and MLSecOps-practices are not. This is a specific entry point for the attackers. According to the CrowdStrike Global Threat Report, the average time of lateral movement within the network continues to decrease (62 to 18 minutes in different reports depends on the year). For the ML infrastructure, where security control is weaker than the main circuit, the response window is still.
MLSecOps (Machine Learning Security Operations) - applying security practices to the ML security lifecycle, from design and data preparation to inference in production. In fact, DevSecOps, adapted to the specifics of secure ML development: data as code, model as a vulnerable artifact, drift as an anomaly.
Machine Learning Safety: TTPs Card Through MITRE ATT&CK
In order for the SOC to work with ML threats, they must be described in TTPs. Below is the mapping attacks on ML-pipeline to MITRE ATT&CK techniques.
ML-Piping stage Attack vector MITRE ATT&CK Technique Tactics
Data collection Dataset substitution through compromised source T1195.002 - Compromise Software Supply Chain Initial Access
Dependencies Malware package in requirements.txt T1195.001 - Compromise Software Dependencies and Development Tools Initial Access
Data storage Access S3/GCS-boat with training data T1530 - Data from Cloud Storage Collection
Model Repository Weight theft from Git LFS / MLflow T1213.003 - Code Repositories Collection
Training Deployment of a malicious container in a cluster T1610 - Deploy Container Execution
Configuration API keys in config.yaml / .env files T1552.001 - Credentials In Files Credential Access
Inference Server Exfiltration through model inversion T1005 - Data from Local System Collection
Preparation of the attack Using ML tools to generate adversarial examples T1588.007 - Artificial Intelligence Resource Development
This mapping is not an academic exercise. Part of the ML threats (T1530, T1552.001, T1610) already covered with standard Sigma packs. It is worth separately T1588.007 (Artificial Intelligence, Resource Development) - using attacking ML tools to prepare adversarial attacks. The SOC command remains to specify the rules for ML-specifics: data poisoning, model substitution, anomalies in the inference API. Half of the work has already been done - it is necessary to finish the remaining.
Classification OWASP LLM Top 10 (2025) is also coming here. LLM04 - Data and Model Poisoning describes manipulation of training data and embeddings. LLM01 - Prompt Injection covers attacks on inference through crafted inputs, bypassing safety controls. LLM10 - Unbounded Consumption - DoS and model extraction through resource exhaustion. Binding to OWASP allows you to use these identifiers when auditing the security of AI models and generating reports for the regulator.
AI system data protection: poisoning and insider threat
Data poisoning is the most insidious attack on the ML system. The consequences are delayed in days and weeks. The model is retrained, passes validation (metrics can even improve on a clean sample) and goes to the prod. The detective happens when the business has already suffered losses. In fact, a time bomb that no scanner can see.
Business logic of attack
Why would an attacker poison data? Three typical motives:
Financial benefit - the scoring model begins to approve fraudulent applications. Classic fraud through compromised pipeline.
Insider sabotage - data scientist with access to the DVC repository and MLflow discreetly replaces the dataset or adjusts the hyperparameters so that the model degrades on a certain segment of the input data in the prod.
The second motive - insider threat - is practically not covered in Russian-language materials on MLSecOps. In vain. Data scientist with write rights in the data-repository and model registry - a privileged user with access to critical artifacts. For SOC, it is an analogue of DBA with rights to the prod database: a person whose actions require separate monitoring and baseline. Understanding the motives of the attacker (external attacker or insider) is the starting point for detection.
What to monitor in SIEM
To detect ML data poisoning and insider attacks, baseline normal data pipeline security behavior is needed:
The volume of commits in the data-repository is an abnormal increase in records or a change in the distribution of tags
Learning Pailine Start Time - Reassembly off schedule or from a non-standard branch
Hyperparameter changes - diff between current and previous experiment in MLflow
Data storage access - access to the S3-boucket with datasets from uncharacteristic IP or service accounts (T1530)
Example of a Sigma rule for detecting an abnormal commit in a data repository:
YAML:
title: Anomalous Data Commit to ML Dataset Repository
status: experimental
logsource:
category: vcs
product: git
detection:
selection:
EventType: push
Repository|contains:
- 'datasets'
- 'training-data'
- 'dvc'
filter_schedule:
User: 'ci-bot'
condition: selection and not filter_schedule
level: medium
The rule will work on a push in a repository with datasets from any user except the CI-bot. Next, the analyst checks the diff: whether the distribution of labels, file size, the format of the data have changed. If the baseline is fixed (the average commit is 500 records, and 12 000) - this is a reason for incident response.
Regulatory context: FZ-152 and negotiable fines
If the training data contain personal data (FZ-152, art. 3), the operator is obliged to ensure their confidentiality (St. 7) and obtain the consent of the subjects for processing with specific purposes (St. 9). The use of PD for model training is a self-directed processing goal that should be explicitly stated in the consent. Leakage or substitution of the dataset with the PD is not only the degradation of the model, but also the basis for negotiable fines. Re-leakage threatens with a negotiable fine on the scale established by law - specific thresholds depend on the volume of the leak and the current version of the Administrative Code. The fine can be millions of rubles, and this is without taking into account the damage from the compromised model.
Protecting ML models from attacks on chain supply
The model is an artifact with the same supply-chain risks as the Docker image. The difference in maturity: for containers already there is Notary/cosign and a culture of signing, and for ML models, the signing of artifacts is not yet everywhere.
According to LegitSecurity, the supply chain ML includes: package managers (PyPI, conda), pre-trained weights (Hugging Face), public datasets, open-source notesbooks, CI/CD for AI models. Each component is a vector for T1195.001 (Compromise Software Dependencies) and T1195.002 (Compromise Software Supply Chain).
Practical control measures by steps:
Step 1. Signing models via cosign (Sigstore). Sign the weight files before placing the registry in the model. If - MLflow, screw the signing as post-hook in CI/CD. Team cosign sign --key cosign.key model-v1.2.onnx takes seconds, but blocks the substitution of the artifact.
Step 2. AI Bill of Materials. Fix for each model: DVC-hash dataset, dependencies with lockfile, git commit pipeline, hash final artifact. It is an analogue of SBOM, but for ML. Without AI BOM, the investigation of the incident turns into archaeology - it is impossible to say on which data the model of semi-annual limitation was trained.
Step 3. Scanning of basic images. ML-papelines use heavy Docker images (nvidia/cuda, tensorflow/serving). Scan them in the same way as any other: trivy image --severity HIGH,CRITICAL <image> will show CVE in the dependencies of the basic image.
Step 4. OPA policies for model registry. Limit the publication of models through policy-as-code. Do not rely on conventions and verbal agreements - they do not experience the rotation of the team.
Example of OPA policy (Rego) for the control of the model deploy:
Code:
package mlsecops.model_deploy
deny[msg] {
input.model.signed != true
msg := "Model artifact must be signed"
}
deny[msg] {
input.pipeline.branch != "main"
msg := "Deploy only from main branch"
}
Two controls - signature and branch - cut off most of the random and part of the purposeful substitution of the artifact. The policy is embedded through OPA/Gatekeeper in the Kubernetes cluster of training and deploy. MLOps security at the policy-as-code level is reproduced, audited and does not depend on whether the ML-engineer remembered about security at the next push.