Skip to main content
AI provenance is Earnie’s third domain of risk, alongside open source and cryptography: identifying the AI your code depends on. It’s available where your organisation has the AI module enabled, next to Cryptography in scanner selection. The AI scanner produces two kinds of finding:
  • Model findings — an AI model file in your repository or uploaded directly, with what’s known about its licence, lineage, and datasets. This page is mostly about these.
  • AI-usage findings — code that calls an AI service or SDK, rather than a model file itself. These are identified and triaged the same way as model findings (see AI Model Findings), but this page doesn’t cover them in detail.
This page explains how a model is identified, the two ways a model reaches Earnie, and how to write policies against what’s found.

Identifying a Model File

When the AI scanner runs, a model file committed to your repository — for example .safetensors, .gguf, .bin, .pt/.pth (PyTorch), .onnx, .h5/.keras, or a similar weights format — is identified as a model, not just skipped as a binary. Earnie reports:
  • A purl identifying the model.
  • How it matched, and a confidence band shown as text, never as a raw score.
  • The nearest known relatives the matcher couldn’t fully separate it from, when applicable.
  • What the SCANOSS model catalogue records: licence, lineage (base models it was derived from), datasets, and the model card.
A model that doesn’t match anything, or can’t be matched, still produces a finding rather than disappearing silently. Its status explains why:
A model that could not be matched is still counted as scanned, not skipped. Coverage stays full, so the absence of a matched model is never silently swept away as if the file had never existed.

Uploading a Model

Not every model your organisation uses lives in a scanned repository. From a project’s Models page, you can upload one directly:
  • One file, or a Hugging Face snapshot as a tar archive — you don’t need to download and re-upload a model by hand.
  • Up to your organisation’s configured cap (3 GB by default).
  • The upload resumes if your connection drops or you reload the page: select the same file again, and Earnie continues from where it left off instead of starting over.
An uploaded model is scanned the same way a committed one is: identification, licence and lineage lookup, triage, policy evaluation, and inclusion in the SBOM are all unchanged. It appears on the project’s Models page with the same finding detail a committed model gets, linked into the scan that identified it.
The weights aren’t kept. Earnie deletes the uploaded file once its scan finishes, whatever the outcome. No fingerprint is retained either, so re-identifying the same model against a newer catalogue means uploading it again. A model over the size cap is refused before any upload starts, with a note that fingerprinting a model that large isn’t available yet.
Deleting a model from the Models page marks its finding absent and removes the stored bytes; its triage history is kept, as any resolved finding’s is.

Git LFS Models

A repository that keeps its models in Git LFS commits only pointers, so by default an AI scan reports each one as LFS pointer, with no way to fingerprint it. An Operator or Admin can opt a project in to fetching them: Project → Scan Configuration → AI provenance. Once enabled, a repository scan that runs the AI scanner fetches each pointed-to model (up to the same size cap as an upload) and identifies it like a committed file.
Off by default. Fetching an LFS object spends the repository’s own LFS bandwidth quota, so this only happens for a project that explicitly opts in. A model that fails to fetch is left as a pointer rather than failing the scan.

Writing Policies Against AI Findings

Every finding carries an ai field group once the AI scanner has run, with kind distinguishing a model finding ("model") from an AI-usage finding ("usage"): code that calls an AI service or SDK, rather than a model file. These are triaged separately too — see AI Model Findings in Triaging Findings. The model fields most policies need: An AI-usage finding carries its own fields instead, such as finding.ai.vendor, finding.ai.sdk, and finding.ai.provider, which identify the AI service or library the code calls. Because these fields are empty on findings they don’t apply to, a rule that reads them evaluates false rather than erroring on other kinds of finding. The Sentence builder offers no negation on AI fields, because “AI SDK is not openai” would also match every open-source finding. AI Vendor is the exception: it also offers is not and is not one of, for an approved-vendors list. Keep a negated vendor beside AI Finding Kind is usage and AI Vendor is not empty, as the Enforce approved AI vendors template does, or the rule also matches every finding that names no vendor.

Starter Templates

Seven built-in templates cover the AI risks most organisations start with. Six evaluate model findings; Enforce approved AI vendors evaluates AI-usage findings instead. Add them from New policy → Pick templates, the same as any other template:
Every AI template waits on the AI scanner having run. A scan that skipped it reports the policy as not evaluated, never a silent pass.
A CRA — Annex I Part II(1) regulatory template set bundles the unidentified-model and restricted-licence templates in one action, the same way the cryptography regulatory template sets work. See Setting Policies for how templates, the rule builder, and template sets work in general.

In the SBOM

An identified model is exported as a CycloneDX machine-learning-model component: its purl, SHA-256, licence, lineage (as pedigree ancestors), and model card. There’s no separate include switch for models, unlike cryptographic assets — a model ships with the SBOM by default whenever your organisation has the AI module.

What’s Next

An AI finding is triaged, reviewed, and gated the same way an open-source or cryptography finding is. See Triaging Findings and Remediation for what happens once one is found.