- Model findings — an AI model file in your repository or uploaded directly, with what’s known about its licence, lineage, and datasets. This page is mostly about these.
- AI-usage findings — code that calls an AI service or SDK, rather than a model file itself. These are identified and triaged the same way as model findings (see AI Model Findings), but this page doesn’t cover them in detail.
Identifying a Model File
When the AI scanner runs, a model file committed to your repository — for example.safetensors, .gguf, .bin, .pt/.pth (PyTorch), .onnx, .h5/.keras, or a similar weights format — is identified as a model, not just skipped as a binary. Earnie reports:
- A purl identifying the model.
- How it matched, and a confidence band shown as text, never as a raw score.
- The nearest known relatives the matcher couldn’t fully separate it from, when applicable.
- What the SCANOSS model catalogue records: licence, lineage (base models it was derived from), datasets, and the model card.
A model that could not be matched is still counted as scanned, not skipped. Coverage stays full, so the absence of a matched model is never silently swept away as if the file had never existed.
Uploading a Model
Not every model your organisation uses lives in a scanned repository. From a project’s Models page, you can upload one directly:- One file, or a Hugging Face snapshot as a tar archive — you don’t need to download and re-upload a model by hand.
- Up to your organisation’s configured cap (3 GB by default).
- The upload resumes if your connection drops or you reload the page: select the same file again, and Earnie continues from where it left off instead of starting over.
The weights aren’t kept. Earnie deletes the uploaded file once its scan finishes, whatever the outcome. No fingerprint is retained either, so re-identifying the same model against a newer catalogue means uploading it again. A model over the size cap is refused before any upload starts, with a note that fingerprinting a model that large isn’t available yet.
Git LFS Models
A repository that keeps its models in Git LFS commits only pointers, so by default an AI scan reports each one as LFS pointer, with no way to fingerprint it. An Operator or Admin can opt a project in to fetching them: Project → Scan Configuration → AI provenance. Once enabled, a repository scan that runs the AI scanner fetches each pointed-to model (up to the same size cap as an upload) and identifies it like a committed file.Off by default. Fetching an LFS object spends the repository’s own LFS bandwidth quota, so this only happens for a project that explicitly opts in. A model that fails to fetch is left as a pointer rather than failing the scan.
Writing Policies Against AI Findings
Every finding carries anai field group once the AI scanner has run, with kind distinguishing a model finding ("model") from an AI-usage finding ("usage"): code that calls an AI service or SDK, rather than a model file. These are triaged separately too — see AI Model Findings in Triaging Findings.
The model fields most policies need:
An AI-usage finding carries its own fields instead, such as
finding.ai.vendor, finding.ai.sdk, and finding.ai.provider, which identify the AI service or library the code calls.
Because these fields are empty on findings they don’t apply to, a rule that reads them evaluates false rather than erroring on other kinds of finding.
The Sentence builder offers no negation on AI fields, because “AI SDK is not openai” would also match every open-source finding. AI Vendor is the exception: it also offers is not and is not one of, for an approved-vendors list. Keep a negated vendor beside AI Finding Kind is usage and AI Vendor is not empty, as the Enforce approved AI vendors template does, or the rule also matches every finding that names no vendor.
Starter Templates
Seven built-in templates cover the AI risks most organisations start with. Six evaluate model findings; Enforce approved AI vendors evaluates AI-usage findings instead. Add them from New policy → Pick templates, the same as any other template:Every AI template waits on the AI scanner having run. A scan that skipped it reports the policy as not evaluated, never a silent pass.
In the SBOM
An identified model is exported as a CycloneDXmachine-learning-model component: its purl, SHA-256, licence, lineage (as pedigree ancestors), and model card. There’s no separate include switch for models, unlike cryptographic assets — a model ships with the SBOM by default whenever your organisation has the AI module.