> ## Documentation Index
> Fetch the complete documentation index at: https://docs.scanoss.com/llms.txt
> Use this file to discover all available pages before exploring further.

# AI Provenance

> How Earnie identifies AI model files in a repository or uploaded directly, what it records about each one, how AI-usage findings differ, and how to write policies against both.

**AI provenance** is Earnie's third domain of risk, alongside open source and cryptography: identifying the AI your code depends on. It's available where your organisation has the **AI** module enabled, next to **Cryptography** in scanner selection.

The AI scanner produces two kinds of finding:

* **Model findings** — an AI model file in your repository or uploaded directly, with what's known about its licence, lineage, and datasets. This page is mostly about these.
* **AI-usage findings** — code that calls an AI service or SDK, rather than a model file itself. These are identified and triaged the same way as model findings (see [AI Model Findings](/en/latest/earnie/using-earnie/triaging-findings#ai-model-findings)), but this page doesn't cover them in detail.

This page explains how a model is identified, the two ways a model reaches Earnie, and how to write policies against what's found.

## Identifying a Model File

When the **AI** scanner runs, a model file committed to your repository — for example `.safetensors`, `.gguf`, `.bin`, `.pt`/`.pth` (PyTorch), `.onnx`, `.h5`/`.keras`, or a similar weights format — is identified as a model, not just skipped as a binary. Earnie reports:

* **A purl** identifying the model.
* **How it matched**, and a **confidence band** shown as text, never as a raw score.
* **The nearest known relatives** the matcher couldn't fully separate it from, when applicable.
* What the SCANOSS model catalogue records: **licence**, **lineage** (base models it was derived from), **datasets**, and the **model card**.

A model that doesn't match anything, or can't be matched, still produces a finding rather than disappearing silently. Its status explains why:

| Status | Meaning |
| - | - |
| **Identified** | Matched against the SCANOSS model catalogue. |
| **Unidentified** | Matching ran but found nothing. |
| **Matching unavailable** | Matching couldn't run this time (no SCANOSS key, the key was rejected, the service is rate-limited, or the scan ran offline). Earnie falls back to an offline guess where it can, and never fails the scan. |
| **Too large** | The file is over the size Earnie will fingerprint (see [Uploading a Model](#uploading-a-model) for the same limit on uploads). |
| **LFS pointer** | The repository keeps this model in Git LFS, and Earnie has only the pointer. See [Git LFS Models](#git-lfs-models). |

<Note>
  A model that could not be matched is still counted as scanned, not skipped. Coverage stays full, so the absence of a matched model is never silently swept away as if the file had never existed.
</Note>

## Uploading a Model

Not every model your organisation uses lives in a scanned repository. From a project's **Models** page, you can upload one directly:

* **One file**, or a **Hugging Face snapshot as a tar archive** — you don't need to download and re-upload a model by hand.
* Up to your organisation's configured cap (**3 GB by default**).
* The upload **resumes** if your connection drops or you reload the page: select the same file again, and Earnie continues from where it left off instead of starting over.

An uploaded model is scanned the same way a committed one is: identification, licence and lineage lookup, triage, policy evaluation, and inclusion in the SBOM are all unchanged. It appears on the project's Models page with the same finding detail a committed model gets, linked into the scan that identified it.

<Note>
  **The weights aren't kept.** Earnie deletes the uploaded file once its scan finishes, whatever the outcome. No fingerprint is retained either, so re-identifying the same model against a newer catalogue means uploading it again. A model over the size cap is refused before any upload starts, with a note that fingerprinting a model that large isn't available yet.
</Note>

Deleting a model from the Models page marks its finding absent and removes the stored bytes; its triage history is kept, as any resolved finding's is.

## Git LFS Models

A repository that keeps its models in Git LFS commits only pointers, so by default an AI scan reports each one as **LFS pointer**, with no way to fingerprint it.

An Operator or Admin can opt a project in to fetching them: **Project → [Scan Configuration](/en/latest/earnie/using-earnie/scan-configuration) → AI provenance**. Once enabled, a repository scan that runs the AI scanner fetches each pointed-to model (up to the same size cap as an upload) and identifies it like a committed file.

<Note>
  **Off by default.** Fetching an LFS object spends the repository's own LFS bandwidth quota, so this only happens for a project that explicitly opts in. A model that fails to fetch is left as a pointer rather than failing the scan.
</Note>

## Writing Policies Against AI Findings

Every finding carries an `ai` field group once the AI scanner has run, with `kind` distinguishing a model finding (`"model"`) from an **AI-usage finding** (`"usage"`): code that calls an AI service or SDK, rather than a model file. These are triaged separately too — see [AI Model Findings](/en/latest/earnie/using-earnie/triaging-findings#ai-model-findings) in Triaging Findings.

The model fields most policies need:

| Field | CEL Path | Meaning |
| - | - | - |
| **Kind** | `finding.ai.kind` | `"model"` for an identified or unmatched model file, distinct from AI-usage findings. |
| **Model status** | `finding.ai.model_status` | `identified`, `unidentified`, `matching_unavailable`, `too_large`, or `lfs_pointer`. |
| **Model licence** | `finding.ai.model_license` | The licence the catalogue records for the model. |
| **Model licence category** | `finding.ai.model_license_category` | How Earnie classifies that licence, for example non-commercial or use-restricted. |
| **Model publisher** | `finding.ai.model_publisher` | The organisation that published the model. |
| **Model format / transform** | `finding.ai.model_format`, `finding.ai.model_transform_kind`, `finding.ai.model_is_adapter` | The file format (for example pickle-based), and whether the model is fine-tuned, merged, distilled, or an adapter on top of another model. |
| **Match method / confidence band** | `finding.ai.match_method`, `finding.ai.confidence_band` | How the match was made, and how confident Earnie is in it. |

An AI-usage finding carries its own fields instead, such as `finding.ai.vendor`, `finding.ai.sdk`, and `finding.ai.provider`, which identify the AI service or library the code calls.

Because these fields are empty on findings they don't apply to, a rule that reads them evaluates false rather than erroring on other kinds of finding.

The Sentence builder offers no negation on AI fields, because "AI SDK is not openai" would also match every open-source finding. **AI Vendor** is the exception: it also offers **is not** and **is not one of**, for an approved-vendors list. Keep a negated vendor beside **AI Finding Kind is usage** and **AI Vendor is not empty**, as the **Enforce approved AI vendors** template does, or the rule also matches every finding that names no vendor.

### Starter Templates

Seven built-in templates cover the AI risks most organisations start with. Six evaluate model findings; **Enforce approved AI vendors** evaluates AI-usage findings instead. Add them from **New policy → Pick templates**, the same as any other template:

| Template | Action | Default |
| - | - | - |
| **Require review for unidentified models** | Require | `unidentified`, `matching_unavailable`, `lfs_pointer`, `too_large`, or a filename match |
| **Block non-commercial model licences** | Block | Non-commercial licence terms |
| **Require review for restricted model licences** | Require | Use-restricted or unknown licence terms |
| **Enforce approved AI vendors** | Require | An allowlist you set — starts empty, so the first evaluation lists every vendor in use |
| **Block unsafe model formats** | Block | Pickle-based formats |
| **Require review for restricted publishers** | Require | A publisher list you set — starts empty |
| **Warn on derived models** | Warn | Fine-tuned, adapter, merged, or distilled models |

<Note>
  Every AI template waits on the AI scanner having run. A scan that skipped it reports the policy as **not evaluated**, never a silent pass.
</Note>

A **CRA — Annex I Part II(1)** regulatory template set bundles the unidentified-model and restricted-licence templates in one action, the same way the [cryptography regulatory template sets](/en/latest/earnie/using-earnie/setting-policies#regulatory-template-sets) work. See [Setting Policies](/en/latest/earnie/using-earnie/setting-policies) for how templates, the rule builder, and template sets work in general.

## In the SBOM

An identified model is exported as a CycloneDX `machine-learning-model` component: its purl, SHA-256, licence, lineage (as pedigree ancestors), and model card. There's no separate include switch for models, unlike cryptographic assets — a model ships with the SBOM by default whenever your organisation has the AI module.

## What's Next

An AI finding is triaged, reviewed, and gated the same way an open-source or cryptography finding is. See [Triaging Findings](/en/latest/earnie/using-earnie/triaging-findings) and [Remediation](/en/latest/earnie/using-earnie/remediation) for what happens once one is found.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.