source, dependency_info, finding_id, and optional occurrence_key). The finding-centric reachability slices such as call_chains, entry_call, and crypto_call are emitted by the dedicated --export-callgraph artifact rather than embedded in the interim report.
Overview
When--scan-dependencies is enabled, crypto-finder goes beyond scanning the user’s source code. It resolves the project’s dependency tree, scans each dependency for cryptographic usage, builds a cross-package call graph, and traces each finding back to the user’s code to answer: “Does my code actually reach this crypto function?”
The Six-Step Pipeline
The full pipeline lives inDependencyScanner.ScanWithDependencies(). Here’s what each step does and why.
Step 1: Resolve Dependencies
TheResolver interface discovers all dependencies and locates their source code on disk. Each ecosystem has its own resolver implementation (see Supported Ecosystems below). Each Dependency carries:
The
RootModule (e.g. github.com/myorg/app for Go, com.myorg for Java) is the default user-code prefix. For Java, packages of functions whose files live in the scanned project tree are also user code. That matters for Gradle projects that omit group and therefore export a project name rather than a Java package prefix.
Step 2: Load & Filter Rules
Rules are pre-loaded once and filtered to the ecosystem’s language(s). For a Go project, onlygo rules are kept; for Java, only java rules. This avoids running irrelevant rules against source code, significantly reducing scanner overhead.
The kept rules are then written to one merged file, which every dependency scan reuses. OpenGrep loads each config file separately on every run, so one file loads faster than hundreds. The merged file is named merged-rules.yaml but holds JSON text, one rule per line: JSON is valid YAML, and OpenGrep parses it about seven times faster. Each merged rule ID carries its source file’s directory, the prefix OpenGrep derived from the file location before, so finding rule IDs do not change. Files the merge cannot represent exactly are passed to the scanner unchanged. These are invalid YAML, *.test.yaml fixtures, hidden paths, and YAML with no JSON form of the same meaning, such as a number written 0x10 or an alias.
Step 3: Scan Dependencies in Parallel
Dependencies are deduplicated bymodule@version and processed in a stable order (module, version, dir) so repeated scans produce deterministic report and call graph inputs. The findings cache is consulted per dependency; the misses are then grouped into batches and each worker runs one scanner process per batch, over every root in it, instead of one process per dependency. OpenGrep loads the rules once per process, and that load is the dominant cost of scanning a small dependency (about 20 s for the JavaScript and TypeScript rules), so a 50-package npm project pays it a handful of times instead of 50. Each process gets --jobs sized to its share of the cores (see the scannerJobs log field).
With an API key, published findings stand in for a dependency’s detection. The SCANOSS mining service scans each package version once and publishes its findings. When an API key is configured, a dependency the findings cache does not hold is looked up by versionless purl and version (POST /v3/cryptography/reachability/component, purl and requirement only). The published findings are taken as they are, including the closest mined patch the API serves for that version and the rules version it was mined with, and are kept to the dependency’s own files. The dependency is then not scanned. A dependency the API has not mined, a lookup error, a refused key, or an API that does not offer the endpoint (HTTP 404) scans the dependency locally; a refused key or a missing endpoint stops further lookups for that scan, with one warning. Published findings are not written to the findings cache. The call graph, conditioned findings and reachability are always built from the local sources, so findings derived from calls between dependencies are kept. Each lookup sends the package’s purl and version to the API. --no-dependency-findings-api scans every dependency locally (internal/engine/dependency_findings_source.go).
Each root’s report is what a scan of that root alone produces. The OpenGrep adapter (ScanRoots, internal/scanner/opengrep/batch.go) attributes every result and error to the root holding its file (longest prefix) and runs the same transformation per root, with that root as the target, so paths, deduplication, rule IDs and finding IDs are unchanged. An error without a file, such as a memory limit reported for the whole process, marks every root in the batch incomplete. The dependency scanner then stamps the ruleset and enriches each report as it does for a single scan, and writes each dependency to the findings cache under its own key, still skipping a dependency whose scan stopped at a limit.
Batch shaping (internal/engine/dependency_batches.go):
- Nested roots never share a process. An npm dependency’s scan excludes its own
node_modules/, where another dependency may be installed, and an exclusion applies to the whole process. Roots are assigned a nesting level (0 when no other root holds them, otherwise one more than the root that does; a repeated directory counts as nested) and only roots of the same level are batched together. A scoped root that names none of another root’s files does not hold it (scanner.Disjoint): Python namespace siblings, and a Python distribution rooted atsite-packagesbeside the distributions installed below it, share a level. Scoped roots carry no exclusion anchored at their own directory, so sharing a process hides none of the other root’s files. - Balanced by weight. Each dependency is weighed by the bytes of its source files for the ecosystem’s languages (a cheap walk that skips nested
node_modules). Within a level, the misses are split intomax(ceil(n / 16), min(workers, n))batches: heaviest dependency first, each into the lightest batch, so the workers finish together and a very large dependency keeps a process to itself. Batches are queued heaviest first. - Process timeout. A batch gets the single-scan
--timeoutonce perceil(roots / jobs): OpenGrep analyzesjobsfiles at a time, so the roots consume about that many single-scan budgets of wall time. With the default 10 minutes and 4 jobs, a 16-root batch may run for 40 minutes before it is treated as failed. - Failure isolation. When a batch’s process fails (exit status above 1, unreadable output, timeout), every dependency in it is scanned alone, as before batching, so one faulty dependency fails only itself and the others keep their findings and their cache entries. A canceled scan fails the batch’s dependencies and rescans nothing.
- A batch of one root, which includes every dependency when the scanner cannot batch (the Semgrep adapter), goes through the single-scan path unchanged.
NpmResolver (internal/dependency/npm_resolver.go) drops the lockfile entries marked "dev": true (in a v1 lockfile, a dev entry together with the entries nested under it), so an installed development-only package is neither scanned nor put in the dependency graph. devOptional packages stay, because npm installs them without dev when an optional production dependency needs them, and workspace members stay because they are the project’s own code. --include-dev-dependencies (WithNpmDevDependencies) restores the development-only packages. Absent dev packages, as after npm install --omit=dev, are never an error either way.
Go dependencies scan only the files their imported packages compile. A Go program contains only the packages it imports, and of each package only the files its build compiles, so the Go resolver lists, for each package in the production import closure, the Go files go list -e -deps reports in GoFiles and CgoFiles (Files), and the scan reads only those files: not the module’s other packages, not the subdirectories of an imported package, which are packages of their own, and not the files that build constraints exclude. Build constraints apply file by file, for the scanning host (see Host build configuration under Go-specific): a file for another GOOS or GOARCH (*_windows.go, //go:build darwin), one whose build tag is unset, a //go:build ignore generator, and with CGO_ENABLED=0 a file that imports "C" are not scanned. Only Go files are listed, because a Go dependency scan runs only Go rules and the Go call graph parses only Go files: a package’s C, assembly and other files, and its _test.go files, never yielded a finding. modernc.org/libc, for example, ships one 4 MB ccgo_<os>_<arch>.go per platform, and the host build compiles one of them. The dependency’s root stays the module directory, so attribution, result paths, finding IDs and the batch nesting levels are those of a whole-module scan; only the targets change. The scanner receives the files as named targets (a DetectionScope, the same mechanism as --detect-paths-from, with --force-exclude so the dependency’s skip patterns still apply to them). Files are leaves, so the packages of one module, such as foo/ and foo/bar/, never hide one another, and a batch over many files is split into several processes when the file list exceeds the command-line budget. OpenGrep keeps its language and size filters for named files, so a file in an imported package yields the findings it yielded in a whole-module scan. The findings cache key includes the scoped file list, so a whole-module entry is never served for a package-scoped scan, or one set of packages for another. A scanner that cannot limit detection to files (the Semgrep adapter) still scans the whole module, under the whole-module key, and the dependency keeps only the findings in its files.
Python namespace distributions scan only their own files. Several distributions can install into one namespace directory: protobuf, googleapis-common-protos, google-auth and google-api-core all resolve to site-packages/google. When two or more distributions resolve to the same Dir, the pip resolver gives each one the files under it that its *.dist-info/RECORD lists and that exist (Files). A distribution without a RECORD owns the files no sibling’s RECORD lists; top_level.txt names import roots, not files, so it can choose Dir but cannot split it. A distribution with a directory of its own keeps Files unset and scans its whole Dir as before. The shared directory stays the root, so a finding’s path is relative to it, as in a whole-directory scan (auth/_helpers.py for google-auth), and the files go to the scanner as named targets through the same DetectionScope as Go’s packages. Siblings whose files are disjoint share a batch: the scanner attributes each result by the exact file named in a scope, so they are not nesting levels of one another. Siblings whose files overlap (two RECORDs listing one file, or two distributions without a RECORD) are still scanned in separate processes. Each file is therefore scanned once and reported under the one distribution that installed it. Previously every sibling scanned the whole namespace in its own process and reported every finding in it. A scanner that cannot limit detection to files (the Semgrep adapter) still scans the whole directory for each sibling, and each sibling keeps only the findings in its own files.
Python distributions scan every top-level package and module they installed. packages_distributions() maps import names to distributions, so one distribution can have several: configobj installs configobj/ and validate/, PyYAML installs yaml/ and _yaml/, setuptools installs setuptools/, pkg_resources/ and _distutils_hack/. The resolver takes each name that is a directory in site-packages, or a module file name.py there (six.py), drops a name inside another one (googleapiclient/discovery_cache) names that are neither (compiled extensions), and __pycache__, which importlib infers from the bytecode of module files and which all module files share, and sorts the rest. A distribution with exactly one package directory keeps it as Dir and ImportPath, so its paths, finding IDs and cache keys are unchanged. Any other distribution is rooted at site-packages, the one directory that holds all its roots, with an empty ImportPath and Files listing every file under each of its package directories and each module file; its share of a namespace directory it also installs into is the files its RECORD lists, as above. Its finding paths therefore start with the package (validate/__init__.py, six.py). Previously the resolver kept whichever name it met first in a Go map, a different one from run to run: configobj scanned configobj/ in one run and validate/ in the next, setuptools one of its three packages, and a distribution whose first name was a module file, such as six, was skipped as a single-file module. Findings and call-graph functions of these distributions changed between runs of the same binary on the same environment; they no longer do. The scope of a distribution at site-packages names none of the other distributions’ files, so it batches with them (see above) and never reports their findings.
Dependency scans use their own exclusions. buildDepScanOptions (internal/engine/dependency_scanner.go) replaces the primary scan’s skip patterns with skip.OnlyDefaultTestPatterns, which keeps only the built-in test patterns (DefaultSkippedTestPatterns in internal/skip/source_defaults.go: test/, tests/, src/test/, src/tests/, __tests__/, **/*Test.java, **/*Tests.java, **/*_test.go, **/test_*.py). For npm it adds an anchor on the dependency’s own node_modules/, because the dependency directory usually sits inside a node_modules tree itself and each child package is scanned under its own identity. It also sets IncludeGitIgnored, which passes the OpenGrep flag that disables .gitignore handling. As a result, none of these apply inside a dependency root:
--excludeandsettings.skip.patterns.scanningfromscanoss.json.--excludeis not a way to hide dependency findings.- The built-in skipped directories (
dist,build,vendor,target,docs,shaded, …) and the generated-stub globs (*.pb.go,*_pb2.py, …).--no-default-exclusionshas nothing to remove here. .gitignorefiles.
--include-tests leaves the test patterns out of the list, so a dependency scan then reads its test files too. The reason for the narrow list: a published package is not a checkout, and often ships its only code in dist/ or build/, which the source-tree defaults would hide. Go dependencies narrow the scan further to the Go files each imported package compiles for the host build (see above), which never include _test.go files.
Dependencies without a usable local source directory are not sent to the scanner. They are logged as Skipping dependency source scan: no local source directory instead of triggering empty-path scanner failures. For Java, those dependencies still proceed to step 4 as type-only inputs as long as module@version can be resolved to a compiled JAR.
Step 4: Build the Call Graph
This is where the architecture gets interesting. The call graph builder uses syntactic parsing to process source files, which means it works on raw source without needing a full Go toolchain, Java compiler, Python interpreter, or Rust toolchain. TheParser interface abstracts all language-specific behavior:
NewParserForEcosystem() factory selects the right parser (GoParser, JavaParser, PythonParser, or RustParser) based on the detected ecosystem.
Performance optimization: Only dependencies with crypto findings get full source parsing. Dependencies without findings contribute only their bytecode type signatures (class names, method signatures, return types, interface hierarchy). This preserves 100% type resolution accuracy for fluent chains while skipping expensive source parsing for ~80% of dependencies.A dependency’s
Files becomes the IncludeFiles of its PackageDir. For Go, the builder parses only the files the host build compiles in the imported packages (the Go parser skips an unlisted file before reading it, so modernc.org/libc’s seven other ccgo_linux_<arch>.go files cost nothing) and still walks through their parent directories so every package keeps the import path a whole-module walk gives it. This loses no reachability. A Go function can call only functions of packages its own package imports, which are in the closure too, and an interface value can only hold a type of a package the program contains, so the module’s other packages hold no call edge and no dispatch target of the program. Likewise a file that build constraints exclude is not part of the program the host builds.
For a Python namespace distribution, the builder parses only its own files of the shared namespace directory, under the import path a whole walk gives them (google.auth._helpers), and the Python dependency type resolver indexes only them. Each function of a namespace directory is parsed once and belongs to the distribution that installed it, so containing-function lookups, synthesized entry-point findings and the call-graph export name that distribution. A distribution rooted at site-packages parses under the empty import path, so each package and module keeps its own name (configobj, validate.check, six), and both the builder and the type resolver walk only the directories that hold its files.
What the parser extracts
Each language parser extracts the same semantic information into the sharedFileAnalysis / FunctionDecl structures. For Go (.go files, excluding _test.go):
For Java (
.java files):
For Python (
.py files, excluding test_*.py and *_test.py):
For Rust (
.rs files, excluding *_test.rs and tests.rs):
The two data structures
TheCallGraph holds two maps:
FunctionsmapsFunctionID.String()→*FunctionDecl(forward: function → its outgoing calls)Callersmaps callee →[]callerID(reverse: who calls this function?)
Edge resolution kinds
While building the caller index, the builder records how each edge was resolved inCallGraph.EdgeResolutions: exact (receiver type known, unique
target or overload set on that type), interface_dispatch (an interface/abstract
method expanded to concrete implementations by name+arity within a namespace
root), or name_only (a fluent-fallback guess with no receiver-type anchor).
This classification is emitted on the graph-fragment export (graph-fragment-1.3,
see Output Formats) so downstream
consumers can fail closed on over-broad dispatch instead of reporting it as
typed reachability. The dispatch heuristics below (interface expansion, fluent
fallback) are exactly the edges tagged non-exact. The same fragment also
exposes crypto_entry_points[] and supporting_calls[] so consumers can use a
stable API index rather than the removed entry_point_index projection.
Type resolution
After building the caller index, the builder runs additional resolution passes to improve type accuracy:-
TypeResolver(language-specific): For Java, a bytecode-based resolver reads.classfiles from resolver-supplied dependency JARs (Maven or Gradle) plus the selected JDK platform archives to extract fully-qualified method signatures. JAR indexing runs in parallel and uses a per-artifact bytecode cache under~/.scanoss/crypto-finder/cache/bytecode/, keyed by exact artifact identity. This provides accurate parameter types (e.g.,io.jsonwebtoken.SignatureAlgorithminstead of genericK) and return types for fluent chain resolution. The Java resolver can be configured per scan withjava_jdk_major(auto,8,11,17,21) andjava_jdk_homes, which also makes Java dependency resolution JDK-aware. TheTypeResolverinterface is extensible — each language can implement its own approach (Go:go/types, Python:.pyistubs, Rust:rust-analyzer). -
Fluent chain resolution: For chained calls like
Jwts.builder().setId(id).signWith(algo, key), return types are propagated through the chain. Ifbuilder()returnsJwtBuilder, thensetId()is resolved asJwtBuilder.setId. Interface inheritance is also followed (e.g.,JwtBuilderextendsClaimsMutator, sosetIdresolves toClaimsMutator.setId). -
Argument source tracing: For each function call, the parser traces where argument values come from — literal constants, local variables, class fields, method parameters, or call results. This produces recursive
source_nodesshowing the data flow into each argument.
Step 5: Trace & Attribute Findings
This is the core of the attribution system. For each crypto finding in a dependency:How TraceBack works (BFS)
- Start with the target function (where the crypto finding was detected)
- Look up all callers via the reverse index (
graph.Callers[targetKey]) - Prepend each caller to the chain being built (so the chain grows backward:
[caller, ...existing]) - Terminate when a root function is reached (a function with no callers, e.g.,
main) - Validate that the complete chain passes through at least one user-package function
- Cycle detection prevents infinite loops in recursive call graphs
- Return all complete chains (BFS finds all paths from entry points to the crypto call site)
[program_entry_point, ..., intermediate, ..., crypto_call_site] — array position [i] calls position [i+1].
Attribution output
After tracing, each finding gets structured attribution metadata in the interim report, and the detailed reachability slices are emitted by the separate call graph export. Dependency finding:file_path values are relative to the dependency root. The artifact identity stays in dependency_info, so consumers do not need to parse module@version back out of the path string.
User code finding:
Step 6: Merge Reports
The current implementation merges dependency findings into the interim report and defers reachability slicing to the call graph export. In other words, interim-report inclusion is not gated on whether a finding later produces one or more exported call chains. User code findings are always included and marked withsource: "direct".
Practical Walkthrough
This section traces the entire pipeline using the test project attestdata/projects/go_with_crypto_dep/. Every value shown is real — produced by running crypto-finder scan --scan-dependencies against this project.
The Source Code
Three files make up the project:go.mod — declares one direct dependency:
main.go — the entry point. Does not use crypto directly:
mypkg/crypto.go — user’s wrapper package. Uses crypto from a dependency:
Step 1: Scan User Code (the normal scan)
The orchestrator runs Semgrep/OpenGrep rules against user code. It finds 3 assets inmypkg/crypto.go:
This produces the user report. At this point there are no
source, call_chain, or dependency_info fields — just raw findings.
Step 2: Resolve Dependencies
The Go resolver lists the main module withgo list -m -json, then runs go list -e -deps over the main module’s packages and keeps every module that provides a package in that non-test import closure. Build constraints are evaluated for the scanning host: on amd64, chacha20poly1305 imports golang.org/x/sys/cpu, so both modules below are in the closure, while on arm64 only golang.org/x/crypto is. On amd64 it returns:
RootModule (example.com/crypto-test) is the Go user-code prefix — any package whose import path starts with it is user code. Java additionally treats packages declared in project source files as user code, even when RootModule is a Gradle project name rather than a package prefix. Everything else is a dependency.
Step 3: Scan Dependencies in Parallel
Each dependency gets scanned with the same rules, limited to Go rules only, and only in the packages the program imports:chacha20poly1305, chacha20, internal/alias and internal/poly1305 of golang.org/x/crypto, and cpu of golang.org/x/sys. Measured on 2026-10-02 with the current rule set:
Total: 4 dependency findings. Only
golang.org/x/crypto has findings, so it proceeds to step 4 together with the user code.
Step 4: Build the Call Graph
The builder receives the user code and the crypto-bearing dependency, limited to its imported packages:.go file and extracts function declarations with their calls.
From main.go:
mypkg/crypto.go:
golang.org/x/crypto/...: hundreds more function declarations.
Then buildCallerIndex() creates the reverse index (callee → who calls it):
Step 5: Trace & Attribute
The system now traces each finding back through the call graph to user code. Three different scenarios play out:Scenario A: User finding — chacha20poly1305.New at line 13
This finding is in mypkg/crypto.go (user code). The enrichment flow:
5a-1. Find the containing function:
FindContainingFunction("mypkg/crypto.go", 13) iterates all FunctionDecls, looking for one whose FilePath matches and whose StartLine..EndLine spans line 13. It finds SecureEncrypt (lines 12–25).
5a-2. Trace back to entry point:
TraceBack(SecureEncrypt, userPackages={"example.com/crypto-test"}, maxDepth=0):
main’s line is 14 (the line where main calls SecureEncrypt), not line 10 where main is declared. This is because findCallLine() searches main’s Calls list for the specific call to SecureEncrypt and returns that call-site line number.
Scenario B: User finding — chacha20poly1305.New at line 29
Same logic, but traces through SecureDecrypt:
main’s line is now 19 — the line where main calls SecureDecrypt, not 14.
Scenario C: Dependency finding — deep x/crypto internal functions (UNREACHABLE)
Take any internal function in golang.org/x/crypto, say ssh.newAESCTR:
call_chains is empty.
In the current implementation, this means the finding may still exist in the interim report, but it will not contribute a useful reachability slice to the call graph export.
This is why the interesting downstream signal stays narrow even when a dependency contains a large amount of internal crypto usage. The vast majority of golang.org/x/crypto’s internal crypto usage is not reachable from the user’s main().
Step 6: Merge
chacha20poly1305.New directly inside mypkg/crypto.go — a file in the user’s own module. The crypto usage is already captured as source: "direct". There’s no intermediate dependency wrapper that the user calls which then reaches crypto.
If the project had a longer chain — e.g. main → mypkg.Encrypt → someMiddleware.Process → chacha20poly1305.New where someMiddleware is a dependency — then we’d expect a dependency-backed reachability slice to appear in the call graph export.
Actual Interim Report Output
The final interim report looks like this:finding_id.
Visual Summary
Walkthrough 2: Multi-Hop Dependency Chain
The first walkthrough showed a case where user code calls crypto directly, so the useful reachability slices are anchored to direct findings. This second walkthrough usestestdata/projects/go_with_dep_chain/ to demonstrate a multi-hop chain where crypto usage is buried inside a dependency and dependency-backed reachability slices become the interesting artifact.
The Source Code
The key difference: user code never toucheschacha20poly1305 directly. Instead, it calls through a wrapper dependency (cryptowrapper_dep/), which has an internal function layer.
main.go — entry point, calls mypkg:
mypkg/crypto.go — user code, delegates to the dependency. No crypto imports.
../cryptowrapper_dep/wrapper.go — the dependency (separate module). Has a public API and an internal function:
What the Scan Produces
Running with--scan-dependencies:
x/crypto and x/sys internals may still be scanned and identified, but they do not help downstream stitching unless the graph can connect them back to component-owned or user-owned entry points.
Tracing the 3-Step Chain
The finding atchacha20poly1305.New (line 59 of wrapper.go) produces a 3-step chain. Here’s the BFS trace:
call_chains:
Full Trace to main
The BFS walks all the way to root functions (functions with no callers, like main). This means the full chain main → SecureEncrypt → Encrypt → newAEAD is preserved. A chain is valid if it passes through at least one user-package function, so chains that only traverse dependency code are discarded.
Visual Summary
- User code has 0 crypto findings —
mypkghas no crypto imports - 2 dependency findings survive reachability because the call graph proves user code reaches them
- 492 dependency findings dropped — deep
x/cryptointernals unreachable from user code mainappears at the head of each chain — BFS walks to root functions (no callers)- Both Encrypt and Decrypt paths preserved — all chains stored in
call_chains
Interim Report Contract (v1.7)
Version 1.7 keeps the attribution fields needed to join findings to the separate reachability export and adds an optional AST-anchored structural identity when callgraph evidence is available. Dependency-backed paths are dependency-root-relative;dependency_info remains the canonical place for dependency module, version, and package URL. Direct findings may additionally expose a valid rule package URL at the asset-level purl; it is version-enriched only when one unambiguous direct dependency match exists.
Call Graph Export
When--export-callgraph is enabled, Crypto Finder emits a finding-centric JSON export that uses the same relative-path convention as the main report.
Schema note: call graph export version 6.12 is current and carries optional direct finding purl values at the finding level, optional occurrence_key structural identities, canonical dependency package URLs inside dependency_info.purl, and contract-scoped supporting_calls[].supporting_call.resolved_key_length evidence for structurally derived Java key-generation configuration calls, including the rule_declared_bits value and rule_conflict marker when a rule-declared key length disagrees with the resolved one. Java runtime provenance remains available in scan_metadata for JDK-aware platform signature enrichment.
- Each top-level record stays keyed by
finding_id, which is the join key back to the interim report.occurrence_key, when present, is a separate structural identity and does not replace that join. call_chainsis the primary value-flow structure. Each chain is ordered from the first reachable caller to the function that contains the matched crypto call.- Each chain node contains a fully qualified
function_name, a normalizedfile_path,start_line, optionaldependency_info, and optionalentry_call. entry_calldescribes how execution entered the current function from the previous step. Itsfile_pathandlineare the call-site location in the previous node’s source file.- The last node in a chain carries
crypto_call, which is the matched crypto-relevant call that triggered the finding. entry_call.parameters[]andcrypto_call.parameters[]both exportparameter_index(always0-based), best-efforttype,argument_expression,resolved_value,variable_namefor simple identifiers only, and recursivesource_nodes.- For Java scans,
scan_metadatamay also includejava_requested_jdk_major,java_runtime_version,java_platform_signatures_used,java_platform_signature_source, andjava_platform_signature_unavailable_reasonto show which JDK major was requested and whether JDK platform signatures were available for enrichment. source_nodescan now carry interprocedural provenance across wrapper hops, for examplePARAMETER -> PARAMETER -> VALUE, and propagated nested nodes keeplocation.file_pathpluslocation.linewhen known.- Method-call expressions are preserved as
CALL_RESULTnodes instead of flattening away their receivers. When the invoked method can be resolved, the node also exportscall_target, and receiver provenance stays nested under theCALL_RESULT(for exampleCALL_RESULT -> PARAMETER alg -> VALUE SignatureAlgorithm.HS256). - Findings that cannot be resolved to a containing function or a specific crypto call remain in the export with
finding_locationandunresolved_reason.
Call Chains Ordering
Thecall_chains field in the call graph export contains all traced paths from program entry points to the crypto call site. Each inner array is one complete path, ordered from program entry point (index 0) to crypto call site (last index). Entry [i] calls entry [i+1].
Example:
Findings Cache
Dependency scanning is dominated by opengrep execution time (~93% of pipeline time). Sincemodule@version produces identical scan results with the same ruleset, caching eliminates redundant work entirely. On a second scan with the same dependencies and rules, the dependency scanning phase drops from minutes to near-zero.
How It Works
The cache sits between Step 2 (rule loading) and Step 3 (parallel scanning) in the pipeline. Before scanning,lookupDependency checks each dependency for a cached result. The misses are scanned in batches (Step 3) and each dependency’s report is stored under its own key, unless its scan stopped at a time or memory limit.
Cache Key Design
The key captures everything that affects scan output:- Module + version: e.g.,
org.bouncycastle:bcprov-jdk18on@1.78 - Rules hash: First 16 hex chars of SHA-256 over sorted rule file contents — if any rule is edited, the cache invalidates automatically
org.bouncycastle:bcprov-jdk18on@1.78:a3f8b2c1d4e5f678
The rulesHash is computed once per scan (not per-dep), so I/O cost is negligible.
Storage Layout
The default implementation (DiskFindingsCache) stores results as JSON files:
golang.org/x/crypto) are replaced with _ for filesystem safety. Writes use temp file + atomic rename to prevent corruption from interrupted scans.
FindingsCache Interface
context.Context on both methods to support network-backed implementations with timeouts and cancellation. The pipeline doesn’t know or care which backend is behind the interface.
Distributed Extensibility
TheFindingsCache interface is the extension point for multi-node scanning:
Each just implements
Get/Put. The scanning pipeline is completely agnostic about the storage backend.
Architecture Map
Supported Ecosystems
The extensible architecture makes adding a new language a matter of implementing two interfaces and registering them. Currently supported:Go
- Resolver:
GoResolver— usesgo list -m -jsonfor the main modules andgo list -e -depsfor the modules in their production import closure and the Go files their imported packages compile for the host build (GoFiles,CgoFiles), the only ones scanned and parsed - Parser:
GoParser— syntactic parsing of Go source - Manifest:
go.mod - Module format: Go import path (e.g.,
golang.org/x/crypto) - Package separator:
/ - Source location: Go module cache (
$GOPATH/pkg/mod/)
Java (Maven / Gradle)
- Resolver:
JavaResolver— auto-detects Maven vs Gradle at the project root - Parser:
JavaParser— syntactic parsing of Java source - Manifest:
pom.xml,build.gradle,build.gradle.kts,settings.gradle,settings.gradle.kts - Module format:
groupId:artifactId(e.g.,org.bouncycastle:bcprov-jdk18on) - Package separator:
. - Source location: Source JARs resolved by the active build tool and extracted to
~/.scanoss/crypto-finder/cache/sources/
Maven Resolution Details
TheMavenResolver uses a three-tier fallback strategy to maximize dependency recovery, especially for multi-module projects:
Every Maven invocation receives -Dmaven.repo.local=<user-home>/.m2/repository, so Maven writes artifacts to the same repository that Crypto Finder reads even when the JVM reports a different user.home. The operating-system user home is HOME on Unix and USERPROFILE on Windows.
Tier 1 — Reactor with --fail-never (always attempted):
- Runs
mvn dependency:list --fail-never -DappendOutput=true -DincludeScope=compile - The
--fail-neverflag continues past module failures;-DappendOutput=trueaccumulates results from all succeeding modules into a single output file - If some modules resolve successfully, their dependencies are collected even if other modules fail
- Detects modules from
<modules>in the parentpom.xml - Runs
mvn dependency:list -pl <module>for each module independently - Modules that fail are skipped; dependencies from succeeding modules are deduplicated and collected
- Runs
mvn install -DskipTests --fail-neverto build all modules locally, populating~/.m2/repositorywith inter-module artifacts - Retries Tier 1 after install
- This is expensive (requires compilation) but is the only way to resolve inter-module transitive dependencies
mvn dependency:sources— downloads-sources.jarfiles to~/.m2/repository/(best-effort; ~65% of Java libraries publish source JARs)mvn dependency:tree --fail-never -DappendOutput=true— builds the dependency graph adjacency list (best-effort)
~/.m2/repository, Java bytecode indexing can still use it as a type-only dependency.
Gradle Resolution Details
TheGradleResolver asks Gradle itself for a machine-readable dependency model via a temporary init script:
- Prefers
./gradlewand falls back togradlefromPATH - Uses the caller’s
GRADLE_USER_HOMEwhen set; otherwise defaults it to<user-home>/.gradleusing the same operating-system user home - Resolves the main Java compile classpath for single-project and multi-project builds
- Treats included Gradle subprojects as
WorkspaceMembersrather than external dependencies - Captures external module coordinates, versioned dependency edges, compiled JAR paths, and best-effort source archive paths
- Reuses the shared source extraction cache so Gradle and Maven dependencies flow through the same Java scanning pipeline
- Reachability still traces from application Java packages when the project has no
groupandRootModuleis the Gradle project name
Multi-Module Project Support
Multi-module Maven projects (parent POM with<modules>) are automatically detected. When detected:
- All modules are registered as
WorkspaceMembers, meaning they are treated as user code for call chain tracing (same as Cargo workspace members) - The
WorkspaceMember.Namefollows the formatgroupId:moduleDirName - The three-tier fallback strategy handles common multi-module failures:
- Inter-module dependencies (e.g.,
eladmin-loggingdepends oneladmin-common) — resolved via Tier 3 - HTTP mirror blocks (Maven 3.8.1+ blocks insecure HTTP repositories) — partial results collected via Tier 1
- Missing parent POMs or private repositories — gracefully degraded via Tier 1/2
- Inter-module dependencies (e.g.,
Java Call Resolution
TheJavaParser resolves method calls through import analysis:
- Explicit imports:
import javax.crypto.Cipher;→Cipher.getInstance(...)resolves to packagejavax.crypto - Wildcard imports:
import java.security.*;→ class names matched against wildcard packages - Local variable types:
Cipher c = Cipher.getInstance(...)→c.doFinal()resolvescto typeCiphervia local variable tracking - Field types: Class fields are tracked similarly to local variables
- Fallback: Unresolved calls default to the current package (same as Go’s behavior for unresolved variables)
Python (pip)
- Resolver:
PipResolver— usespython -m pip list --format=json+python -m pip showto resolve packages with the same interpreter used for metadata lookups - Parser:
PythonParser— syntactic parsing of Python source - Manifests:
pyproject.toml,requirements.txt,Pipfile,setup.py - Module format: Python package name (e.g.,
cryptography) - Package separator:
. - Source location: Site-packages directory (e.g.,
~/.local/lib/python3.x/site-packages/)
Python Resolution Details
ThePipResolver executes the following steps:
- Root module detection — reads
pyproject.tomlfor[project] nameor[tool.poetry] name; with neither the root module is empty and the project’s own symbols are keyed at the scan root, never by the directory name python -m pip list --format=json— lists all installed packages with versionspython -m pip show <packages>— gets location and dependency info for each package (batched in groups of 50)- Distribution-to-import mapping — uses that SAME interpreter’s
importlib.metadata.packages_distributions()(Python 3.10+) to map distribution names to import names. Falls back to scanning*.dist-infodirectories (top_level.txt→RECORDfile) for older Python versions - Package root resolution — uses the import mapping, then heuristic name normalization, to find every top-level package directory and module file of the distribution (see “Python distributions scan every top-level package and module they installed”). C-extension packages are skipped
Python Call Resolution
ThePythonParser resolves calls through import analysis:
import X:X.func()resolvesXvia importsfrom X import Y:Y()resolves to packageX, treated as constructorY.<init>()- Chained attributes:
a.b.c.func()— first segment resolved via imports, rest chained selfcalls:self.method()resolves to the current package- Wildcard imports:
from X import *recorded for fallback resolution - Aliased imports:
import X as Y—Ymaps toX - Fallback: Unresolved calls default to the current package
Rust (Cargo)
- Resolver:
CargoResolver— usescargo metadata --format-version=1 - Parser:
RustParser— syntactic parsing of Rust source - Manifest:
Cargo.toml - Module format: Crate name (e.g.,
ring) - Package separator:
:: - Source location: Cargo registry cache (e.g.,
~/.cargo/registry/src/.../<crate>-<version>/)
Rust Resolution Details
TheCargoResolver runs cargo metadata --format-version=1 which provides:
- All packages with name, version, and manifest path
- Resolve graph with dependency edges between packages
- Workspace detection — packages with
source: null(local/workspace crates) are treated as user code; all others are dependencies
Rust Call Resolution
TheRustParser resolves calls through use declaration analysis:
- Scoped identifiers:
Aead::new(...)— resolvesAeadthroughuseimports - Qualified paths:
ring::aead::new(...)— first segment resolved via imports - Scoped use lists:
use ring::aead::{Aead, AeadCore}— each item registered separately - Wildcard use:
use ring::aead::*recorded for fallback resolution selfcalls:self.method()resolves to the current modulesrc/transparency: Thesrc/directory is transparent in module paths (e.g.,ring/src/aead/→ring::aead, notring::src::aead)- Impl blocks: Methods in
impl Type { fn method() {} }are extracted with their type association - Fallback: Unresolved calls default to the current module
Adding a New Language
To add support for a new ecosystem:- Implement
callgraph.Parser— with syntactic parsing for the target language - Implement
dependency.Resolver— shells out to the ecosystem’s package manager - Register the parser in
parser_registry.go— add onecase - Register the resolver in
scan.go— add onedepRegistry.Register()call - Add manifest detection in
detectEcosystem()— add oneifchecking for the manifest file
builder.go, tracer.go, dependency_scanner.go, entities, or schemas.
Performance
Two-Phase Call Graph Build
The call graph build is the most expensive step in the dependency scanning pipeline. To minimize cost while preserving 100% type resolution accuracy, the builder uses a two-phase approach: Phase 1 — Source parsing (targeted): Only dependencies with crypto findings + user code modules get full source parsing viaParser.ParseDirectory(). This builds FunctionDecl entries with call sites, parameters, and return types.
Phase 2 — Bytecode type indexing (comprehensive): ALL dependencies (including those without findings) are indexed via JavaBytecodeTypeResolver. This reads .class files from Maven JARs to extract class names, method signatures, return types, and interface hierarchy. The type index is used to resolve fluent chains and enrich parameter types across dependency boundaries.
Why both phases are needed: Java fluent APIs (e.g., Jwts.builder().signWith(key)) require knowing return types from one dependency to resolve calls in another. A dependency without crypto findings may define the return type that bridges a call chain from user code to a crypto finding. Skipping its type information would break backward tracing.
Benchmarks (eladmin — 160 deps, 27 with findings, 269 crypto assets)
Current warm-run numbers with findings cache + Java bytecode cache enabled:
The largest recent improvements came from three changes:
- Two-phase call graph build: only findings-bearing dependencies get full source parsing
- Parallel Java JAR indexing: exact-version JARs are indexed concurrently
- Per-artifact bytecode cache: repeated scans avoid reparsing unchanged JARs
Current Bottlenecks
The pipeline has three main time consumers:-
Source parsing for graph packages (~23s on warm
eladmin): Parses.javafiles from 32 packages (user code + 27 deps with findings) to build 168K function declarations with call sites. -
Dependency scanning with opengrep: Still dominates cold scans. On warm scans most dependencies hit the findings cache; on first scans the cost depends on dependency source size and worker count (
--dep-workers). Batching removed the per-dependency rule load, so the remaining cost is the analysis itself. - Call graph post-processing (~2-3s): Caller index construction, bytecode merge/rewrite, and fluent-chain resolution are no longer dominant but still scale with graph size.
Future Optimization Opportunities
Source parsing and graph size reduction
The bytecode resolution bottleneck has largely been removed. Remaining performance opportunities are now upstream:- Reduce graph packages further: Keep shrinking the set of dependencies that require full source parsing without breaking attribution accuracy.
- Smarter pre-scan eligibility: Skip dependencies that have source directories but no scannable files before invoking opengrep.
- Containing-function lookup index:
Tracer.FindContainingFunction()still scans functions linearly; indexing by file could cut repeated lookups on large reports. - Selective bytecode indexing: Only index JARs whose types appear in unresolved calls from graph packages. More complex, but now one of the few remaining bytecode-side wins.
Opengrep scanning
- Result caching: The existing
DiskFindingsCachecaches dependency scan results bymodule@version:rulesHash. On repeated scans of the same project, most dependencies hit the cache. - Cold-scan throughput: First-scan performance still depends heavily on opengrep throughput over large dependency sources. Scanning only the directories a Go module actually imports, rather than the whole module, is the next lever there.
Call graph export
- Export is fast (~10-15s for 269 findings) and not currently a bottleneck. The finding-centric export only traces paths reachable from findings, producing a compact JSON (~9.5K lines for eladmin) regardless of full graph size.
Limitations
General
- Static analysis — The call graph is built from syntactic call expressions. It cannot resolve interface dispatch, reflection-based calls, or function values passed as arguments.
- One module at two versions — A finding’s
dependency.relationshipanddependency.pathcome from the resolved dependency graph. When a module resolves at more than one version (npm installs a package twice when versions conflict), a resolver that records versioned edges (npm, Maven, Gradle, Cargo) gives each copy its own route. For a resolver that does not, the finding in such a module, and any dependency reached through one, carries norelationshiporpath; itspurlandversionstay. - All paths stored — When multiple call chains exist (BFS finds all paths), all are stored in
call_chains. This ensures no reachability information is lost.
Go-specific
- Production import closure only — The inventory holds the modules that provide a package imported, directly or transitively, by a non-test package of a main module. Modules required only by
_test.gofiles or by build-tagged tool files such astools.go, and requirements nothing imports, are not resolved or scanned.--include-testsdoes not widen the inventory. - Host build configuration —
go list -depsevaluates build constraints for the scanning host’sGOOS,GOARCHand build tags (including any inGOFLAGS), so a module imported only on another platform is not inventoried, and within an imported package only the files that build compiles are scanned and parsed.CGO_ENABLEDdecides whether files that import"C"count. - Packages that fail to load — A package the go tool cannot load, such as a directory that mixes two package names or imports a module missing from the module cache while offline, is skipped with a warning. Modules imported only through it are not inventoried; the rest of the closure is.
- No cross-module method resolution — Method calls on variables (e.g.,
cipher.Encrypt()) are recorded with the variable name as the type, not the resolved type. Cross-package type resolution would require full type analysis.
Java-specific
- Gradle source archives are best-effort — Gradle dependency resolution provides binary artifact paths deterministically, but source archive availability still depends on what upstream repositories publish.
- Missing source JARs — Dependencies without sources are skipped for source scanning, but they can still contribute Java bytecode types if the compiled JAR is present locally. They cannot produce source-level findings until sources are available. Their code has no calls in the graph, so no call chain crosses them: in the callgraph export such a dependency on a finding’s
dependency.pathcarrieswithout_source: true, and a finding whose every route from the application crosses one, with no chain reaching it, readsunknown(dependency_without_source) instead ofunreachable. - Wildcard import resolution — When multiple wildcard imports could match a class name, resolution is best-effort.
- Limited inheritance/polymorphism — Variable types are tracked syntactically, but interface method call sites and abstract-class method call sites are fanned out to same-name/arity concrete implementations and subclass overrides within the same namespace root (see
expandInterfaceDispatch/expandAbstractClassDispatchininternal/callgraph/builder.go). This is a name+arity heuristic, not full type resolution, so it can over-match unrelated overloads across sibling implementations. - Multi-module Maven partial resolution — Multi-module Maven projects are supported via a three-tier fallback strategy. Tier 3 (
mvn install -DskipTests) requires compilation and may fail if the project needs specific JDK versions or build tools not available in the scan environment.
Python-specific
- Requires a Python interpreter with
pipavailable — The resolver now runspython -m pipandimportlib.metadatathrough the same interpreter. IfVIRTUAL_ENVis set, that environment’s Python is preferred; otherwise it falls back topython3thenpythonfrom PATH. - C-extension packages skipped — Packages without Python source on disk (compiled C extensions) cannot be scanned.
- Distribution-to-import mapping — Relies on
importlib.metadata(Python 3.10+) or*.dist-infofallback; packages with non-standard layouts may not be resolved. - No dynamic dispatch — Calls resolved through
getattr,__getattr__, or metaclass magic are not tracked.
Rust-specific
- Requires
cargoin PATH — The resolver shells out tocargo metadata. - No trait dispatch — Method calls on trait objects (e.g.,
dyn Cipher) are resolved syntactically by type name; trait implementations are not followed. - Macro-generated code — Functions generated by macros (e.g.,
proc_macro) are invisible to syntactic parsing. src/transparency assumption — The parser assumessrc/is the crate root; non-standard[lib] pathconfigurations may produce incorrect module paths.