A preprint on detecting one-day vulnerabilities in forks maps 2.2 million public forks as descendants of known-vulnerable commits listed in the OSV advisory database. If you maintain a fork of an open-source project, that number is the practical problem: a fork created after a vulnerability was introduced but before upstream patched it can carry that vulnerability indefinitely, even though the original repository has been clean for months. The paper puts it directly: “a code repository forked after the introduction of a vulnerability, but before it is patched, may remain vulnerable long after the vulnerability has been fixed in the initial repository.” The same paper notes that the history-analysis methods used to track vulnerable software versions “fail to track vulnerabilities in forks, leaving fork maintainers to identify them manually.”
The practical consequence: forking moves CVE tracking from GitHub’s scanners to you. Dependabot, the advisory database, and dependency alerting are all built around package versions and tracked dependencies, not around whether a specific patch commit ever landed in the specific commit graph you ship. The fix is to audit forks on commit history rather than alert silence: enumerate where your fork diverges from upstream, diff it against upstream fixing commits where advisory data links to them, re-sync low-divergence forks, cherry-pick verified fixes into heavily modified ones, and treat any fork that no scanner has flagged as unverified rather than safe.
Everything capability-related in that framing comes from the preprint, which I will label as paper-reported throughout. The registry entry lists a conference venue in its journal reference (Scored 2026), but it shows no released tool you can run. The audit procedure below stands instead on measured constraints from independent studies: how complete advisory data actually is, how fast advisories arrive, and what update automation does and does not cover.
The fork window: why a patched upstream leaves you exposed
The failure mode is timing, not negligence. A bug exists in some upstream commit. You fork the repository while the bug is present. Upstream later merges a fix, issues an advisory, and its own dependents get patched through normal channels. Your fork sits on the old commit graph, and none of those channels fire, because the advisory describes an upstream package version, not your copy.
The preprint, titled “Did You Forkget It?”, estimates the exposure by propagating vulnerability information at the commit level across the global graph of public code captured by the Software Heritage archive, seeded from the OSV advisory database. Paper-reported results from the study: “Starting from 7162 repositories with vulnerable commits listed in OSV, we propagate vulnerability information to 2.2 million forks.” That is not a claim that 2.2 million forks are exploitable; it is the population of forks the method maps as descendants of vulnerable commits. But it gives a sense of how much public code sits outside the tracked-dependency model entirely.
The same study reports identifying 135 high-severity one-day vulnerabilities at a precision of 0.69, with 9 confirmed by maintainers. Those numbers matter twice: once as evidence the problem is real, and once as a calibration for anyone who wants to automate detection. I will come back to the precision figure, because it caps how much you can trust any automated flagging, including your own.
What your alerting stack cannot see
Fork owners typically lean on three signals, and each has a documented blind spot that matters for this audit.
Dependabot covers version and security updates, not fork divergence. A September 2026 study of Dependabot cooldown adoption reports that after its observation window, “GitHub made a three-day cooldown the default for Dependabot version updates without requiring explicit configuration, while leaving security updates immediate and allowing repositories to customize or opt out of the default.” That is useful for dependency manifests, and the study attributes the three-day default to a GitHub changelog entry, so the detail is paper-reported. But neither path answers the fork question. Dependabot reasons about declared dependencies and published versions; it does not compare your fork’s commit history against upstream fixing commits. If you vendored or hard-forked the code, Dependabot’s security update path has nothing to attach to.
The advisory database is not a complete patch map. Any audit that starts from “find the fixing commit, then check if my fork has it” depends on advisories that name a fixing commit. The VFCFinder study measured how often that link exists: “around 63% (6,159/9,764) of GHSA advisories do not have VFC links,” and the same study reports roughly 2,129 of those (34.6%) contain no source code link at all. So for a clear majority of GitHub Security Advisories, there is no recorded patch commit to diff against, and for a substantial minority there is not even a repository pointer. “No advisory data found” is not evidence of no exposure; it is often evidence that the advisory record is thin.
Advisory timing varies sharply by channel. A study of the GitHub Security Advisories review pipeline found that “GRAs are reviewed significantly sooner after patches than NVD-based advisories, with a median of 2 days versus 28 days.” If your audit watches NVD as its source of truth, you are working with a median four-week lag between patch and reviewable advisory. Watching GitHub repository advisories narrows that window considerably, which changes how soon a fork audit can run after a fix lands.
Manual reporting does not backstop the gaps. The last resort assumption is that if a fork shipped something dangerous, someone would have filed an issue. A study of security issue reporting in npm package repositories undercuts that: “many user-reported security issues might not be addressed by developers—they are not tagged as security-related issues and might be closed without valid justification.” Absence of reports in your fork’s issue tracker is not evidence the fork is clean. It may be evidence nobody is looking.
The audit, step by step
This runbook is my synthesis of the constraints above, not a procedure the preprint validated end to end. It borrows the preprint’s central idea, commit-level comparison rather than version-level matching, and wraps it in steps a fork owner can execute with git and the public advisory databases.
-
Inventory your forks and pin the divergence point. For each fork you ship code from, record the upstream repository and the merge-base commit: the point where your history last matched upstream.
git merge-baseagainst an upstream remote gives you this. Everything after that commit is your exposure surface. -
Seed the CVE list from OSV and GHSA. Pull advisories affecting the upstream project from OSV and the GitHub Advisory Database. This is the same seeding strategy the preprint describes, and it is deliberately advisory-driven: you are checking known, published vulnerabilities, not hunting for unknown ones.
-
Locate the fixing commit for each advisory. Where the GHSA record carries a vulnerability-fixing-commit link, this is trivial. Remember the 63% figure: most advisories lack one. For those, options include searching upstream commit messages and tags around the patch date, or using automated fixing-commit identification. VFCFinder’s authors report “96.6% recall for finding the correct VFC within the Top-5 commits, and an 80.0% recall for the Top-1 ranked commit,” and they used the tool to backfill over 300 missing fixing commits into GHSA, all accepted. Top-5 recall means the right commit is usually in a short candidate list, which makes manual confirmation practical, but the Top-1 number warns against trusting the first match blindly.
-
Determine the introducing window. A fixing commit tells you the end of the exposure window; you also need the start. If your merge-base predates the commit that introduced the bug, the fork never contained it, and you can close that CVE. The MAS-SZZ work on vulnerability-inducing commit identification describes the primitive here: “iteratively backtracking the commit history guided by the root cause, and dynamically determines the earliest vulnerability-contained commit.” That is a research system, but the manual equivalent is real work: trace the vulnerable code back through history until it appears. A cheaper approximation is to check whether the vulnerable function or file exists at your merge-base at all; if it does not, the fork predates the vulnerability.
-
Diff your fork against each fixing commit. For every CVE inside your exposure window, check whether the fix’s changes are present in your fork:
git cherryfor equivalent patches, or a direct diff when the code has drifted enough that cherry-pick equivalence fails. A fork that renamed the file or refactored the function will defeat commit-identity checks, so this step needs a human read for anything non-trivial. -
Record the verdict per fork, per CVE. Patched, not affected (predates introduction), affected and queued, or undetermined because no fixing commit could be located. That last category is where the advisory-data gap lands in practice, and it should be tracked as an open risk, not silently dropped.
Re-sync, cherry-pick, or contain
Once you know which forks are missing which fixes, the maintenance decision depends on how far the fork has drifted and how confidently you can isolate the patch. This table is a decision aid grounded in the constraints above; the fork-count and precision numbers behind it are paper-reported, and the recommended actions are inference from the measured data, not tested outcomes.
| Fork state | Advisory has VFC link | Exposure window overlaps fork | Recommended action |
|---|---|---|---|
| Low divergence (few local commits, mostly tracking upstream) | Yes | Yes | Re-sync to an upstream release that includes the fix; lowest long-term cost |
| Low divergence | Yes | No (merge-base predates introduction) | No action; record the finding |
| High divergence (heavily modified, re-sync would lose local work) | Yes | Yes | Cherry-pick the fixing commit; verify the diff applies to your modified code |
| High divergence | No (part of the ~63% without VFC links) | Suspected | Identify the fixing commit manually or with Top-5 candidate ranking; treat as open until resolved |
| Any divergence, no active maintenance capacity | Any | Confirmed affected | Contain: stop shipping from the fork or restrict its exposure; a vulnerable fork you cannot patch is a liability, not an asset |
Two judgment calls sit inside that table. First, re-syncing is almost always the better economic choice when the fork has drifted little, because every cherry-picked patch deepens the divergence that makes the next patch harder. Cherry-picking is the right tool when the fork’s local changes are the product, not an accident. Second, the last row is the one teams resist: if you lack the capacity to audit and patch a fork, the honest options are to resource it, contain it, or retire it. An unaudited fork with confirmed exposure is not neutral technical debt.
Reading the detection numbers honestly
It is tempting to hear “automated detection of one-day vulnerabilities in forks” and wait for a scanner. I would not plan around one, for three reasons grounded in the paper’s own results.
The reported precision is 0.69, which means roughly a third of flagged findings are false positives under the study’s evaluation. That is respectable for a research method operating over millions of repositories, but as an operational signal it implies heavy triage per alert. Only 9 of the 135 high-severity findings were confirmed by maintainers, so the confirmations validate the pipeline’s direction without individually validating the other 126 flags, a count derived from the reported 135 and 9. And the method is paper-reported with no tool release in the registry: you cannot run the Software Heritage-scale propagation yourself, and anyone marketing a “fork scanner” derived from this work is selling something the paper does not ship.
The useful takeaway is methodological, not product-shaped. Commit-level comparison catches what version-level matching structurally misses, because a fork’s vulnerability status lives in its commit graph, not its manifest. You can apply that principle with ordinary git tooling today, bounded by the advisory-data gaps and by your own triage capacity. That is the inference this article is built on, and it is worth stating plainly: the preprint demonstrates the problem’s scale and one detection approach; the procedure here is assembled from the independent measurements of what advisory data, automation, and manual reporting actually deliver.
Making it a standing policy
A one-time audit decays the moment upstream cuts another patch. Two parameters from the evidence should set the cadence. Watch GitHub repository advisories rather than NVD alone, because the median review latency of 2 days versus 28 days determines how soon an audit can run after a fix. And re-run the audit on a schedule matched to upstream’s patch tempo, since every new fixing commit upstream reopens the question for every fork in your inventory.
Assign ownership explicitly. The preprint’s core finding is that no platform system does this tracking for forks, which means “someone will notice” is a policy of hoping. The team that decided to fork owns the divergence, the audit, and the re-sync-or-cherry-pick call.
There is one honest limit to close with. History-based detection can confirm that a known, published vulnerability with an identifiable patch did or did not reach your fork. It says nothing about vulnerabilities that have no advisory, advisories in the 63% without a fix link that you fail to resolve, or bugs introduced by your own fork’s changes. A clean audit result means “no known, traceable exposure found,” and that is a far better position than scanner silence. It is not a guarantee, and treating it as one recreates the same assumption this whole exercise exists to break.
Frequently Asked Questions
How many public forks are mapped as descendants of known-vulnerable commits?
Starting from 7162 repositories with vulnerable commits listed in OSV, we propagate vulnerability information to 2.2 million forks.
What is the median review latency for GitHub repository advisories compared to NVD-based advisories?
GRAs are reviewed significantly sooner after patches than NVD-based advisories, with a median of 2 days versus 28 days.
What percentage of GitHub Security Advisories lack a vulnerability-fixing-commit link?
around 63% (6,159/9,764) of GHSA advisories do not have VFC links

Join the discussion
Share a useful perspective or ask a question about this article.