groundy
Developer Tools

Static Analysis vs the Manual Data Map: Mapping Code Types to W3C's DPV

PrivDev maps Bearer CLI types to W3C DPV using a hybrid LLM pipeline. The authors report moderate agreement, suggesting it serves as a draft generator rather than a final GDPR

Published 4 references
A skeptical green resin dinosaur holds a yellow puzzle piece above an ivory tray with one fitted green piece and empty cavities. A separate winding green strand lies behind it, all casting shadows on warm ivory.
On this page10 sections

On September 11, 2026, at the 14th Workshop on Software Visualization, Maintenance and Evolution in São Paulo, three PUC-Rio researchers presented PrivDev, a pipeline that converts Bearer CLI scan output into a personal-data map expressed in the W3C Data Privacy Vocabulary. The PrivDev preprint (arXiv 2610.03518) reached arXiv days later, and everything known about it comes from that single author-reported, unreplicated paper. It is still the most concrete answer available to a question privacy engineers keep asking: can a static-analysis scan replace the spreadsheet-and-interview process that produces most GDPR data maps?

The short answer is no. The useful answer is more specific. PrivDev’s own design draws the line for you: 43 of Bearer CLI’s 122 data types map deterministically to DPV categories, a retrieval-grounded LLM proposes the remaining 79 with explicit high/medium/low confidence labels, and any flow the scanner cannot observe stays outside the map entirely. That three-part structure is a decision framework even before anyone replicates the numbers, because each tier fails differently and deserves a different level of trust.

What PrivDev actually does

The pipeline’s stated scope, in the authors’ words: “PrivDev maps 122 Bearer CLI data types to Data Privacy Vocabulary Personal Data (DPV-PD) categories and links them to potentially relevant GDPR provisions.”

Bearer CLI supplies the scan; DPV supplies the target vocabulary. DPV comes out of the W3C Data Privacy Vocabularies and Controls Community Group, which the SPECIAL H2020 project established in 2018, and it was built as a machine-readable interoperable vocabulary for exchanging legally relevant metadata. Its personal-data taxonomy, the pd:* classes, comprises 235 classes in DPV 2.3. That number matters later, because it defines how much room the mapping has to miss things.

The mapping itself is hybrid: “Our approach combines deterministic mapping for 43 exact-label matches with a retrieval-grounded Large Language Model (LLM) to resolve the remaining 79 non-trivial mappings.” For the LLM tier, the authors used deepseek-v4-flash, selected over GPT-4o-mini for its higher rate limit and lower per-token cost. Read that correctly: it is a throughput and cost choice. The preprint reports no accuracy comparison between the two models, so the cheaper model’s mapping quality relative to alternatives is simply unknown.

The output is not a spreadsheet. PrivDev emits an RDF knowledge graph containing 118 ODRL policy resources, structurally validated with SHACL. ODRL is the W3C standard for modelling policies and agreements, and per the DPV 2.0 paper the two vocabularies overlap but are complementary: DPV describes what the data is, ODRL expresses permissions, prohibitions and duties around it. For an audit trail, machine-readable artifacts with structural validation are a genuine step up from a wiki table. Structural validity, though, only means the artifact is well-formed. It says nothing about whether the labels inside it are right.

What a code scan can classify on its own, and what it cannot see

The 43 deterministic mappings exist because some scanner labels match DPV-PD labels exactly. When the scanner says a field is an email address and DPV-PD has a class for email addresses, no judgment is involved. That is the tier where automation is safest, and it is also the tier that tells you the least you did not already know: these are the declared, obviously-named fields any competent reviewer would catch.

The harder question is what sits outside all 122 types. Static analysis reads what code declares. It does not see what a third-party API actually returns at runtime, what a dynamic schema carries in production, or what gets derived downstream from fields that looked innocuous at the call site. None of those flows appear in a scan, so none of them can appear in a scan-derived map, and an unmapped type is not evidence that no personal data flows.

This is not a PrivDev-specific weakness. Independent work on annotation-based static analysis for personal data protection states it flatly: “It is necessary to point out the obvious: it is not possible to achieve compliance with the GDPR merely with just static analysis.” The same paper notes that the GDPR itself lays down no exact technical requirements, so there is no fixed checklist a scanner could be certified against. The ceiling is structural, not a matter of better tooling.

Worth holding both sides of this: the manual alternative is not reliable either. Empirical work on privacy-relevant source code found that “privacy code reviews, however, struggle with identifying personal data due to unclear definitions and varied contexts, increasing reliance on these tools despite their limitations in recognizing diverse personal data types”. The real choice is not automation versus a trustworthy human process. It is between two fallible processes, and the question is how to combine them so their failure modes do not overlap.

Reading the confidence labels as a triage scheme

For the 79 non-trivial mappings, PrivDev’s LLM does not just emit a label. “Instead, it returns a JSON object with three fields: pd_iri, confidence (high/medium/low), and notes (the verbatim DPV definition supporting the selection). The response is validated against a predefined schema before the artifact is generated.”

The notes field is the practically important part. Each proposed mapping carries the verbatim DPV definition the model relied on, which means a human reviewer can check the proposal against the source text without rerunning anything. That turns the LLM tier from a black box into a reviewable work queue, and the confidence label gives you the queue’s sort order. The preprint, at least as reported, does not publish how the 79 mappings distribute across the three confidence levels, so you cannot pre-commit to a review workload; treat the split as unknown until you run it on your own codebase.

Combining the pipeline’s design with the structural limits above gives a three-tier handling scheme. The tiering is my inference from the pipeline’s mechanics, not a validated error-rate finding from the paper.

TierWhat it coversBasis for the labelRecommended handling
Deterministic43 exact-label matchesLabel equality between scanner and DPV-PDAutomate; spot-check for scanner-label ambiguity
LLM high confidenceSubset of the 79Model confidence + verbatim DPV definition in notesSample-based human review against the notes field
LLM medium/low confidenceRemainder of the 79Model confidence + notesFull human review before the map is relied on
Runtime-invisible flowsThird-party payloads, derived data, dynamic schemasAbsent from scan output entirelyRuntime evidence or manual documentation; never inferred from the map

The trust ceiling

The authors did run a human evaluation, and it is the right kind of evidence to ask for. “In human evaluation, nine annotators produced 711 judgments, yielding a raw agreement of 0.72, Gwet’s AC1 of 0.68, and Gwet’s AC2 of 0.88.” Raw agreement of 0.72 and AC1 of 0.68 are moderate, not audit-grade. (The higher AC2 reflects the statistic’s treatment of category distance; it does not rescue the raw number.) If nine people judging the same mappings only agree at that level, the mapping task itself is genuinely ambiguous, which is exactly why the confidence-gated review above is not optional.

The authors say as much. Their own summary: “Our results indicate that the proposed mappings are plausible and reproducible, while also revealing ambiguities in scanner-defined data-type labels and coverage gaps in DPV-PD.” Two concessions hide in that sentence. Scanner-defined labels are ambiguous, so even the deterministic tier deserves spot checks. And DPV-PD has coverage gaps within its 235 pd:* classes, so some real personal-data categories have no mapping target at all. “Plausible and reproducible” is the authors’ framing, and the reproducibility claim is one nobody outside the team has tested.

Keep the provenance straight before any of this reaches an audit file: one preprint, one scanner (Bearer CLI), one vocabulary version (DPV 2.3), nine annotators, no independent replication, and no per-mapping error rate. Accuracy on any other scanner is untested, and nothing in the paper establishes that the output satisfies a Records of Processing Activities obligation or would be accepted by an auditor.

A pre-audit checklist for DPV-labeled output

If you adopt this style of pipeline, I would treat the following as the minimum before the map informs anything external. Each item traces to a specific property of the pipeline or its evaluation.

  1. Split the output by tier. The 43 deterministic mappings and the 79 LLM-proposed mappings fail differently. Automate the first; never bulk-approve the second.
  2. Review against the notes field, not the label. Every LLM mapping ships the verbatim DPV definition it used. Check that the definition actually covers your field’s semantics, and route every medium- and low-confidence entry to full human review.
  3. Treat unmapped scanner types as unknown, not clean. The authors document coverage gaps in DPV-PD itself. An unmapped type means the pipeline had no target, not that the field is safe.
  4. Confirm SHACL validation ran on the artifacts you actually ship. The paper validates its 118 ODRL resources structurally; your regenerated output needs the same check, since structural validity is the only automated guarantee you get.
  5. Re-derive mappings if your scanner is not Bearer CLI. All 122 types and both tiers are specific to one scanner’s label set. Portability to other scanners is untested.
  6. Add runtime evidence for flows the scan cannot see. Third-party payloads, derived data and dynamic schemas need dynamic observation or manual documentation, because GDPR compliance is not achievable by static analysis alone.
  7. Do not present the map as audit evidence in its raw form. With annotator agreement at raw 0.72 and AC1 0.68, the artifact is a draft with a known error surface. The human review record is what makes it defensible.

Verdict

Adopt the pipeline as a first-draft generator, not as a documentation system of record. Automate the deterministic tier, gate every LLM-proposed mapping on its confidence label, and keep runtime or manual verification for the flows no scan can observe. Done that way, the toolchain absorbs the most tedious part of data mapping, the part humans are demonstrably bad at, while the judgment calls stay with people whose names can go next to the decision.

What would change this recommendation is replication. If independent teams reproduce the mapping quality on other scanners and codebases, or publish per-mapping error rates, the review gates can loosen where the error rates justify it. Until then, the honest summary is the authors’ own: plausible and reproducible, with documented ambiguity and documented gaps. That is enough to save you real time. It is not enough to sign an audit on.

Frequently Asked Questions

How many Bearer CLI data types map deterministically to DPV categories in PrivDev?

43 of Bearer CLI’s 122 data types map deterministically to DPV categories, a retrieval-grounded LLM proposes the remaining 79 with explicit high/medium/low confidence labels, and any flow the scanner cannot observe stays outside the map entirely.

What does the notes field in PrivDev’s LLM output contain?

Instead, it returns a JSON object with three fields: pd_iri, confidence (high/medium/low), and notes (the verbatim DPV definition supporting the selection). The response is validated against a predefined schema before the artifact is generated.

What are the human evaluation agreement scores for PrivDev’s mappings?

In human evaluation, nine annotators produced 711 judgments, yielding a raw agreement of 0.72, Gwet’s AC1 of 0.68, and Gwet’s AC2 of 0.88.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. PrivDev preprintarxiv.orgAccessed
  2. DPV 2.0 paperarxiv.orgAccessed
  3. Empirical work on privacy-relevant source codearxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy