groundy
Industry & Business

Pricing LLM Training Data: Measured Gain vs Flat Token Rates

A new preprint proposes pricing LLM training data by measured gain rather than flat token rates, offering a hybrid contract structure for buyers and sellers.

Published 5 references
A skeptical green resin dinosaur examines a clear block containing ivory chips, resting one foot on a closed capsule containing a yellow disc. Hard shadows fall across the warm ivory background.
On this page10 sections

Flat per-token rates tell neither side of a training-data deal what a corpus actually contributes to a model. A new arXiv preprint, Utility-Aware Data Pricing, proposes pricing data by token-level quality plus empirically measured training gain instead of raw volume. The framework is early and unproven at production scale, but it gives buyers and sellers a concrete vocabulary for a better deal structure: a flat base rate plus a measured-gain kicker, with the measurement protocol written into the contract.

Why a flat per-token rate prices the wrong thing

A per-token or per-document rate assumes every unit of text contributes equally to the model that consumes it. The data-valuation literature has been dismantling that assumption for years. The Data Shapley work by Ghorbani and Zou established the foundational point: the value of a datum is not a universal property. It depends on three things: the learning algorithm, the performance metric, and the rest of the training data. A corpus that helps one buyer’s model on one benchmark may be redundant, or worse, for another buyer training on different data.

Worse is a real outcome, not a rhetorical one. In the same paper’s patient-data acquisition experiment, adding data points with low estimated value actively hurt the trained model’s performance. A flat rate pays for that harm at the same price as the most useful records in the deal. Anyone who has watched targeted Wikipedia edits leave a detectable signal in Llama pretraining already knows that raw volume ratios say nothing about per-token influence during training; pricing by volume inherits the same blindness.

Flat rates survive because they are cheap to compute and hard to dispute. Both parties can count tokens. That convenience is the entire argument for the status quo, and the new preprint is an attempt to make the alternative measurable enough to negotiate over.

What the preprint actually proposes

Utility-Aware Data Pricing, observed October 1, 2026, builds a valuation framework on three pillars and explicitly positions itself against row-based counting:

  1. Token-level information density. Tokens are weighted by information entropy, semantic richness, and syntactic coherence rather than counted raw.
  2. Empirical training gain. Data contribution is quantified through influence functions, lightweight proxy-model strategies, and Data Shapley values, with a small proxy model standing in as a tractable estimator for the target LLM.
  3. Cryptographic verifiability. Data contribution records are made auditable and tamper-evident.

The authors report preliminary experiments in which proxy gain was the strongest valuation signal. The qualifier matters: it was strongest under the proxy-defined value function, while other signals varied across tasks. That is a circularity risk. If you define value in terms of proxy-model improvement and then find that proxy-model improvement is the best predictor of value, the result is nearly tautological. This is a single-group, author-reported preprint with no independent replication, and the authors themselves list evaluation on larger models and datasets as future work.

Two implementation details deserve more attention than they will get in most coverage.

The tokenizer is not your tokenizer. All text in the framework is tokenized with a lightweight regex-based tokenizer matching alphanumeric words and individual punctuation marks from lowercased input, with no pretrained subword vocabulary. Its “tokens” are roughly words, not LLM-vocabulary tokens. Any comparison between the paper’s quality-weighted token counts and the per-token rate on a vendor invoice requires a conversion the paper does not supply, and that conversion is itself a negotiation surface.

Gain is not attributed per token. The paper states plainly that the framework “uses token-level information as a basis for quality assessment, but does not attempt to attribute training contribution to individual tokens.” The token-level signals are quality proxies. The measured training gain lives at the corpus or dataset level, estimated through proxy models. A seller quoting per-token gain figures is claiming something the underlying method does not produce.

Three price structures, three measurement bills

The practical question for procurement teams is not whether gain-based pricing is “better” in the abstract. It is what each structure costs to measure, what each leaves disputable, and who carries which risk.

Flat per-token / per-documentQuality-weightedGain-based (corpus level)
Pricing basisRaw token or document countToken counts weighted by entropy, semantic richness, coherenceProxy-model training gain estimated via holdout runs
Measurement costNear zero; both parties can countCheap once the tokenizer and quality metrics are fixedMultiple proxy training runs per estimate; Monte Carlo averaging
Model-dependenceNone (and that is the flaw)Low; quality scores are model-agnosticHigh; gain is specific to the buyer’s algorithm, metric, and co-training data
Attribution granularityPer token, triviallyPer token, as quality proxyCorpus level only; per-token gain claims unsupported
What can be disputedWhether volume equals valueWhich tokenizer and quality metrics applyEverything: proxy choice, sample count, value function, approximation validity
Re-evaluation triggerNone neededWhen quality metrics changeWhenever the buyer’s model, metric, or training mix changes
Who holds the riskBuyer pays for low- or negative-value dataBuyer pays for quality, not outcomesSeller must fund or share measurement; both sides dispute estimates

The rows are not symmetric. Flat pricing’s advantage is real: it is the only structure where the price is verifiable by either party in an afternoon. Quality weighting keeps most of that convenience while refusing to treat a boilerplate token and a dense technical token as equivalent. Gain-based pricing is the only structure that measures what buyers actually care about, and it is the most expensive and contestable to operate.

The gain estimate is an approximation, and the guarantees do not survive it

Two constraints from the underlying literature should be on the table before any gain number is quoted in a negotiation.

The first is cost. In the Data Shapley method, a single estimate of a data source’s marginal contribution is one Monte Carlo sample of its Shapley value. The authors describe the remedy: repeat the process and average the marginal contributions. Each repetition is another training run, even if only of a proxy model. This is the same economic logic that makes predicting a fine-tune’s payoff before training finishes valuable: every measurement run has a compute bill attached, and the question of who pays it, buyer, seller, or split, is a contract term, not a technicality. The preprint’s use of lightweight proxy models is precisely an attempt to push this bill down, but it does not eliminate the repetition requirement, and it adds a new dispute: whether the proxy’s gain predicts the target model’s gain. The paper’s own evidence for that link is the proxy-defined value function result, which is the circularity noted above.

The second is validity. A survey of the Shapley value in machine learning notes that most applications use approximations, and that the axiomatic properties that justify Shapley values, the fairness guarantees that make the method more than a heuristic, do not hold under those approximations. The survey adds that this is often overlooked. The consequence for a licensing deal is direct: any Shapley-derived gain figure a seller brings to the table is a contestable estimate, not a quantity with provable fairness properties. A buyer’s analysts can rerun the approximation with different sampling and get a different number. Neither number is wrong in a way a court or arbitrator could easily resolve.

This does not make gain measurement useless. It makes it a negotiated instrument rather than a meter. The price it supports should look like an earn-out or a performance bonus, bounded and conditional, not like a metered utility rate.

Audit trails prove the record exists, not that the number is right

The preprint’s third pillar, cryptographic verifiability, is easy to oversell. The paper is explicit about the boundary: the cryptographic layer enables tamper-evident and traceable records but “does not guarantee computational correctness.”

That distinction has a clean contract analog. Tamper-evidence answers “was this measurement record altered after the fact?” It does not answer “was the measurement done correctly, with the agreed proxy model, the agreed number of Monte Carlo repetitions, and the agreed tokenizer?” The first question is what the cryptography handles. The second requires audit rights: access to the measurement pipeline, the proxy model specification, the random seeds or sampling protocol, and the holdout data. A deal that adopts gain-based pricing without those rights has replaced an unverifiable price (flat volume) with an unverifiable price that carries a cryptographic seal.

Buyers should also note what tamper-evident provenance records are good for independently of gain pricing: establishing which data was delivered, when, and in what form. That has value even in a pure flat-rate deal, because it grounds the diligence questions in the next section.

Provenance risk is the sleeper cost in flat-rate deals

The case for weighing tokens by quality and gain usually focuses on model performance. There is a second axis that flat rates ignore entirely: legal provenance, where the evidence is older but concrete.

A survey of 201 high-reputation Stack Overflow answerers found that 138 of them, 69%, never check for licensing conflicts between their copied code snippets and Stack Overflow’s CC BY-SA 3.0 license, according to the Toxic Code Snippets study. The same study surveyed 87 Stack Overflow visitors and found 85% were not aware of the CC BY-SA 3.0 license and 66% never check for license conflicts when reusing code. The researchers also identified 214 code snippets that could potentially violate the licenses of their original software, appearing 7,112 times across 2,427 GitHub projects.

An independent study, Stack Overflow: A Code Laundering Platform?, found 1,219 Stack Overflow posts with potential license violations, and observed that developers tend to share entire classes in question posts without license information.

Two cautions on these numbers. They date from 2017 and 2018, they measure license awareness and conflicts under CC BY-SA 3.0 rather than licensing rates, and they say nothing about Stack Overflow’s current terms. What they establish is narrower and still useful: community-sourced code corpora, among the most valuable training data categories, have documented, quantified provenance hygiene problems. A flat per-token rate prices a snippet with clean provenance identically to one copied out of a GPL project by an answerer who never checked. The buyer inherits that risk whether or not the price reflects it.

This is where the preprint’s audit pillar connects to an existing practice problem. Groundy’s coverage of synthetic data eating AI training makes a related point about provenance from the content side: pipelines that cannot keep human and model-generated populations separable cannot bound what the model later leaks or inherits. Provenance risk and quality risk compound. Both are invisible to volume pricing.

When the model changes, the price changes

The Data Shapley result that value depends on the learning algorithm, the performance metric, and the co-training data has a consequence that neither flat nor quality-weighted pricing faces: a measured-gain price expires.

A gain estimate produced for one buyer’s model, on one metric, against one training mix, does not transfer. It does not transfer to the buyer’s next model generation, and it certainly does not transfer to a different buyer. This cuts in both directions in a negotiation. Sellers cannot claim a permanent, portable gain figure for their corpus. Buyers cannot amortize one measurement across future models. Any gain-based deal needs explicit re-evaluation triggers: a new model generation, a changed performance metric, a materially changed training mix. And each trigger raises the question the parties should settle at signing rather than at renewal: who pays for the re-measurement runs.

There is an asymmetry worth exploiting. The seller has the stronger incentive to fund initial measurement, because the burden of proof has shifted to them: under flat pricing, sellers of low-value data get paid the same as everyone else, and gain measurement is what ends that subsidy. The buyer has the stronger incentive to control re-evaluation timing, because re-measurement is where the buyer’s model roadmap leaks into a vendor relationship. Splitting initial measurement costs to the seller and re-evaluation costs to the buyer is a defensible starting position, though the evidence here supports the structure of the problem, not any particular split.

The defensible deal: base plus kicker

No supplied evidence shows any real licensing deal using gain-based pricing. The research packet contains no deal terms, no rates, and no market data of any kind, and vendor claims of “utility-priced data” should be treated as unverified until a contract shows otherwise. What the evidence does support is a contract-design inference, and it is a bounded one.

The defensible structure, given everything above, is a hybrid:

  • A flat per-token base rate. Both parties can compute it, it covers the seller’s costs, and it anchors the deal in something verifiable. The tokenizer used for counting should be specified in the contract, because the preprint’s own regex tokenizer demonstrates that “a token” is a definitional choice, not a fact.
  • A measured-gain kicker, at corpus level. Estimated via proxy-model holdout runs, with the measurement protocol, the proxy model, the number of Monte Carlo repetitions, and the value function all contracted explicitly. The kicker should be bounded and conditional, in recognition that the underlying estimate is an approximation whose axiomatic guarantees do not hold.
  • Re-evaluation triggers and audit rights. Tamper-evident records as a baseline, plus access rights sufficient to check computational correctness, because the cryptography alone does not provide it.
  • A hard refusal on per-token gain claims. The framework itself does not attribute training contribution to individual tokens. A seller quoting per-token gain is selling a number the method does not produce.

The strongest limitation bears repeating without softening: the preprint is author-reported, single-group, and preliminary, its headline result held only under its own proxy-defined value function, and its authors list larger-model evaluation as future work. Whether proxy-model gain predicts target-LLM gain at production scale is the open question the whole structure depends on. Until someone answers it with an independent replication, gain-based pricing components should be sized like what they are: performance bonuses built on contested estimates, not metered measurements of value.

Frequently Asked Questions

Does the proposed framework attribute training gain to individual tokens?

The paper states plainly that the framework “uses token-level information as a basis for quality assessment, but does not attempt to attribute training contribution to individual tokens.” The token-level signals are quality proxies. The measured training gain lives at the corpus or dataset level, estimated through proxy models. A seller quoting per-token gain figures is claiming something the underlying method does not produce.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Utility-Aware Data Pricingarxiv.orgAccessed
  2. Data Shapleyarxiv.orgAccessed
  3. Survey of the Shapley Value in Machine Learningarxiv.orgAccessed
  4. Toxic Code Snippetsarxiv.orgAccessed
  5. Stack Overflow: A Code Laundering Platform?arxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy