groundy
Ethics, Policy & Safety

Can One Aligned LLM Serve Everyone? Portfolios vs Single-Policy Tuning

PALM preprint shows compact policy portfolios reduce approximation gaps versus single models, but production routing and audit trails remain unproven.

Published 6 references
A skeptical translucent green dinosaur raises a scuffed yellow boot beside an unused green boot and an ivory tile bearing a split footprint. Hard shadows fall across a warm ivory background.
On this page10 sections

In April 2026, a preprint called PALM reported that a compact portfolio of alignment policies can preserve near-optimal performance across all reward weightings, and that portfolios it constructed generally achieved smaller approximation gaps than same-size portfolios built from uniformly spaced or randomly sampled weights. For a governance reviewer, the practical finding is not the benchmark result itself. It is what the result implies about the question you should be asking a vendor. The question is no longer “is this model aligned?” It is “which alignment policy served which user, and can you show me the log?”

That reframing matters because a single aligned model is, by construction, a compromise. When your user base disagrees about what the assistant should value, one policy encodes one answer and silently overrules everyone else. The research reviewed here does not settle whether portfolios fix that problem in production. It does something more useful for a buyer: it shows that the averaging is documented, that runtime selection between policies is technically feasible, and that the audit infrastructure for recording which policy fired now exists as an architecture, if not yet as a deployed product. As a reviewer, you can demand the disclosure artifacts before you accept any “aligned” claim.

The fixed-aggregation critique: why one policy averages people away

The clearest statement of the problem comes from the authors of Directional Preference Alignment, who describe how multi-objective reward training actually works in practice:

“However, the multi-objective rewards are then combined in a fixed way (Wu et al., 2023b; Touvron et al., 2023, e.g.,), mainly to represent a preference averaged over different human groups, failing to capture the user-dependent preference.”

Read that sentence as a governance finding, not merely a methods critique. If the reward is a fixed weighted combination, the weights were chosen once, by someone, and every user inherits them. Two users with genuinely different preferences, say one who wants maximum brevity and one who wants exhaustive caveats, get the same compromise behavior. Neither is served; both are averaged.

The same DPA work proposes one alternative: encoding user preferences as unit vectors to get fine-grained, user-dependent control over generation. That is a per-user steering mechanism rather than a fixed policy. It matters here because it establishes that the field has been working on the heterogeneity problem from two directions: steer one model per user, or maintain several policies and route among them. The portfolio approach, which is what PALM tests, takes the second path.

What PALM actually showed, and what it did not

PALM (“Many Preferences, Few Policies: Compact Portfolios for Multi-Objective LLM Alignment”) trained and evaluated portfolios of policies using three acquisition methods with different cost profiles: full fine-tuning, parameter-efficient fine-tuning (LoRA), and decoding-time alignment. The experiments span four tasks across three models: full fine-tuning of Qwen2.5-3B-Instruct for RLVR-GSM, LoRA fine-tuning of the same model for HelpSteer, and MOD decoding on ALPACA-7B and LLaMA-2 7B for Safety Alignment and Helpful Assistants. The authors report:

“In terms of multiplicative approximation, PALM achieves the best or comparable multiplicative approximation across all configurations. Additive results follow the same trend, with a few exceptions at small sizes. PALM’s advantage becomes more consistent as portfolio size grows, with the best performance at larger sizes across all tasks.”

Three things in that sentence deserve separate treatment.

First, know what the coverage numbers cover. “Best or comparable” is a result the preprint’s authors report on four tasks across three models, none larger than 7B. The paper does quantify coverage: it proves an explicit bound on how large a portfolio must be for a given approximation tolerance, and its results tables report worst-case multiplicative and additive gaps at each portfolio size. But that coverage is over the space of reward-weight vectors. Weight vectors are mathematical objects, not people, and nothing in the paper measures how a user population distributes across them. If a vendor cites PALM-style work to claim that a handful of policies covers everyone, ask for coverage over measured preference profiles from users like yours. That is the figure this evidence does not provide.

Second, read the exceptions for what they are. The additive-metric exceptions at small portfolio sizes are cases where PALM does not beat same-size portfolios built from uniformly spaced or randomly sampled weights; they are not comparisons against a single policy. In the reported tables, worst-case multiplicative gaps shrink from one policy to two: 0.0777 to 0.0247 on RLVR-GSM, and 0.4448 to 0.0538 on Safety Alignment. The paper’s own numbers therefore support a two-policy portfolio reducing worst-case loss on these tasks. What the exceptions narrow is the construction claim: PALM’s advantage over uniform or random placement is not consistent at small sizes and stabilizes as portfolios grow. For procurement, that shifts the question from how many policies to how they were chosen.

Third, the three acquisition paths are the decision-relevant detail. Full fine-tuning per policy is expensive; LoRA adapters and decoding-time alignment are cheaper. A portfolio is only practical if most policies ride the cheap paths. The PALM experiments demonstrate all three, but on modest models: a 3B instruct model for the fine-tuning tasks and 7B models for decoding-time alignment. Whether the cost arithmetic holds at frontier sizes is not addressed in the preprint.

Routing at serving time: the operational shift

A portfolio is only useful if something decides which policy answers which request. PALM’s own deployment proposal makes the selection problem explicit:

“A provider can instead deploy a small, fixed portfolio of fine-tuned LLMs, with the assurance that every weight vector (and hence, every user) has a near-optimal option in the portfolio.”

Read that assurance closely: it concerns what the portfolio contains. Every user has a near-optimal option in the set; the preprint does not evaluate the mechanism that would match a given request to that option. Runtime routing is therefore an inference from this work, not a reported result: if preference heterogeneity is real and runtime selection is feasible, the pluralism problem moves from training-time averaging to serving-time routing.

Under single-policy tuning, the value choice happened once, at training time, inside the vendor’s pipeline. You could argue about it, but you could not observe it per request. Under runtime routing, the value choice happens per request, in your serving stack or your vendor’s. That is better, because it becomes observable in principle. It is also worse, because now there is a decision to log, a selector to audit, and a new failure mode: the router misclassifying which preference profile a user belongs to.

Nothing in the evidence reports a production deployment of portfolio routing with audit trails running under live traffic, so treat the routing half of the story as feasible engineering, not demonstrated practice.

Proving which policy served which user

Routing without records is just a new way to be unaccountable. Two September 2026 preprints supply the missing layer.

Policy-as-Skill (PaS) packages policy as “a reusable, versioned, executable capability rather than only as text.” A versioned, executable policy is exactly what an audit trail needs as its subject: you can log that policy v3.2, not some unversioned system prompt, governed this response. The Unified Policy Architecture (UPA) goes further, specifying a governance kernel designed to stay independent of any particular model provider, agent framework, or execution platform. Its policy evaluation emits governance obligations that, per the paper, “may include plugin execution, audit logging, compliance verification, governance evidence generation, notification, policy obligations, and approval metadata.”

That list is the concrete logging surface a reviewer should name in a contract: audit logging, compliance verification, governance evidence generation, approval metadata. If a vendor’s portfolio offering cannot produce those artifacts per decision, the routing is not auditable, and the portfolio is governance theater.

But auditability and accuracy are different properties, and the strongest counter-evidence in this whole stack sits inside PaS itself. The authors report that the audited configuration, PaS+Audit, reached 0.538 exact accuracy (95% bootstrap CI 0.497–0.577), macro-F1 0.346, and review F1 0.854, described as the strongest balanced PaS profile without deterministic decision override. Those numbers come from the paper’s own evaluation, and the macro-F1 of 0.346 is weak by any reading. The lesson is not that PaS fails; it is that “we log every policy decision” is fully compatible with “the policy execution is often wrong.” Demand both the audit trail and per-policy performance evidence. Neither substitutes for the other.

The EU AI Act documentation surface

If you operate in EU markets, routing disclosure is a compliance surface. COMPL-AI, a technical interpretation and benchmarking suite for the EU AI Act, identifies three technical-requirement categories for foundation models under the Act’s robustness and safety principle: Robustness and Predictability, Cyberattack Resilience, and Corrigibility. A runtime router touches all three. Predictability in particular becomes harder to argue when the same user input can reach different policies depending on routing state.

COMPL-AI also supplies a necessary dose of cold water. Its authors report that none of the examined models are fully compliant with the EU AI Act’s requirements, and that certain technical requirements “cannot be currently assessed with the available set of tools and benchmarks, either due to a lack of understanding of relevant model aspects (e.g., explainability), or due to inadequacies in current benchmarks (e.g., privacy).” Two consequences follow. Do not let a vendor equate “aligned” with “EU AI Act compliant”; the compliance bar itself is partly unmeasurable today. And expect any per-policy documentation you commission to have gaps where assessment tools do not exist, gaps that should be documented as gaps rather than smoothed over.

On the documentation format itself, AI Cards maps machine-readable AI documentation onto the EU AI Act’s technical and risk-management documentation provisions, in a modular form the authors describe as applicable across jurisdictions. For a portfolio deployment, the natural reading is one card per policy plus one for the router: each policy’s training provenance, intended preference profile, and evaluation results documented separately, with the router’s selection criteria documented as its own component. The AI Cards authors note the framework’s modular design makes it scalable beyond the EU context, which matters if you answer to more than one regulator.

A decision framework for the reviewer

The choice between accepting a vendor’s single aligned model and commissioning a portfolio is not ideological. It turns on measurable properties of your user base and your regulatory exposure. The table below maps the evidence to the decision.

SituationReasonable choiceArtifacts to demand
User preferences appear homogeneous in your measurements, low regulatory exposureAccept single-policy modelDocumentation of the aggregation choice; COMPL-AI-style evaluation against the three requirement categories
Measured preference heterogeneity across user groupsCommission compact portfolio plus routingPolicy-selection audit logs, routing disclosure, per-policy AI Cards documentation
Heterogeneity plus EU AI Act obligationsPortfolio with governance layerAll of the above, plus governance evidence generation and approval metadata per UPA’s obligation list
Any configurationDemand audit trails and performance evidence togetherPer-policy performance evidence alongside the audit log, because PaS+Audit’s reported 0.538 exact accuracy shows audited execution can still be inaccurate

Two rules sit underneath the table. First, the heterogeneity measurement is yours to make or commission; nothing in the cited research tells you your users disagree. The DPA critique establishes that averaging fails user-dependent preferences as a matter of construction, but how much that costs you depends on how much your users actually differ. Second, treat size and construction as separate questions. PALM’s small-size exceptions are against same-size uniform or random portfolios, not against a single policy, and its tables show worst-case multiplicative gaps shrinking from one policy to two on the reported tasks. The live risk at small sizes is a badly built portfolio: uniformly spaced or randomly sampled weights, the standard alternatives, can spend slots on near-duplicate behaviors.

My judgment, given this evidence: I would treat any vendor’s single “aligned” model as one averaged policy with documented user-dependent failure modes, because that is what the fixed-aggregation literature says it is. I would commission a portfolio only with contractual policy-selection audit logs, routing disclosure, and per-policy documentation attached, because the routing decision is where the values now live. And I would pilot rather than commit, because every performance number cited here comes from the papers’ own evaluations, none of it from production.

What this evidence cannot settle

The portfolio premise currently rests on one preprint, evaluated by its own authors across four tasks and three models. Before you publish any claim derived from it, internally or externally, three questions need answers the present evidence does not provide.

Coverage over people. How many of your users’ measured preference profiles does a portfolio of a given size actually serve well? The preprint reports worst-case approximation gaps over the reward-weight space and a provable bound on portfolio size; it does not measure coverage over any human population, because weight vectors are not users.

Preference provenance. Whose preferences defined the profiles the policies were trained to serve? Nothing in the cited work answers this, and it is the question a trust-and-safety lead will be asked first. A portfolio of policies derived from the vendor’s own annotator pool is a different governance object than one derived from your users.

Production auditability. UPA and PaS specify architectures for governance evidence and versioned policy execution. Whether routing plus logging survives contact with real latency budgets, real traffic, and real incident response is untested in this evidence base. The PaS results suggest the honest answer today is “partially”: the paper reports that audit/validation “can be added without changing the model decision” (PaS+Audit at 0.538 exact accuracy versus 0.533 for retrieval alone), while deterministic control reaches 0.612 at the cost of a lower macro-F1 and effects the authors describe as strongly task dependent.

Those are research-plan questions, not reasons to wait passively. The decision you can make now is narrower and more durable: stop accepting “aligned” as a property of a model, and start requiring it to be a property of a documented, logged, per-policy system. The vendors who can produce those artifacts have earned the claim. The rest are asking you to take the average on faith.

Frequently Asked Questions

What specific artifacts should a reviewer demand from a vendor offering a portfolio of alignment policies?

That list is the concrete logging surface a reviewer should name in a contract: audit logging, compliance verification, governance evidence generation, approval metadata. If a vendor’s portfolio offering cannot produce those artifacts per decision, the routing is not auditable, and the portfolio is governance theater.

Does the PALM preprint prove that a small portfolio of policies covers all real users?

But that coverage is over the space of reward-weight vectors. Weight vectors are mathematical objects, not people, and nothing in the paper measures how a user population distributes across them. If a vendor cites PALM-style work to claim that a handful of policies covers everyone, ask for coverage over measured preference profiles from users like yours. That is the figure this evidence does not provide.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. PALMarxiv.orgAccessed
  2. Directional Preference Alignmentarxiv.orgAccessed
  3. Policy-as-Skillarxiv.orgAccessed
  4. Unified Policy Architecturearxiv.orgAccessed
  5. COMPL-AIarxiv.orgAccessed
  6. AI Cardsarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy