Skip to content

One Token Can Be Enough

Bridging Prompting and Activation Steering
with Prefix Steering

arXiv preprint · 2026

PaperCode arXiv BibTeXThread In progress
TL;DR

Like a prompt, a brief steering intervention at the start can shape what follows. Prefix Steering, even over a single token, often retains much of full steering’s control while better preserving general capabilities.

The idea

Steering need not last all generation

Prompting adds a token at the start; full steering intervenes on every token, although its first intervention already reaches later tokens through attention; Prefix Steering intervenes only on a short span starting at the final prompt token, whose effect reaches later tokens through attention. INPUT GENERATION Promptingadds control text p Full steeringevery token possibly redundant Prefix Steeringfirst few tokens
  1. A prompt acts once. It sits at the start of the context, yet every later token reads it through attention.
  2. Steering usually acts on every token. Yet its first intervention also reaches later tokens through attention, an effect rarely exploited, so the later interventions may be redundant.
  3. Steering can match a prompt’s attention output. Under fixed-state attention assumptions, steering existing tokens by some \(r\) can make the attention output equal to the one a prompt induces: \(o_{\text{steer}}(h) = o_{\text{prompt}}(h)\).

What we find

Control saturates early; capability keeps falling

  • Prefix Steering
  • Full steering
  • Prompting
  • Unsteered

Enable JavaScript to see the chart; the values are in the table below.

OLMo 3 7B · HarmBench control and MMLU-Pro capability (%), by number of steered tokens. Steering a few tokens reaches most of full steering’s control, while capability declines as the span grows.

View data
Tokens steeredControlCapability
180.139.2
281.139.2
382.638.9
482.237.6
585.538.6
685.437.5
785.237.4
885.237.3
984.336.2
1084.536.1
1182.536.0
1284.535.8
1384.835.7
1484.235.5
1585.535.2
Full86.334.2
Prompting75.739.1
Unsteered36.339.3

Connecting prompting and steering

Prompting adds tokens; steering modifies existing ones. Both reach later tokens only through causal attention, so we ask when modifying an existing representation can reproduce the attention output an added prompt induces. The results below hold under fixed-state attention assumptions, holding the other tokens’ representations fixed, for one head at a fixed layer.

Result 1

Steering can reproduce a prompt

Shifting one existing token by a suitable vector gives exactly the attention output that an added prompt token would.

Result 2

The match carries over nearby

The same vector keeps working for a whole family of inputs, and the error grows at most linearly as the input drifts away from them.

Result 3

More tokens make it harder

Each extra prompt or steered token adds a constraint, so the match holds for fewer inputs; steering more positions need not make it more robust.

Theory and evidence How does the theory work? The exact statements behind the three results, the evidence in OLMo 3 7B’s own attention, and how the constructed direction relates to DiM.

Page 1 / 5

Notation, for one attention head at a fixed layer: \(h\) is the representation whose attention output we compare, \(h_{t_n}\) the final input token (the one steered), \(h_p\) an appended prompt token, and \(W_Q, W_K, W_V\) the head’s query, key, and value projections. Steering replaces \(h_{t_n}\) with \(h_{t_n} + r\).

Trading duration for strength

The connection motivates Prefix Steering: apply an existing steering direction over a short span starting at the final prompt token, then stop intervening. Prefix-k steers k positions; directions come from the Difference-in-Means (DiM) of activations on positive and negative examples. Two knobs remain, how long to steer and how hard.

Example 1 / 5

Strength

Strength can stand in for duration

  • Prefix-1
  • Prefix-5
  • Full
  • Prompting
  • Unsteered

Enable JavaScript to see the chart; the values are in the table below.

OLMo 3 7B · HarmBench control against MMLU-Pro capability, for strengths \(\alpha\) from \(10^{-4}\) to \(10\) (larger marks are stronger). Every full-steering setting is beaten on both axes by some Prefix setting; shading marks the region Prefix-5 at \(\alpha = 10\) beats.

View data
MethodαControlCapability
Prefix-11080.12739.123
Prefix-1179.34139.124
Prefix-10.170.21439.187
Prefix-10.0148.17339.286
Prefix-10.000141.84239.298
Prefix-51092.23835.591
Prefix-5185.15437.476
Prefix-50.173.28638.714
Prefix-50.0150.12739.183
Prefix-50.000147.10639.294
Full1091.31722.042
Full189.42434.487
Full0.188.28335.216
Full0.0174.19538.374
Full0.000149.52738.741
Prompting—75.73339.1
Unsteered—36.27139.3
Lemma 5 · Why it works Why can strength replace duration? One stronger push at a single position approximates weaker pushes spread over many positions.
Many weak

add \(\lambda r\) at every position in \(S\)

≈
One strong

add \(\beta r\) at a single position \(j\)

The two attention outputs nearly match when

\[\beta = \beta^{\star} = \frac{\lambda}{a_j} \sum_{i \in S} a_i\]

with \(a_i\) the pre-intervention attention weights: the single position makes up for the attention the other steered positions would have received.

Matching full steering exactly could also bring back its capability cost, so in practice we choose strength by the control–capability trade-off.

Enable JavaScript to see the chart.

OLMo 3 7B, HarmBench, 100 examples: attention-output error between full steering (\(k = 128,\ \lambda = 0.1\)) and single-token steering at strength \(\beta\), normalized by the prediction \(\beta^{\star}\). The band shows one standard deviation; the error is lowest at the predicted \(\beta^{\star}\).

Takeaway

Duration and strength trade off: a short span can make up in strength what it lacks in length, and on OLMo 3 7B every full-steering setting is beaten on both control and capability by some Prefix setting.

Across models, tasks, and methods

We compare one- and five-token Prefix Steering with full steering, prompting, and the unsteered model on Qwen3 1.7B and 14B and OLMo 3 7B and 32B, using two steering operators: additive steering, which adds the direction to the activation, and COAST, which turns the activation toward it while keeping its norm. Tasks cover safety, sentiment, politeness, and two answer-formatting tasks on MATH-500.

  • Prefix-1
  • Prefix-5
  • Full
  • Prompting
  • Unsteered
  • Prefix-1 → Prefix-5 → Full

Enable JavaScript to see the charts.

Control (vertical) against capability (horizontal), %. Capability is MMLU-Pro for safety, sentiment, and politeness, and MATH-500 for the two formatting tasks. Each model has its own axes; up and to the right is better. Paper figure (PDF)

Operators

The same pattern with both

Switch between Additive and COAST above: with either operator, Prefix-1 and Prefix-5 keep more capability than full steering in all 20 model–task settings.

Versus prompting

It depends on the task

Prompting does well on sentiment, politeness, and boxed answers. Steering gives stronger control on safety and plain-text answers, where prompting alone is less reliable.

After reasoning

Control outlasts intervention

In IF-Boxed and IF-Plain the format matters only at the final answer, long after the prefix ends, yet Prefix Steering still shapes it.

Stop early, or keep steering more gently?

On OLMo 3 7B we also compare Prefix Steering with constant full steering, linear and exponential decay, DAS, and ACT. Longer or adaptive schedules can reach stronger control, but a short prefix keeps the most capability on every task.

  • Prefix-1
  • Prefix-5
  • Full
  • Linear decay, exponential decay, DAS, ACT

Enable JavaScript to see the charts; the values are under View data.

Control (vertical) against capability (horizontal), % · OLMo 3 7B. Capability is MMLU-Pro for concept tasks and MATH-500 for formatting. Each task has its own axes; up and to the right is better. Hover a point to see its schedule.

View data
Capability (top) and control (bottom), %. Bold marks the highest capability per task.
TaskPrefix-1Prefix-5Linear decayExp. decayDASACTFull
Safety39.080.736.685.435.883.734.984.534.785.135.681.234.386.2
Sentiment39.188.238.789.935.390.434.591.235.390.136.090.234.591.0
Politeness38.986.937.688.735.989.135.990.035.689.235.888.934.289.8
IF-Boxed47.891.645.494.846.792.545.793.245.093.545.391.943.295.0
IF-Plain45.749.746.452.545.851.644.752.442.952.741.151.040.152.2
Takeaway

Across the evaluated settings, short prefixes often offer a more favorable control–capability Pareto frontier than full steering, and remain competitive with prompting and alternative strength policies.

Read the paper

One Token Can Be EnoughPDF · 32 pages

Citation

@misc{zhu2026tokenenoughbridgingprompting,
      title={One Token Can Be Enough: Bridging Prompting and Activation Steering with Prefix Steering},
      author={Xudong Zhu and Zhihui Zhu},
      year={2026},
      eprint={2610.04967},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2610.04967},
}