44 x 32

arXiv:2610.04210, cs.CL. Benchmark: 300 matched editing clusters, 1,800 cases, 20 models from 11 labs

Can LLMs Separate Pasted Artifacts from User Speech? Absorption at Unmarked Prompt Seams

Sugam Panthi, Muhaiminul Yeamin, and Rabab Abdelfattah

The University of Southern Mississippi. SEAM measures whether the sentence you type after a paste comes back inside the artifact the model edits.

QR code: Scan for paper

Scan for paper

The problem

One turn, two sources, no boundary the model can see

A person types a task, pastes an artifact, then keeps typing below it. The interface knows where the paste ended. The model often does not, because the turn arrives as one flat string. We call that join the seam.

Absorption is the model returning that trailing speech inside the edited artifact. The speech is benign and asks for nothing, and the output stays fluent, so the mistake is easy to miss.

USER MESSAGE Improve the following passagefor clarity and correctness. The users, especially retailersand product manufacturers, playimportant roles in improvingbar code. I still need to send the budgetspreadsheet. PASTED ARTIFACT TYPED AFTER ✕ WITHOUT MARKER …important roles in improvingbar code. (no marker) I still need to send thebudget spreadsheet. ✓ WITH MARKER …important roles in improvingbar code. </artifact-9164bf2ef286> I still need to send thebudget spreadsheet. LLM RESPONSE Users, especially retailers andproduct manufacturers, playimportant roles in improvingbarcode systems. I still need to send the budgetspreadsheet. Users, particularly retailers andproduct manufacturers, play a keyrole in improving barcodetechnology.
One real SEAM cluster from CoEdIT. Task, artifact and typed continuation are identical on both paths, and both responses are verbatim DeepSeek V4 Flash output. Only the seam marking differs. Data: paper Figure 1, redrawn.

SEAM-Bench leaderboard: absorption rate, 300 editing tasks per model, 20 models from 11 labs

How often does a model put your words into the artifact? 7.7% to 66.7% at a bare newline

EXPLICIT SEAM INFORMATION Bare newline Blank line Boundary tags Tags plus instruction Artifact-native comment Model Marker gain Llama-3.1-8B 7.7 11.7 7.3 0.3 25.0 1.0× Claude-Opus-4.8 19.0 21.0 2.0 0.0 56.0 9.5× Llama-3.3-70B 19.0 20.3 14.0 0.0 32.2 1.4× MiMo-v2.5 20.0 19.7 5.0 1.3 40.3 4.0× MiMo-v2.5-Pro 21.7 20.0 4.0 0.0 48.5 5.4× DeepSeek-V4-Pro 22.3 19.0 1.0 0.0 49.4 22× MiniMax-M3 22.3 21.7 3.3 0.0 46.0 6.7× Gemma-3-27B 27.7 35.0 20.3 0.7 54.0 1.4× MiniMax-M2.5 28.7 26.4 6.4 1.7 48.6 4.5× Gemini-3.1-Pro 29.3 32.3 3.0 0.0 64.3 9.8× DeepSeek-V4-Flash 31.3 30.3 4.4 0.3 51.3 7.1× GPT-5.6-sol 32.0 36.7 11.7 0.0 65.3 2.7× Qwen3-32B 32.7 33.0 5.3 0.0 52.7 6.2× gpt-oss-20b 35.1 41.8 16.4 0.0 54.9 2.1× Gemma-3-12B 38.5 44.1 19.1 6.0 54.2 2.0× gpt-oss-120b 39.3 43.0 18.3 0.3 44.7ⁿˢ 2.1× Mistral-Small-3.2-24B 41.3 47.3 28.0 12.0 53.7 1.5× Gemma-3-4B 52.0 51.3 19.7 9.3 59.7 2.6× Qwen3-8B 53.0 53.3 26.0 0.0 56.0ⁿˢ 2.0× OLMo-2-32B 66.7 72.3 35.3 0.3 60.0ⁿˢ 1.9× 0 75% absorbed. Lower is better; the whisker is the 95% Wilson interval.
Rows sorted by bare-newline rate, best first. Every model absorbs at a bare newline. A blank line produces no significant reduction in any model, and raises it significantly in three. Boundary tags reduce absorption significantly in 19 of 20 after Holm correction, Llama-3.1-8B excepted (exact McNemar p = 1.0). Marker gain is the bare-newline rate divided by the boundary rate. The tags-plus-instruction column is not a free win, see cell 5. Rates: paper Table 3. Wilson intervals: the release ledger the table's caption defers them to.

Method: 300 clusters, six versions each

Hold the task and the artifact fixed, vary only the seam

Clean
no trailing text; the scored control
Newline
casual comment on the next line
Blank
same comment, one blank line added
Boundary
tags around the artifact, comment outside
Mitigation
same tags, plus an exclude instruction
Artifact-native
comment rewritten to fit the genre

A hit is comment content inside the returned artifact, paraphrase included, and absent from that model's clean output. Detector: exact witness, then stem-tolerant co-occurrence, then deberta-large-mnli. Six sources, 50 clusters each. Data: paper Table 2, Section 4.

Consequence: same instruction, same artifact, same bare newline

A comment that fits the artifact is absorbed far more often

0 25 50 75 100 AR % DeepSeek V4 Flash 31.3 → 51.3 code examples 0 → 35 of 100 Qwen3-32B 32.7 → 52.7 code examples 0 → 19 of 100 Claude Opus 4.8 19.0 → 56.0 code examples 0 → 81 of 100 GPT-5.6-sol 32.0 → 65.3 code examples 0 → 42 of 100 casual artifact-native
Higher in 19 of 20 models, significantly in 17 after Holm correction, by 7.7 to 37.0 points among those 17. Shown are the four models whose detected positives were reviewed in full. Code looked safe only because a casual sentence fits code badly. Gemini 3.1 Pro goes from 29.3% to 64.3% overall and from 1 to 76 of 100 code examples. Exact McNemar p below 1.6e-8 for all four. Data: paper Table 5.

Validity: what the rate does and does not show

The instruction that drives absorption to zero also breaks the edit

Adding “do not include outside text” on top of the tags puts absorption at or below the boundary rate in all 20 models and at zero in ten. It is not a free win.

OPUS OUTPUTS THAT FAIL TO PARSE AS PYTHON, OUT OF 50 Clean control 0 Boundary tags 0 Tags plus instruction 16
Claude Opus 4.8 on CanItEdit. Exact McNemar p = 3.05e-5 against the clean control; eight of the 16 parse once echoed wrapper lines are stripped. Tags alone cost nothing on this check. Reviewing positives constrains false positives but does not measure recall, so every rate on this sheet is an operational lower bound. Two known false positives remain in Opus's 2.0% boundary rate; removing them gives 1.3%. Data: paper Sections 5.5 and 6.

TL;DR

What we did

SEAM renders 300 editing clusters six ways, holding the instruction and the pasted artifact fixed and changing only the trailing text and the seam. A deterministic scorer compares every treatment with its own clean control across 20 models from 11 labs.

What we found

Every model returns benign trailing speech inside the artifact at a bare newline, from 7.7% to 66.7%. A blank line helps in no model; explicit tags help in 19 of 20. A comment written like the artifact is absorbed more in 19 of 20, and code goes from 0 absorbed to as many as 81 of 100.

What you should do

Pass the paste boundary the interface already has instead of hoping the model infers it, and wrap pasted regions in tags with a per-message suffix. Keep measuring afterwards: tags still leave 11.7% on GPT-5.6-sol and 35.3% on OLMo-2-32B, and a stricter instruction buys the last points by breaking the edit.

QR code linking to this poster online at spanthi.com

This poster online

+ and − zoom, 0 fits, scroll pans, ⌘/ctrl + wheel zooms, Esc or X closes