Research Radarcs.CLAug 27, 2026classified

Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers

Ke Shu, Kira Hinderks, Eetu Mäkelä, Mikko TolonenarXivPDF
cs.CL

Paper Guide Brief

Reading Brief

The paper presents a workflow study for recovering essay-scale republication and reuse from fragmented text-reuse evidence in eighteenth-century books and newspapers, focusing on pair-level evidence consolidation rather than fragment retrieval. It compares a staged rule-based workflow, decision tree and direct LLM baselines, and automated rule adaptation, showing that the final workflow achieves a strong precision-recall trade-off and produces compact, auditable candidate spaces for historical inspection.

Central Claim

A pair-level evidence consolidation workflow for essay-scale republication and reuse detection, with comparative evaluation across ECCO books and historical newspapers, demonstrating that auditable staged rule-based consolidation outperforms direct LLM baseli...

Contribution

A pair-level evidence consolidation workflow for essay-scale republication and reuse detection, with comparative evaluation across ECCO books and historical newspapers, demonstrating that auditable staged rule-based consolidation outperforms direct LLM baselines in precision-controlled candidate generation under incomplete ground truth.

Why It Matters

This contribution matters because it shifts the focus from fragment-level retrieval to pair-level evidence consolidation, enabling auditable and precision-controlled recovery of essay-scale text reuse across structurally different historic...

Prerequisites

pair-level evidence consolidation, staged rule-based workflow, decision tree, direct LLM prompting, automated rule adaptation

Atlas Placement

Natural Language Processing (subfield)

Read If

You care about pair-level evidence consolidation, staged rule-based workflow, decision tree.

Skip If

You only care about F1, precision.

Methods
pair-level evidence consolidationstaged rule-based workflowdecision treedirect LLM promptingautomated rule adaptationfeature aggregationcoverage and span featuresfragment chaining
Tasks
essay-scale republication detectiontext reuse detectionhistorical text reusetransmission relation inferencecandidate space reduction
Datasets
ECCO (Eighteenth Century Collections Online)Burney Newspapers CollectionHume essays
Benchmarks
F1precisionrecalldeployment output sizeanchor recovery

Noosaga Placements

  • Rule-Based Computational Linguisticsframework90%
    The core method is a staged rule-based workflow, which is a rule-based computational linguistics approach.
    staged rule-based workflowrule cascadeexplicit rule cascade
  • The paper addresses text reuse detection and uses LLM baselines, aligning with natural language processing tasks and methods.
    direct LLM baselinestext reuse detectionpair-level evidence consolidation
  • Statistical and Corpus-Based Computational Linguisticsframework70%
    The paper compares the rule-based workflow against statistical baselines like a decision tree and LLM-based methods, situating it within statistical and corpus-based computational linguistics.
    decision treedirect LLM settingsautomated rule adaptation
  • The work involves computational analysis of historical texts and text reuse, which falls under computational linguistics.
    historical text reusefragment-level reuse hitsessay-scale republication
  • Symbolic NLPframework60%
    The paper uses LLM baselines, which are part of symbolic NLP in the sense of rule-based prompting, and compares them to the rule-based workflow.
    direct LLM settingsLLM baselines
  • Machine Learningsubfield50%
    The paper compares a decision tree and automated rule adaptation, which are machine learning methods, though the focus is on workflow design.
    decision treeautomated rule adaptationsupervised classification

Abstract

This paper addresses the recovery of essay-scale republication and reuse from fragmented text-reuse evidence, a setting whose central challenge is pair-level evidence consolidation and not fragment retrieval alone. The study focuses on a candidate set centered on essays by eighteenth-century Scottish philosopher David Hume, spanning books from ECCO (Eighteenth Century Collections Online) and historical newspapers. Because the input consists of fragmented reuse hits instead of clean document pairs, and positive coverage is inherently incomplete, we formulate the task as pair-level evidence consolidation into plausible transmission relations and compare three methodological families: a staged rule-based workflow, baselines (a decision tree and two direct LLM settings), and automated rule adaptation. On labeled ECCO--ECCO slices, pair-level feature aggregation alone already reaches 0.948 F1 on the main labeled slice, while the final workflow gives the strongest overall precision-recall trade-off among the tested rule stages. On the full ECCO--ECCO candidate universe, direct LLM baselines flag up to 14,886 pairs as reprints compared to 771 for the final workflow, behaving in this direct-prompt setup as high-recall candidate expanders rather than precision-controlled deployment classifiers. On ECCO--Newspaper, manual audit confirms all 176 predicted positives as genuine cases of republication or reuse, while issue duplication and source-side multiplicity reveal additional provenance structure. Under incomplete ground truth, auditable pair-level evidence consolidation provides a practical way to produce compact candidate spaces for historical inspection.

Paper Context

Source ContextWhole paper
Budget100,000 tokens
Coverage33,116 chars

Classified from the full extracted paper text (33,116 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.

Full-paper context sent 33,116 of 33,116 extracted characters to classification.