Pair-Level Essay-Scale Republication and Reuse from Fragmented Historical Text Reuse: A Workflow Study on Eighteenth-Century Books and Newspapers
Paper Guide Brief
Reading Brief
The paper presents a workflow study for recovering essay-scale republication and reuse from fragmented text-reuse evidence in eighteenth-century books and newspapers, focusing on pair-level evidence consolidation rather than fragment retrieval. It compares a staged rule-based workflow, decision tree and direct LLM baselines, and automated rule adaptation, showing that the final workflow achieves a strong precision-recall trade-off and produces compact, auditable candidate spaces for historical inspection.
Central Claim
A pair-level evidence consolidation workflow for essay-scale republication and reuse detection, with comparative evaluation across ECCO books and historical newspapers, demonstrating that auditable staged rule-based consolidation outperforms direct LLM baseli...
Contribution
A pair-level evidence consolidation workflow for essay-scale republication and reuse detection, with comparative evaluation across ECCO books and historical newspapers, demonstrating that auditable staged rule-based consolidation outperforms direct LLM baselines in precision-controlled candidate generation under incomplete ground truth.
Why It Matters
This contribution matters because it shifts the focus from fragment-level retrieval to pair-level evidence consolidation, enabling auditable and precision-controlled recovery of essay-scale text reuse across structurally different historic...
Prerequisites
pair-level evidence consolidation, staged rule-based workflow, decision tree, direct LLM prompting, automated rule adaptation
Atlas Placement
Natural Language Processing (subfield)
Read If
You care about pair-level evidence consolidation, staged rule-based workflow, decision tree.
Skip If
You only care about F1, precision.
Noosaga Placements
- Rule-Based Computational Linguisticsframework90%The core method is a staged rule-based workflow, which is a rule-based computational linguistics approach.staged rule-based workflowrule cascadeexplicit rule cascade
- The paper addresses text reuse detection and uses LLM baselines, aligning with natural language processing tasks and methods.direct LLM baselinestext reuse detectionpair-level evidence consolidation
- Statistical and Corpus-Based Computational Linguisticsframework70%The paper compares the rule-based workflow against statistical baselines like a decision tree and LLM-based methods, situating it within statistical and corpus-based computational linguistics.decision treedirect LLM settingsautomated rule adaptation
- The work involves computational analysis of historical texts and text reuse, which falls under computational linguistics.historical text reusefragment-level reuse hitsessay-scale republication
- Symbolic NLPframework60%The paper uses LLM baselines, which are part of symbolic NLP in the sense of rule-based prompting, and compares them to the rule-based workflow.direct LLM settingsLLM baselines
- The paper compares a decision tree and automated rule adaptation, which are machine learning methods, though the focus is on workflow design.decision treeautomated rule adaptationsupervised classification
Abstract
This paper addresses the recovery of essay-scale republication and reuse from fragmented text-reuse evidence, a setting whose central challenge is pair-level evidence consolidation and not fragment retrieval alone. The study focuses on a candidate set centered on essays by eighteenth-century Scottish philosopher David Hume, spanning books from ECCO (Eighteenth Century Collections Online) and historical newspapers. Because the input consists of fragmented reuse hits instead of clean document pairs, and positive coverage is inherently incomplete, we formulate the task as pair-level evidence consolidation into plausible transmission relations and compare three methodological families: a staged rule-based workflow, baselines (a decision tree and two direct LLM settings), and automated rule adaptation. On labeled ECCO--ECCO slices, pair-level feature aggregation alone already reaches 0.948 F1 on the main labeled slice, while the final workflow gives the strongest overall precision-recall trade-off among the tested rule stages. On the full ECCO--ECCO candidate universe, direct LLM baselines flag up to 14,886 pairs as reprints compared to 771 for the final workflow, behaving in this direct-prompt setup as high-recall candidate expanders rather than precision-controlled deployment classifiers. On ECCO--Newspaper, manual audit confirms all 176 predicted positives as genuine cases of republication or reuse, while issue duplication and source-side multiplicity reveal additional provenance structure. Under incomplete ground truth, auditable pair-level evidence consolidation provides a practical way to produce compact candidate spaces for historical inspection.
Paper Context
Classified from the full extracted paper text (33,116 characters). The Paper Guide brief above is the user-facing synthesis; raw context is kept out of the page.
Full-paper context sent 33,116 of 33,116 extracted characters to classification.