A generalized solution to verify authorship and detect style change in multi-authored documents
Name
3625007.3627589.pdf
Size
921.72 KB
Format
Adobe PDF
Checksum (MD5)
b573fc41c6f817f8b67c0650e58cef46
Author(s) •
Leekha, Rohan
Vandam, Courtland
Date Issued
November 6, 2023
Publisher
ACM
Citation
Leekha, Rohan and Vandam, Courtland. 2023. "A generalized solution to verify authorship and detect style change in multi-authored documents."
Version
Final published version
Abstract
Identifying changes in style can be used to detect multi-authored social media accounts, plagiarism, compromised accounts, and author contributions in long documents. We propose an approach to recognize changes in authorship using large language models. Our approach leverages sentence-level contextual embeddings and semantic relationships. First we expand the training set by adding adversarial examples to the minority class [5], [13], [17]. Then we fine-tune a sequence classification transformer model to detect style change. Our approach outperforms all baselines of PAN21 with macro F1-scores of 0.80, 0.74, and 0.70 for detecting style changepoint between paragraphs, closed-set author ID per paragraph, and style changepoint between sentences, respectively. Our approach also performs better than the leading competitors in PAN22. Also, we achieved a five percent improvement in macro F1-score (0.78) on the newly introduced DarkReddit+ dataset for authorship verification.
Description
ASONAM '23, November 6–9, 2023, Kusadasi, Turkiye
MIT Department
Lincoln Laboratory
Terms of Use
Creative Commons Attribution
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3625007.3627589