Study on LLMs for Promptagator-Style Dense Retriever Training
Name
3746252.3760960.pdf
Size
553.43 KB
Format
Adobe PDF
Checksum (MD5)
bac79121505fe57e5cf85122ce921105
Author(s) • •
Gwon, Daniel
Jedidi, Nour
Lin, Jimmy
Date Issued
November 10, 2025
Publisher
ACM|Proceedings of the 34th ACM International Conference on Information and Knowledge Management
Citation
Daniel Gwon, Nour Jedidi, and Jimmy Lin. 2025. Study on LLMs for Promptagator-Style Dense Retriever Training. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management (CIKM '25). Association for Computing Machinery, New York, NY, USA, 4748–4752.
Version
Final published version
Abstract
Promptagator demonstrated that Large Language Models (LLMs) with few-shot prompts can be used as task-specific query generators for fine-tuning domain-specialized dense retrieval models. However, the original Promptagator approach relied on proprietary and large-scale LLMs which users may not have access to or may be prohibited from using with sensitive data. In this work, we study the impact of open-source LLMs at accessible scales (≤14B parameters) as an alternative. Our results demonstrate that open-source LLMs as small as 3B parameters can serve as effective Promptagator-style query generators. We hope our work will inform practitioners with reliable alternatives for synthetic data generation and give insights to maximize fine-tuning results for domain-specific applications. Our code is available at https://www.github.com/mitll/promptodile
Description
CIKM ’25, Seoul, Republic of Korea
MIT Department
Lincoln Laboratory
Terms of Use
Creative Commons Attribution
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3746252.3760960