Scatter/Gather: A Cluster-based Approach to Browsing Large Document Collections
Name
ScatterGather_A_Cluster-based_Approach_to_Browsing.pdf
Description
Accepted version
Size
220.96 KB
Format
Adobe PDF
Checksum (MD5)
c6013c8a7501151f469aa9cd7a3963d2
Author(s) • • •
Cutting, Douglass R
Karger, David R
Pedersen, Jan O
Tukey, John W
Date Issued
2017
Journal
ACM SIGIR Forum
Publisher
Association for Computing Machinery (ACM)
Version
Author's final manuscript
Abstract
Document clustering has not been well received as an information retrieval tool. Objections to its use fall into two main categories: first, that clustering is too slow for large corpora (with running time often quadratic in the number of documents); and second, that clustering does not appreciably improve retrieval.
We argue that these problems arise only when clustering is used in an attempt to improve conventional search techniques. However, looking at clustering as an information access tool in its own right obviates these objections, and provides a powerful new access paradigm. We present a document browsing technique that employs docum-ent clustering as its primary operation. We also present fast (linear time) clustering algorithm.
We argue that these problems arise only when clustering is used in an attempt to improve conventional search techniques. However, looking at clustering as an information access tool in its own right obviates these objections, and provides a powerful new access paradigm. We present a document browsing technique that employs docum-ent clustering as its primary operation. We also present fast (linear time) clustering algorithm.
MIT Department
Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
Massachusetts Institute of Technology. Computer Science and Artificial Intelligence Laboratory
Terms of Use
Creative Commons Attribution-Noncommercial-Share Alike
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1145/3130348.3130362