TRECDD
TREC Dynamic Domain
The dataset used for TREC 2017 Dynamic Domain Track consists of two domains: Ebola and New York Times.
1.1 Ebola
The Ebola dataset is crawled by Juliana Friere (NYU, juliana dot freire at nyu dot edu), Kien Pham(NYU), Peter Landwehr (Giant Oak, peter dot landwehr at giantoak dot com) and Lewis McGibbney (JPL, Lewis dot J dot Mcgibbney at jpl dor nasa dot gov).
The Ebola dataset contains records related to the Ebola outbreak in Africa in 2014-2015. The original dataset includes tweets relating to the outbreak, web pages from sites hosted in the affected countries as well as PDF documents from websites such as World Health Organization, Financial Tracking Service and The World Bank. Such information resources are designed to provide information to citizens and aid workers on the ground.
1.2 New York Times
The New York Times dataset is published by Evan Sandhaus in 2008 under LDC Catalog No. LDC2008T19.
The New York Times dataset consists of articles published in New York Times from January 1, 1987 to June 19, 2007 with metadata provided by the New York Times Newsroom, the New York Times Indexing Service and the online production staff at nytimes.com. Most articles are manually summarized and tagged by professional staffs. The original form of this dataset is in News Industry Text Format (NITF). This dataset can aid the research in Document Categorization, Information Retrieval, Entity Extraction and etc.