Awesome Information Retrieval 
Curated list of information retrieval and web search resources from all around the web.
Curated list of information retrieval and web search resources from all around the web.
Information Retrieval involves finding relevant information for user queries, ranging from simple domain of database search to complicated aspects of web search (Eg - Google, Bing, Yahoo). Currently, researchers are developing algorithms to address Information Need of user(s), by maximizing User and Topic Relevance of retrieved results, while minimizing Information Overload and retrieval time.
C.D. Manning, P. Raghavan, H. Schütze. Cambridge UP, 2008. (First book for getting started with Information Retrieval).
Bruce Croft, Don Metzler, and Trevor Strohman. 2009. (Great book for readers interested in knowing how Search Engines work. The book is very detailed).
R. Baeza-Yates, B. Ribeiro-Neto. Addison-Wesley, 1999.
B. Croft, D. Metzler, T. Strohman. Pearson Education, 2009.
S. Chakrabarti. Morgan Kaufmann, 2002.
W.B. Croft, J. Lafferty. Springer, 2003. (Handles Language Modeling aspect of Information Retrieval. It also extensively details probabilistic perspective in this domain, which is interesting).
Ed Greengrass, 2000. (Comprehensive survey of Conventional Information Retrieval, before Deep Learning era).
C.T. Meadow, B.R. Boyce, D.H. Kraft, C.L. Barry. Academic Press, 2007 (library/information science perspective).
Matthew Lease (University of Texas at Austin).
Chris Manning and Pandu Nayak (Stanford University).
Raymond J. Mooney (University of Texas at Austin).
Vagelis Hristidis (University of California - Riverside).
Ray R. Larson (UC berkeley).
Jamie Callan (CMU).
David Yarowsky (John Hopkins University).
Andrea LaPaugh (Princeton University).
Dr. Jilles Vreeken , Prof. Dr. Gerhard Weikum (MPI).
Prof. ChengXiang Zhai (University of Illinois at Urbana-Champaign).
Open Source Search Engine that can be used to test Information Retrieval Algorithm. Twitter uses this core for its real-time search.
The Lemur Project develops search engines, browser toolbars, text analysis tools, and data resources that support research and development of information retrieval and text mining software.
Linked data web.
This is one of the first collections in IR domain, however the dataset is too small for any statistical significance analysis, but is nevertheless suitable for pilot runs.
TREC is the benchmark dataset used by most IR and Web search algorithms. It has several tracks, each of which consists of dataset to test for a specific task. The tracks along with suggested use-case are:
This is one of the largest Web collection of documents obtained from crawl of government websites by Charlie Clarke and Ian Soboroff, using NIST hardware and network, then formatted by Nick Craswel.
This is collection of wide variety of dataset ranging from Ad-hoc collection, Chinese IR collection, mobile clickthrough collections to medical collections. The focus of this collection is mostly on east asian languages and cross language information retrieval.
It contains a multi-lingual document collection. The test suite includes:
The corpora is now available through NIST. The corpora includes following:
This data set consists of 20000 newsgroup messages.posts taken from 20 newsgroup topics.
This data set is a comprehensive archive of English newswire text data including headlines, datelines and articles.
Past newswire/paper datasets (DUC 2001 - DUC 2007) are available upon request.
Manik Verma (Microsoft Research)
Tim Berners-Lee (Ted Talk) [Tim Berners-Lee invented the World Wide Web. He leads the World Wide Web Consortium (W3C), overseeing the Web's standards and development].
Gary Flake, Technical Fellow at Microsoft (TED Talks).
Jeff Dean (WSDM Conference, 2009).
David Wilne (The University of Waikato, 2008).
Steve Tjoa (RackSpace Developers) [This talk shows that IR is not just text and images].
Liron Shapira (Box Tech Talk).
Doug Imbruce (Techcrunch Disrupt)[Doug Imbruce is the Founder of Qwiki, Inc, a technology startup in New York, NY, acquired by Yahoo! in 2013].
Dr. Alma Whitten (Google Brussels Tech Talk).
Andreas Ekström (Swedish Author & Journalist, TED Talk).
Eli Pariser (Author of the Filter Bubble, TED Talk).
Andy Yen (CERN, TED Talk) [This talk talks about privacy, which Search Engines intrude into, and how can people protect it].
Michael Douglas [TEDx SouthBank].
Christopher "moot" Poole" (Ted Talks) [Christopher "moot" Poole is founder of 4chan, an online imageboard whose anonymous denizens have spawned the web's most bewildering and influential subculture].
Google Research.
Dr. Edel Garcia.
Information Extraction.
Information Retrieval from Lip Reading.
Sketch-based search.
Information Extraction.