Transcription of The Unreasonable Effectiveness of Data
{{id}} {{{paragraph}}}
EXPERT OPINION81541-1672/09/$ 2009 IEEEiEEE iNTElliGENT SYSTEMSP ublished by the IEEE Computer SocietyContact Editor: Brian Brannon, as f = ma or e = mc2. Meanwhile, sciences that involve human beings rather than elementary par-ticles have proven more resistant to elegant math-ematics. Economists suffer from physics envy over their inability to neatly model human behavior. An informal, incomplete grammar of the English language runs over 1,700 Perhaps when it comes to natural language processing and related fi elds, we re doomed to complex theories that will never have the elegance of physics equations. But if that s so, we should stop acting as if our goal is to author extremely elegant theories, and instead embrace complexity and make use of the best ally we have: the Unreasonable Effectiveness of of us, as an undergraduate at Brown Univer-sity, remembers the excitement of having access to the Brown Corpus, containing one million English Since then, our fi eld has seen several notable corpora that are about 100 times larger, and in 2006, Google released a trillion-word corpus with frequency counts for all sequences up to fi ve
language models that are used in both tasks consist primarily of a huge data-base of probabilities of short sequences of consecutive words (n-grams). These models are built by counting the num-ber of occurrences of each n-gram se-quence from a corpus of billions or tril-lions of words. Researchers have done a lot of work in estimating the prob-
Domain:
Source:
Link to this page:
Please notify us if you found a problem with this document:
{{id}} {{{paragraph}}}