Deniz Yuret. Nov. 2012. Signal Processing Letters, IEEE. Volume: 19, Issue: 11, Pages: 725-728. DOI, URL. (Download the pdf, code, and data (1GB). Our EMNLP-2012 paper uses FASTSUBS to get the best published result in part of speech induction.).
Update of Jan 9, 2014: The latest version gets rid of all glib dependencies (glib is broken, it blows up your code without warning if your arrays or hashes get too big). It also fixes a bug: The previous versions of the code and the paper assumed that the log back-off weights were upper bounded by 0 which they are not. In my standard test of generating the top 100 substitutes for all the 1,222,974 positions in the PTB, this caused a total of 66 low probability substitutes to be missed and 31 to be listed out of order. Finally a multithreaded version, fastsubs-omp is implemented. The number of threads can be controlled using the environment variable OMP_NUM_THREADS. The links above have been updated to the latest version.
Abstract: Lexical substitutes have found use in areas such as paraphrasing, text simplification, machine translation, word sense disambiguation, and part of speech induction. However the computational complexity of accurately identifying the most likely substitutes for a word has made large scale experiments difficult. In this paper I introduce a new search algorithm, FASTSUBS , that is guaranteed to find the K most likely lexical substitutes for a given word in a sentence based on an n-gram language model. The computation is sub-linear in both K and the vocabulary size V . An implementation of the algorithm and a dataset with the top 100 substitutes of each token in the WSJ section of the Penn Treebank are available at http://goo.gl/jzKH0.
Full post...
January 09, 2014
January 07, 2014
NLP Course Slides
Here are the slides I used in my NLP course last semester. More to come...
Full post...
- Learning to map strings to meaning (introductory survey; Apr 12, 2014)
- Learning to map strings to probabilities (ngram models, smoothing; Oct 1, 2013)
- Interlude on probabilistic inference (Bayesian inference, common distributions; Oct 20, 2013)
- Learning to map strings to classes (text classification, generative, conditional, nonprobabilistic models; Nov 12, 2013)
Full post...
December 03, 2013
November 13, 2013
Learning to map strings to meaning
This is the talk I gave at the Univ 101 Freshman Seminar and the introductory lecture of my NLP class. The PDF Presentation contains slides with links to relevant material and papers.
Full post...
Full post...
November 09, 2013
A Language Visualization System
Ünal, Emre and Yuret, Deniz. In the Proceedings of the 2nd Workshop on Games and NLP (GAMNLP-13). November, 2013. İstanbul, Turkey. (Download PDF, Presentation, see our previous post, go to the workshop page).
Abstract:
A novel language visualization system is presented here that is capable of understanding some concrete nouns, visualizable adjectives and spatial prepositions in full natural language sentences to generate 3D scenes. The system has a rule based question answering component and it can answer spatial inference questions about the scene created by these sentences. This work is the first step of our greater vision of building a sophisticated text-to-scene and scene-to-text conversion engine.
Full post...
Abstract:
A novel language visualization system is presented here that is capable of understanding some concrete nouns, visualizable adjectives and spatial prepositions in full natural language sentences to generate 3D scenes. The system has a rule based question answering component and it can answer spatial inference questions about the scene created by these sentences. This work is the first step of our greater vision of building a sophisticated text-to-scene and scene-to-text conversion engine.
Full post...
Subscribe to:
Posts (Atom)
