Rob Voight Meeting
Rob Voight Meeting What are the gradient boosting not standard. standard NLP would be random forest, svm, perceptron (simple feed forward)
- scikit-learn
- neural stuff
How to get it working in neural models.
Goto: Use coreNLP (python wrapper) to parse it all. That returns a JSON dict. Inherent in it is sentence, word segmentation. POS tags, dependency arcs, Throw out (or not) stop words. Traditionally throw out stopwords.
Paper: What’s the best way to build a baseline?
Trigram SVM is very good. Using binary features is already almost good enough.
Log-transformed or proportional
Use dependency parsing —> arc
How to use dependency trees as features. Dependency ngrams.
Could we take advantage of the multiple headlines?
Autoencoder to reduce the dimensionality of the model. Topic models.
Paragraphs are asking about discourse structure; people do not do that in a principled way. You could imagine putting emphasis on discourse conjunctions.
CNNs?