Note to Chris Potts
Note to Chris Potts HI Chris,
I’m also lukewarm about the final paper, so I think it’s a fair assessment. I spent too much time trying to finish additional analyses in time, and not enough time presenting what I had done in the paper. Even though I don’t have fundamentally new results, I would like to defend my work for this class. Rather than considering the course final projects as potentially self-plagiarizing, I think it makes sense to think of them as milestones toward having something worth publishing. From that perspective, I think I made a lot of progress in this class.
Code changes (comparing the final commit in previous quarter to : https://github.com/cproctor/language_games/compare/7ed97a38ed602a57e4a9d1914d439e63d5f231ca…master
Features/analyses implemented:
-
Predicted comment popularity using other features, as a binary classification problem with five different cases: at least 1/3/5/10/20 upvotes. No inspiring results so far.
-
Wrote code to generate co-occurrence matrix, generated co-occurrence matrices for the corpus in preparation for retrofitting.
-
Analyzed previous analysis, re-wrote most of it, and re-ran it. In particular, the way I was indexing into monthly embeddings to create comment bag-of-words representations was incorrect, invalidating all my prior results.
-
Analyzed error in Mittens training, wrote a proof showing it should scale with corpus size.
-
Implemented (but did not use) a distance metric for comments <-> GloVe embeddings. Studied GenSim source code for approaches.
-
Provisioned GCloud, then AWS server, set up long-running GloVe/Mittens training (estimated 12 days, currently 60% complete.)
New literature:
- GloVe, retrofitting
- Approaches to studying linguistic change in subspaces, and validating against external evidence: WEAT, HistWords.
Hours spent: [Graph]
What do you think about me taking an incomplete, and then submitting a revision within a month?