NLU Lit Review Feedback
NLU Lit Review Feedback
Comments:
====================================================================== Is the general problem/task definition clearly articulated?
Yes, it’s clear, I think: model linguistic change and user lifecycles using distributed representations instead of traditional language models. This is an interesting, innovative idea, and you’re building on your prior work, which is the best way to achieve deep, lasting results. Excellent! May 22 is very soon! If that doesn’t work out, check out the EMNLP workshops: http://emnlp2018.org/workshops/ The deadlines are all in July, I believe. W-NUT and WASSA both seem appropriate at first blush.
====================================================================== Do the summaries articulate the major contributions of each article and identify informative similarities and differences between the papers?
On my reading, the lit review is focused on ways that we can improve your word2vec-based approach and add nuance to the findings. There’s another direction that is worth thinking about: different techniques for measuring the underlying ideas related to linguistic alignment and community conventions. A bunch of the additional work included below heads in this direction. Srivastava et al. and Goldberg et al. are perhaps particularly relevant because of their focus on social/professional outcomes. For how vectors change over time, the HistWords project is highly relevant.
====================================================================== Are there existing, high-quality publicly available datasets for working in this area? If not, is data availability likely to be an obstacle to progress on the project?
I believe the project is in good shape here. You have the Hacker News dataset, and we have a number of other longitudinal community datasets that would also support the relevant work. Let me know if the Hacker News dataset starts to feel confining, and we can explore other options.
====================================================================== Are the future work ideas feasible given the available data and time allotted? If yes, do any stand out as really good paths to follow? If not, can we formulate a workable path forward using the ideas in this lit review?
I am curious about your word2vec approach to modeling user lifecycles on Hacker News, and the finding that users start to drift from the norms as soon as they join. This would seem to suggest that there just aren’t any norms to adhere to, but perhaps I am missing something about the findings. In any case, do you know what the (updated) word2vec spaces look like over this period? If, for example, the community gets larger over the time period, then the embedding space is likely to get denser, and there might be other global properties that emerge like that. Could this be getting in the way of seeing how norms evolve? I am really keen to see whether Mittens might prove a softer-touch method for performing updates over time using only text!
====================================================================== Other relevant work, or tips on where to look for relevant work?
HistWords https://nlp.stanford.edu/projects/histwords/
Danescu-Niculescu-Mizil, Cristian, Lee, Lillian, Pang, Bo, and Kleinberg, Jon. 2012. Echoes of power: language effects and power differences in social interaction. In Proceedings of the 21st World Wide Web Conference, 699–708. New York: ACM.
Doyle, Gabriel, Yurovsky, Dan, and Frank, Michael C.. 2016. A robust framework for estimating linguistic alignment in Twitter conversations. In Proceedings of the 25th International World Wide Web Conference, 637–648. Republic and Canton of Geneva, Switzerland: International World Wide Web Conferences Steering Committee.
Goldberg, Amir; Sameer B. Srivastava; V. Govind Manian; Will Monroe; and Christopher Potts. 2016. Fitting in or standing out? The tradeoffs of structural and cultural embeddedness. American Sociological Review 81(6): 1190-1222. Srivastava, Sameer B.; Amir Goldberg; V. Govind Manian; and Christopher Potts. 2016. Enculturation trajectories: language, cultural adaptation, and individual outcomes in organizations. Management Science.
Action Items:
MODELING USER LONGITUDINAL CHANGE — Start Mittens training asap. — Create a model that performs better on the initial task.
MODELING COMMUNITY CHANGE — My focus will be on establishing cultural fit; this is particularly important for learning. — Could I use a measure of cultural embeddedness? — Or could I show that distributional representations enhance findings on cultural fit.
— I could use LIWC categories as anchors — HistWords could be used to characterize community language change? https://nlp.stanford.edu/projects/histwords/ — Use visualization technique…
— One possible way of validating word change is by aligning with known changes
Are there even community norms? — I could show that there are by computing the relationship between upvotes and distance from community!
https://news.ycombinator.com/item?id=15507821 How to find a user’s favorite comments: https://news.ycombinator.com/favorites?id=tptacek&comments=t
I’m thinking this paper is really about using distributional word representations in this task?