CS224U Notes
CS224U Notes Laurens van der Maaten Google Tech Talk Look into: Parallel coordinates
Could I use PCA to reduce the dimensionality of word points to care about their place on a couple of relational axes?
I may have a similar challenge to the challenge presented in “Visualizing Data Using t-SNA" isomap Locally linear embedding
t-Distributed Stochastic Neighbor Emedding
Linguistic distance does not satisfy the triangle inequality. So use multiple maps, and assign an importance weight to each point. Define the similarity between two points as the weighted sum of their similarities in the multiple maps. Check out his d3 tool for exploring maps.
Can you control how the relationships are distributed across multiple maps? Gotta read the paper. If there some way of creating multiple maps for various relational axes, then perhaps I could minimize some kind of objective function such that the maps were maximally informative about word usage that will shape future participation. (In other words, I’m looking for ideologies which people use to establish power relations.)
Probably, it works a bit like circuit board layout, developing relations wherever they are most valuable.
IDEA for an art piece: Take all the writing Zuz and I have ever done, and compute TFIDF for words, using TF from our own writing and IDF from global values. This will help identify characteristic ways we use language. Then project this onto a large wall using t-SNE. Filter out words based on overlap in their bounding boxes, always choosing the word with the higher frequency count. This should produce an image
For the bakeoff, what about using gigaword and imdb as multiple maps? How could they be combined?
FOR LANGUAGE GAMES PAPER Emphasize that I use word2vec because it’s an online algorithm.
Is there some sense in which GloVe is an autoencoder? Yes, you could consider the probability ratios -> PMI as learning some kind of identity function, with the GloVe vectors as the intermediate representations. (Specifically, a score-matching autoencoder)
=================================== Idea: What if I retrofitted my embedding instead of online training? Or if I used the discourse to retrofit concepts?
PAPERS TO READ Retrofitting Word Vectors to Semantic Lexicons Improving Word Representations via Global Context and Multiple Word Prototypes Multi-Prototype Vector-Space Models of Word Meaning