Language Games in Communities of Practic

Language Games in Communities of Practice Learning happens within communities of practice, bounded by the shared pursuit of goals. Communities of practice tend to overlap with discourse communities, distinctive language games in which meaning-making happens. In order to play a role within a community of practice, it is necessary to have an identity. Discoursal identities are one form of learner identity. In this paper, we demonstrate a method for building an unsupervised model of the language games played by a discourse community. Comparing the distinctive way words interact with meaning in a discourse community with comparable games in the broader culture allows us to articulate the distinctive outlook shared by the community. Sometimes this includes systematic bias. We compare the way various adjectives and roles (programmer; teacher) attach to gender within Computer Science. 

Method: 

  1. Start with global word vectors as “global meanings.” Train word vectors on a corpus. These are “local meanings” Then choose several anchor terms (man; woman; boy; girl) and define a transformation from local meanings to global meanings. 

  2. Hacker news corpus

  3. Transcripts of Computer Science classes

  4. Dialogue captured from live computer science classes. 

  5. Select the 100 most common adjectives, as well as a bunch of roles. For each, compute the distance using (L2?) from “man” and “woman” or some triangulated anchor point. 

How is this useful in thinking about epistemic/ontological iconicity? And how is epistemic/ontological iconicity different from rhematization?

NEXT STEPS FOR RESEARCH

It seems like the task proposed in Danescu et al is a good baseline: can I predict who will stay and who will leave? 

So far, I haven’t been very successful in doing so. One strategy going forward might be to retrain using a different strategy. 

Currently, I’m training just using Google’s default word2vec settings. 

In terms of my argument, first I need to show that the embedding works. I can validate it using the Danescu et al task.  Then I can explore what the embedding means. 

A few possibilities:

  1. How do I generate co-occurrence counts from a corpus of weighted posts?
  2. How do I