Committing to algorithmic modeling
What terms would be appropriate to describe Breiman’s perspective and the perspective of mainstream statistics?
Searched for 1 pattern, read 1 file (ctrl+o to expand)
⏺ Breiman’s own terminology, laid out right at the start of the paper (lines 43–86), is exactly the right vocabulary here:
Breiman’s two terms
- “Data Modeling Culture” — mainstream statistics (his estimate: 98% of statisticians). Assumes the inside of the “black box” generating the data is a known stochastic model: response = f(predictors, parameters, random noise) — e.g., linear regression, logistic regression, the Cox model. The model’s form is treated as a hypothesis about how nature actually works. Validated by goodness-of-fit tests and residual analysis. Goal: estimate parameters, do inference, interpret coefficients as meaningful.
- “Algorithmic Modeling Culture” — the minority approach he champions (his estimate: 2%, “many in other fields” — computer science, ML). Treats the inside of the box as unknown and possibly too complex to model explicitly. Doesn’t assume a generative stochastic model; instead searches for a function f(x) — decision trees, neural nets — that predicts well. Validated purely by predictive accuracy on held-out data, not by fit to an assumed mechanism.
Other terms used in the discussion that map onto this same split
- “prediction culture” vs. (implicitly) an inference/explanation culture (line 2012) — Efron’s reframing, emphasizing goals (predict vs. explain) rather than methods.
- generative/model-based vs. algorithmic/black-box — common shorthand in later literature for the same divide.
- interpretable parametric models vs. non-parametric/black-box learners — a methodological framing (data models nearly always parametric; algorithmic models are typically non-parametric).
For your own use, I’d recommend sticking close to Breiman’s own pair — “data modeling culture” vs. “algorithmic modeling culture” — since that’s the paper’s actual organizing vocabulary, and reach for “generative/parametric” vs. “black-box/predictive” only if you want more generic statistical-language synonyms outside the paper’s own framing.
✻ Cogitated for 19s
❯ To what extent does the field of learning sciences or computer-supported collaborative learning adopt the algorithmic modeling culutre?
⏺ Short answer: learning sciences/CSCL is historically rooted in the data modeling culture, but it has a substantial and growing pocket — educational data mining and (parts of) learning analytics — that has explicitly adopted the algorithmic modeling culture, and the resulting tension within the field maps almost exactly onto the Breiman/Efron-Cox debate you’ve been reading.
The core of the field: data modeling culture
Learning sciences grew out of psychology, education, and sociocultural theory, and its dominant epistemics still look like Breiman’s 98%: researchers posit a theoretical mechanism for how learning happens (e.g., conceptual change, distributed cognition, productive struggle) and then build models — ANOVA, regression, mixed-effects models, structural equation models, and qualitative methods like interaction analysis or grounded theory — to test whether that mechanism holds. Validation is by theoretical coherence and statistical fit, not raw predictive accuracy on held-out data. CSCL specifically leans even harder this way: its signature methods (conversation analysis, uptake/transactivity coding, epistemic network analysis) are explicitly built to explain collaborative knowledge construction, not to predict an outcome variable. This is the same posture as design-based research and conjecture mapping more broadly — mechanism and interpretability are the point, which is presumably why that’s your own design idiom too.
The algorithmic-modeling pocket: EDM and learning analytics
Starting roughly with the founding of the EDM conference (2008) and LAK (2011), a sizable adjacent community treats the inside of the “learner box” as unknown and complex, and just fits predictive functions: decision trees, random forests, gradient boosting, and deep nets (LSTMs, transformers) on clickstream/log data to predict dropout, performance, or affect. Validation is held-out predictive accuracy (AUC, RMSE), exactly Breiman’s criterion — not goodness-of-fit to a theorized learning mechanism.
The clean case study: knowledge tracing
This is probably the most legible instance of the Two Cultures split playing out inside learning sciences itself. Bayesian Knowledge Tracing (Corbett & Anderson, 1994) is a textbook data model — a hidden Markov model with four interpretable parameters (prior knowledge, learn rate, guess, slip) meant to represent a real cognitive process. Deep Knowledge Tracing (Piech et al., 2015) is an explicit algorithmic-modeling move: an LSTM black box that improved prediction accuracy but discarded interpretability. The backlash (Xiong et al.’s reproducibility critiques, and a subsequent wave of “interpretable knowledge tracing” papers trying to recover explanatory structure from the black box) is essentially a rerun of Efron’s complaint: the predictive gains are real, but the field had no theory for judging what kind of bias/structure the black box was smuggling in, or whether its internal representations corresponded to anything cognitively meaningful.
Net assessment
If you use Breiman’s own population estimate as a model: learning sciences/CSCL proper is still closer to his 98%/2% split in favor of data modeling, especially in CSCL’s qualitative-discourse tradition. But EDM and learning analytics — increasingly visible, well-funded, and influential in ed-tech — sit much closer to parity, or even tip toward algorithmic modeling when the deliverable is a practical system (early-warning dashboards, adaptive tutoring) rather than a theoretical claim. The unresolved friction in the field is exactly Efron’s: predictive black boxes work, but learning sciences (unlike, say, ad-click prediction) usually wants the model to also produce warranted claims about how people learn — and there’s still no settled theory for extracting trustworthy explanatory content from a biased, high-variance-reducing algorithmic model trained on log data.
❯ Read ~/Desktop/causal_models.txt. To what extent does this paper offer a framework for theoretically-aligned (and theoretically-generative) algorithmic modeling?
⏺ This paper is explicitly framed as an alternative to Breiman’s binary — its whole pitch is a “middle space” that escapes “a largely futile debate between the two cultures” (line 1452). Worth separating your two criteria, because the paper succeeds much more on one than the other.
Theoretically-aligned: strong
This is the paper’s central contribution. A DAG forces the analyst to commit, before touching data, to explicit claims about which variables cause which others, which are confounders, and which are mere proxies for unmeasured constructs (motivation, self-regulation). That directly attacks the “clicks to constructs” problem (line 1303) — the gap, endemic to learning analytics, between what’s logged (clickstream “data exhaust”) and what theory claims is actually happening cognitively. By requiring the causal structure to be specified first, the statistical model — whatever fits it (regression, SEM, in principle even an ML estimator) — is disciplined to answer to a theoretical claim rather than just chase predictive accuracy on whatever variables happen to be in the log. This is close to a direct answer to Efron’s complaint about black-box methods having “so little theory to fall back upon” for judging their bias: the DAG is exactly the missing piece of theoretical scaffolding that tells you what kind of structure a model should respect.
Theoretically-generative: limited, and the paper itself flags this
There’s a real generative mechanism gestured at — causal discovery (lines 843–876): algorithms that infer a causal structure (or class of equivalent structures) from observed conditional-independence patterns in data, exploratory rather than confirmatory, with “uncertain edges” pointing toward where new data or experiments are needed and feeding back into theory. That actually would be theoretically-generative and algorithmic in Breiman’s sense — letting the data, via an algorithm, propose new theoretical structure rather than just testing a pre-specified one.
But the authors explicitly decline this path for their own demonstrations: “Rather than following the causal discovery model of Brokenshire and Kumar, here we will explore a method more closely related to the human centred approach…” (line 880). Their two worked examples (Zimmerman’s SRL model, the RWA example) both start from an existing proto-theory and use the DAG apparatus to formalize and test it — Borsboom’s TCM steps 3–4 (formal model → check explanatory adequacy), not step “discover new explanatory principles from data.” The generativity they do claim is modest: building the DAG can reveal previously implicit assumptions, surface competing models, or expose missing variables (lines 1403–1408) — generative of clarity, not generative of new causal hypotheses pulled from data.
Where it falls short of fusing with Breiman’s algorithmic culture
Note also that the statistical machinery they actually deploy — SEMs, path equations, regression-style structural equations (lines 820–841) — is squarely inside Breiman’s data modeling culture, not his algorithmic one. The paper never shows how to wire a genuinely algorithmic/black-box estimator (random forest, neural net) into a causal-DAG scaffold — the kind of fusion that exists in the broader causal ML literature (causal forests, double/debiased ML, targeted learning) isn’t discussed. ML appears only at the margins, in cited related work (e.g., Barthakur et al.’s ML classifier for reflective-writing depth) rather than integrated into their own framework.
Net assessment: this is a well-developed framework for theory-aligned modeling — forcing algorithmic/statistical work to answer to explicit causal claims — and only a gestural, aspirational framework for theory-generative algorithmic modeling, since the one mechanism that would deliver that (causal discovery) is named, attributed to 15-year-old failed-to-gain-traction work, and then set aside in favor of the more confirmatory, theory-first approach the authors actually demonstrate. It’s better read as a rigorous bridge between proto-theory and testable models than as a recipe for letting algorithms generate new theory from data.
Read Sandoval’s conjecture mapping (/Users/chrisp/Zotero/storage/SPE83ZKG/Sandoval\ -\ 2014\ -\ Conjecture\ Mapping\ An\ Approach\ to\ Systematic\ Educ.pdf). To what extent is conjecture mapping compatible with Kitto et al’s proposal to start with causal hypotheses? That is, you could use conjecture mapping to structure your inquiry and generate hypotheses, and then use Kitto et al’s approach to evaluate them empirically?
Yes, in broad strokes — and the compatibility is closer than you might expect, because both papers are solving the same methodological problem (warranted causal attribution in messy educational settings without RCTs) and structure their answer the same way: stages connected by hypothesized causal arrows. But there are three real points of friction worth flagging before you build a pipeline on it.
Where they line up structurally
Sandoval’s conjecture map is already, visually and logically, a coarse causal diagram: embodiment → mediating processes → outcomes, with design conjectures (“if learners engage in this task/participant structure with these tools, this mediating process will emerge”) and theoretical conjectures (“if this mediating process occurs, it will produce this outcome”) as the two arrow-types. This maps almost directly onto Kitto et al.’s Theory Construction Methodology staging (proto-theory → formal model → check explanatory adequacy): a conjecture map is essentially a proto-theory, already partway formalized into boxes and arrows, but — by Sandoval’s own admission — without a calculus for estimating effects from data. That’s exactly the gap Pearl’s apparatus (DAGs, do-calculus, identification strategies) is built to fill. So your proposed division of labor — conjecture mapping to generate and structure the hypotheses, causal modeling to formally test them — tracks the seam Kitto et al. themselves describe between “proto-theory” and “formal theory ready for testing.”
Genuine complementarity
Conjecture mapping is strong precisely where Kitto et al. is thin: it has a developed practice for generating and revising conjectures through engagement with actual enactments — interaction analysis, artifact analysis, the Figure 2 → Figure 3 revision in Sandoval’s own example, where close observation of classroom discourse surfaced an unanticipated mediating construct (norms for persuasion/consensus) that got added to the model. Kitto et al. is strong precisely where conjecture mapping is thin: once you have a structural hypothesis, they supply the actual machinery (backdoor criterion, conditional independence tests, SEM estimation) for testing it against data rather than just asserting it’s plausible. Chaining them — qualitative, small-N conjecture-mapping cycles to generate a DAG’s nodes and edges, then larger-scale causal-DAG testing to confirm it — is a coherent and fairly natural extension of how Sandoval already describes design research scaling up (“trajectory,” p. 32) from exploratory small studies toward generalizable, exportable explanatory claims.
Three frictions
- Acyclicity. DAGs forbid loops. Kitto et al. had to artificially “epoch-slice” the looping Zimmerman SRL proto-theory into separate time-indexed copies (S0→B0→E0→S1…) to make it DAG-compatible, and they admit this is a real cost. Sandoval’s conjecture maps have no such constraint — mediating processes are explicitly iterative within a single enactment (a teacher introduces a norm, students take it up, this reshapes the next round of interaction, as in Ryu & Sandoval’s findings). Translating a real conjecture map into a DAG will usually demand the same kind of awkward epoch-slicing, which risks losing exactly the in-the-moment, recursive process Sandoval cares about capturing.
- Variables vs. practices. This is the sharper issue. Sandoval explicitly argues, via Bereiter and Cronbach, against treating designed elements as variables: “‘Collaborative inquiry’… is not a variable but a complex form of activity” (p. 29–30), and warns that decomposing designs into variables-and-interactions “tears the design apart… moving further and further away from the design itself.” A DAG, by contrast, requires every node to be a well-defined measurable variable that can enter a joint distribution. Forcing conjecture-map boxes like “discursive practices” or “epistemic reflection” into DAG nodes is structurally the same move as the “clicks to constructs” reduction Kitto et al. themselves warn about in other systems — there’s a real risk the formalization step erases the thing conjecture mapping exists to protect.
- Different theories of causation, not just different tools. This is the deepest tension. Sandoval grounds conjecture-mapping’s causality explicitly in Maxwell’s (2004) “scientific realist” view — causal processes, observed and mechanistic, traced in particular settings — and contrasts this with the Humean “regularity” view (causation inferred only from repeated co-occurrence across cases). He picks the process view because he thinks the regularity view is inadequate for design research. Pearl’s apparatus, however much more sophisticated than naive correlation, is still fundamentally a regularity-type formalism: causal claims are licensed by conditional-independence patterns across a population of measured units. So your pipeline doesn’t just move from “less rigorous” to “more rigorous” — it crosses an actual epistemological boundary, from single-case mechanism-tracing to population-level probabilistic inference. That’s not necessarily a contradiction (Sandoval doesn’t reject regularity-based methods outright, just argues they’re insufficient alone), but it means the DAG-testing phase is answering a different kind of causal question than the conjecture-mapping phase asked, and you should be explicit about that shift rather than treating it as a seamless escalation in rigor.
Net assessment
Compatible as a two-phase research design, and arguably an attractive one — it matches how design research is supposed to scale (small-N theory generation → larger-scale, more formal testing) and gives conjecture mapping the calculus it admits it lacks. But the translation step from conjecture map to DAG isn’t free: it forces acyclic time-slicing of inherently recursive processes, forces reduction of rich situated practices into discrete variables (the very move Sandoval warns against), and quietly swaps a process-based notion of causation for a population-regularity one. Treat the conjecture-map-to-DAG step as real theoretical work in its own right, not just a formatting change.