Note permalink
note-dat...
## Data-Derived Competency Ontology **The opportunity:** If competency category structure was assumed rather than discovered, then the right question is whether the structure can be derived from performance data directly — without pre-specifying categories. Modern data sources (FDM, simulator recordings, communication transcripts, eye-tracking, physiological sensors) make this tractable in ways it wasn't in the 1980s or early 2000s. **What "data-derived" means here:** Instead of expert-consensus categories → OBs → measurement, invert: collect dense behavioral observations → apply dimensionality reduction or clustering methods → discover latent structure → use that structure as the measurement basis. The categories emerge from the data rather than being imposed on it. **Why this matters for the measurement problem:** A data-derived ontology could: - Reveal the actual covariance structure of expert behavior (which categories, if any, are genuinely independent) - Identify dimensions that expert-consensus panels missed or merged inappropriately - Scale measurement bandwidth beyond what's achievable through manual OB grading - Enable assessment coverage at Dreyfus Stage 5, where OB instruments currently break (the behavioral signature is in the data even if the expert can't articulate it) **Connection to AI decomposability angle:** This is a different framing of the same underlying possibility. The question isn't "can we decompose Stage 5 cognition into rules?" — it's "can we build a measurement infrastructure that doesn't require articulation at all?" FDM data from expert pilots doesn't require their introspection. **Human-readability constraint:** A data-derived ontology may not produce categories that are naturally human-interpretable. This is acceptable for automated assessment legs, but creates a challenge for instructor-mediated feedback — the instructor needs to be able to explain something actionable to the trainee. The gap between machine-derived structure and pedagogically deliverable feedback is a real constraint, not a temporary limitation. --- **Open questions:** 1. Has anyone attempted unsupervised discovery of competency structure from aviation FDM or simulator data at scale? What methods were used and what did they find? 2. What is the minimum human-readability requirement for a data-derived competency category to be pedagogically useful? Is there a principled way to constrain the discovery process to produce interpretable outputs? 3. How much performance data (hours of FDM, number of expert pilots) would be required for stable dimensionality discovery? What are the practical data acquisition constraints in the aviation training context? 4. If a data-derived ontology yields a different structure from ICAO's 9 categories — say, 6 dimensions or 14, or a hierarchical structure — how would that interact with the regulatory framework? ICAO PANS-TRG is normative; operators must comply. Is there a path from empirically derived structure to regulatory uptake? 5. What is the relationship between data-derived competency dimensions and individual differences in learning path? If the structure is discovered from aggregated expert performance, does it capture the variation in how novices develop through Dreyfus stages?
Permalink: http://localhost:3000/sensemaking?noteId=note-data-derived-competency-ontology