CURRENT STATE
Study at a glance
Data export required—Malaysia pages selected
—sentence units
—pages represented in sentences
500planned human QA sample
Pipeline
1
Construct definitionComplete
2
Corpus ingestionComplete
3
Malaysia preprocessingComplete
4
Corpus QAHuman review required
5
Lexical-universe annotationnext
6
Lexical candidate extractionnext
8
Human psychometricslater
Research logic
CorpusNatural language before theory seeding
Lexical candidatesWords and short expressions with source context
Lexical familiesSemantic clustering, not factors
PsychometricsOptional downstream validation
Study principle
“Townscape is not the sum of visible objects.”
The computational stage is designed to recover the vocabulary people use to describe the visible scene, preserving object attributes, relations, appearance, evaluation, metaphor, and local expressions before any latent structure is proposed.
Upstream ingestion snapshot: 68,444 source pages, collection run WV-2026-08. The dashboard treats the local DuckDB export as derived data and Parquet as authoritative.