Original, replication-grade datasets on Turkish courts and parliaments — built for cumulative, multi-user computational research and released openly.
DATASET · VERSION 1.0
A speech-level corpus of Turkish Grand National Assembly proceedings covering the full period of competitive multi-party politics in Turkey: 2,418,313 speaker turns across five chambers, of which 702,503 are linked to specific MPs and parties via a roster of 10,740 MP-term records with gender annotations.
An analytical core of 136,118 high-confidence floor speeches (49.4 million words, 2,868 MPs, 26 parties) is provided for quantitative analysis. The corpus extends the Turkronicles OCR archive (1950–2018) with new segmentation and linkage layers, and adds 314,899 turns extracted from official transcripts for the 27th term (2018–2023). Speaker boundaries are detected with a ten-pattern unified parser with decade-aware OCR normalization.
DATASET · FORTHCOMING
A dual-corpus database spanning the Court's full operational history: 5,513 norm review decisions and 16,626 individual application decisions with claim-level outcomes, full decision texts and validated classifications, accompanied by a companion dataset on all 138 justices who have served on the Court.
Built to make judicial disagreement and avoidance measurable at population scale across six decades and four constitutional junctures, the database is released with a validation suite for independent replication and extension. It will be published openly when the companion article — currently revise & resubmit at the Journal of Law and Courts — is accepted.