The brief
The study of slavery in war draws on armed conflict databases, humanitarian datasets, tribunal records, archival sources and qualitative case studies. ACLED and UCDP publish REST APIs. The ILO publishes surveys. ICTY case records sit in tribunal archives. Wikidata answers SPARQL. They differ in format, temporal resolution, geographic reference frame and disciplinary vocabulary.
That heterogeneity is not an accident to be tidied away. It reflects a genuinely multi-disciplinary field, and each source is shaped correctly for the community that produced it. The cost lands on anyone trying to work across them. A researcher comparing patterns of forced labour across civil wars has to assemble and reconcile sources that were never designed to be used together, and has to do it again from scratch for the next question.
CDISAW is not another dataset. It is the layer that makes the existing ones answerable together, holding 23.5 million linked event records drawn from 53 sources behind one query interface and one shared ontology.
The insight
The organising decision was to make classification the entry point rather than a filter applied at the end.
Most spatial and temporal platforms start from where and when, then let you narrow by category. Here it runs the other way. A researcher first selects cells from a 10 by 11 grid of slavery typology against conflict context, and only then constrains time, place and actor. That ordering is a claim about the field: the research question in this domain is the intersection of a kind of slavery with a kind of conflict, and everything else qualifies it.
Building the interface around that claim has a practical payoff. Because the typology is explicit and shared, two researchers asking about forced labour in occupation contexts are demonstrably asking the same question, even when their underlying sources differ completely.
The second decision was to treat a query as a citable object. Serialising the full query state to a persistent CDQ code means a paper can reference the operation that produced its result, not merely the datasets it drew on. That is a small piece of infrastructure addressing a real gap, since a reader can currently check your sources but not your question.
What it does not do
It does not resolve contested classification. Assigning a dataset to cells in the matrix is an interpretive act, and reasonable scholars will disagree about whether a given case is forced marriage or sexual slavery, or whether a conflict counts as civil war or insurgency. The platform makes those assignments explicit and inspectable, which is better than burying them, and it does not make them correct.
It inherits every bias in its sources. Historical atrocity data is recorded unevenly, and the unevenness is not random: better-documented theatres produce more records, and an absence of events in a region is far more likely to mean an absence of recording than an absence of slavery. Counts from this platform describe the evidence base, not the past.
Place-name disambiguation across centuries remains unsolved, actor timelines and a GraphQL endpoint are planned rather than built, and scheduled synchronisation against ACLED and UCDP is not yet running. The Bosnia pilot will be the first full-scale demonstration that the cross-layer query design holds up against a real research programme.










