Dataset and variable discovery

Find the dataset, then find the variable inside it

Dataset catalogues usually stop at the study description. That is rarely the question a researcher actually has, which is not “which surveys exist?” but “which survey measured this, on whom, and can I get it?”

Secondary analysis often begins with a construct rather than a study: trust in institutions, household food insecurity, volunteering frequency. Searching a repository by study title answers a different question, so the usual route is to open promising studies one at a time and read each codebook to see whether the measure is there at all.

miscite indexes archives at both levels. A search can match a study description or the text of the individual variables inside it, so a construct-shaped question can return the surveys that actually carry the measure. Because a great deal of good social science data is not simply downloadable, every result states what it takes to obtain it.

Capabilities

What dataset search can do

The index is built around the hierarchy the data already has — archive, dataset, variable — so a question can be asked at whichever level it belongs to.

Semantic, lexical, or hybrid matching

Describe a construct in your own words and retrieve datasets whose descriptions or variables mean the same thing, or fall back to exact keyword matching when you know the instrument's wording. Hybrid retrieval runs both and combines the rankings.

Criteria you can combine

Add rows joined by and/or, each scoped to a field: anywhere, title or acronym, description, producer, universe or unit, or variable description. Scoping a term to variable text is what separates “the study is about this” from “the study measured this”.

Geography, unit, and access filters

Narrow by country using ISO 3166-1 codes, by unit of analysis, and by access level. Filtering on access is often the practical constraint: a perfect dataset behind an application you cannot complete in time is not a usable dataset.

Variables in context

Open a dataset to search within its own variables, or open an archive to see what it holds. From an article in literature search, a related-datasets action looks for the data behind the paper you are reading.

Process

How a dataset search runs

The sequence is deliberately the reverse of browsing a catalogue: start from the measure, arrive at the study.

  1. Describe the measure, not the study

    Enter the construct you need. Semantic matching does not require you to guess the archive's vocabulary, which is useful when the same idea is called wellbeing in one survey and life satisfaction in another.

  2. Scope the criteria

    Add rows to require a producer, a universe, or a term that must appear in variable text. Filter by country, unit of analysis, or access level when those are real constraints.

  3. Read the access terms

    Each result states whether the data is open, needs registration, needs an application, is restricted, or is available on request. Some archives state nothing, and that is shown as unstated rather than assumed to be open.

  4. Look inside the candidates

    Open a dataset to search its variables directly and confirm the measure exists, its wording, and its coverage before you invest in an access request.

Best fit

Who this fits

Dataset search is aimed at people doing secondary analysis rather than collecting new data.

Secondary data analysts

Establish quickly whether an existing survey already measures your construct, and on which population, before designing around the assumption that it does not.

Graduate students scoping a project

Survey what data exists for a question before committing a chapter to it, including the practical question of what can actually be obtained within a degree timeline.

Methods and replication work

Trace which archives hold a study, compare how different surveys operationalize the same construct, and locate the replication data behind a published paper.

Boundaries

What it does not claim

This is dataset discovery, not a data warehouse. Being clear about the boundary saves a wasted search.

  • miscite does not host or redistribute the data. Results link to the archive that does, and obtaining the data remains subject to that archive's own terms.
  • Variable-level detail exists only where the archive publishes it — through a codebook, a data dictionary, a processed statfile's value labels, or an enumerated questionnaire. Roughly half the indexed sources publish it; for the rest, matching is limited to the study description.
  • Access terms are read from what the archive states. Where an archive states nothing, the result says access is not stated rather than guessing, and the archive remains the authority.
  • Coverage is social science data, not every dataset in existence. A source is indexed only when its records carry, or can be made to carry, the detail needed to answer a variable-level question.

FAQ

Questions about dataset search

How is this different from searching a data repository directly?

A repository search matches study-level metadata within that one archive. miscite searches across archives and, where the archive publishes them, matches the individual variables inside a study. That is the difference between finding studies about a topic and finding studies that measured it.

Can I search for a specific survey question?

Yes, where the archive publishes its codebook. Add a criterion scoped to variable description and enter the wording or the construct. Datasets whose variables match are returned, and you can open one to search within its variables and check the exact phrasing.

Does a result mean I can download the data?

Not necessarily, which is why every result states its access terms. A large share of indexed datasets require free registration, a formal application, or a direct request, and some are restricted. Access is granted by the archive, never by miscite.

Which archives are covered?

Social science repositories and microdata catalogues, including Harvard Dataverse, the CESSDA catalogue, the UK Data Service, GESIS, ODISSEI, the World Bank Microdata Library, the IHSN catalogue, OSF, Zenodo, and Figshare, alongside individually curated sources that publish data without a bulk interface. The Data sources tab lists every source with its coverage.

Do I need an account to search datasets?

No. Dataset and variable search is open, as literature search is. An account adds the surrounding workspace — library, alerts, and manuscript checks — not the search itself.

Related guides

Dataset search sits next to the literature it belongs to.

Find the data behind the question

Search datasets and the variables inside them without an account, and see what each archive asks before you commit to a study.