LGA | Estimated 2025 human population (GHSL) | Sample size of gridded EAs | Sample size of building footprints |
|---|---|---|---|
Sentinel | |||
Gaya | 332,651 | 15,256 buildings | |
Nassarawa | 495,737 | 15,257 buildings | |
Gabasawa | 306,392 | 15,257 buildings | |
Non-sentinel | |||
Gezawa | 246,366 | 13 EAs | 1,485 buildings |
Ungogo | 703,042 | 36 EAs | 3,644 buildings |
Bebeji | 311,263 | 16 EAs | 1,810 buildings |
Dambatta | 397,691 | 20 EAs | 2,273 buildings |
Dawakin Kudu | 463,983 | 24 EAs | 2,592 buildings |
Dawakin Tofa | 398,075 | 20 EAs | 2,248 buildings |
Kiru | 470,655 | 24 EAs | 2,659 buildings |
Kumbotso | 1,181,763 | 60 EAs | 6,639 buildings |
Sumaila | 509,952 | 26 EAs | 2,952 buildings |
Takai | 334,477 | 17 EAs | 1,925 buildings |
Tarauni | 285,035 | 15 EAs | 1,602 buildings |
Tudun Wada | 396,206 | 20 EAs | 2,208 buildings |
Total | 6,833,288 | 291 EAs | 77,807 buildings |
3 Methodology
The Zero-Dose Baseline Survey tested several sampling and estimation methodologies. This report presents statistical results from the stratified multistage probability sampling design, which employs the same methods as outlined in World Health Organization (WHO)’s Vaccination Coverage Cluster Surveys: Reference Manual (WHO 2018). The other methodologies employed concern more recent, potentially more efficient approaches to sampling and estimation, and are covered in the companion report (Mindset 2026).
The population of interest are children aged 0-23 months in Kano State, Nigeria. The survey mode was face-to-face interviews using Computer-Assisted Personal Interviewing (CAPI).
3.1 Data Analysis
This report presents statistical results using the R {survey} package (Lumley, Gao, and Schneider 2024). The analyses employ multi-stage bootstrap replicates to calculate errors, ensuring that the complex sampling design and calibration adjustments are appropriately recognized in all statistical methods used.
The analyses also incorporate Vaccination Coverage Quality Indicators (VCQI) data cleaning standards using the {vcqiR} package. Developed by the WHO and Pan American Health Organization (PAHO), VCQI facilitates the computation of several WHO-standardized indicators related to immunization coverage and zero-dose prevalence (Rhoda 2021). The package also helps identify and resolve potential inconsistencies, missing values and entry errors before analysis, ensuring that only meaningful data points contribute to final estimation.
Sampling Error: Unless otherwise stated, all statistical tests, estimates of sampling variance, Margin of Errors (MOEs), Confidence Intervals (CIs), and confidence bounds are reported at the 95% confidence level. In this report, confidence bounds mean 1-sided confidence limits. The Upper Confidence Bound (UCB) is the upper limit of the \(100 \times (1 – \alpha)\)% confidence interval whose lower limit is 0%; the Lower Confidence Bound (LCB) is the lower end of the \(100 \times (1 – \alpha)\)% confidence interval whose upper limit is 100%. Alpha is usually set to 0.05, so we say that we are 95% confident that the population parameter falls above the LCB, or we say that we are 95% confident that it falls below the UCB. MOEs, CIs, and confidence bounds are provided to accommodate different audience preferences, although they express the same statistical concept. For estimates of proportions, CIs are calculated using the logit method widely used for complex surveys (Korn and Graubard 1998). This method differs from the simpler Wald intervals typically taught in introductory statistics courses, where a 95% confidence interval is constructed around a proportion \(p\) as \(p \pm\) MOE. The logit intervals avoid certain pitfalls of the Wald method, such as inaccuracies when proportions are near 0 or 1, but are otherwise interpreted in the same way.1 Specifically, they represent an interval within which we expect the true proportion to lie, acknowledging that in repeated sampling, we would be wrong about such claims approximately one out of twenty times when constructing CIs in the same manner.
Sampling Efficiency and Design Effects: Several statistical tables in this document report the Design Effect (DEFF), a measure used in survey sampling to quantify the impact of the survey’s sampling design on the precision of statistical estimates. In simple terms, the DEFF reflects how much the variability of an estimate is inflated due to the survey design compared to what it would be under simple random sampling. Complex survey designs—such as stratification, clustering, or unequal weighting—can increase the DEFF, making estimates less precise. A DEFF greater than 1 indicates that the design leads to more variability than simple random sampling, while a design effect of 1 means the design has no additional impact.
Multiple Testing and False Discovery Rate Control: In analyses presenting repeated statistical tests, adjustments are made to p-values and confidence intervals to control for the False Discovery Rate (FDR) using the Benjamini-Hochberg method (1995). This adjustment aims to control the expected proportion of false positives (incorrectly rejected null hypotheses) among all the hypotheses declared significant. Unlike methods that strictly control the family-wise error rate, such as the Bonferroni correction, FDR methods allow for a balance between discovery and error control. This makes them more powerful in contexts with large numbers of hypotheses, increasing the likelihood of identifying true effects while controlling the rate of false discoveries. Unless otherwise stated, all statistical tests for differences between two groups use a two-sided design-based \(t\)-test. Association tests between dependent dichotomous variable and an ordinal independent variable use a design-based Wald regression term test for a quasi-binomial model. All tests use a conventional \(\alpha = 0.05\) cutoff level for significance.
Adherence to Data Presentation Standards: This report adheres to the National Center for Health Statistics (NCHS) Data Presentation Standards for Proportions to ensure the reliability of statistical estimates (2017). Proportions with large sampling errors or potentially unreliable confidence intervals are flagged to indicate potential instability and prevent misinterpretation. Specifically, a single asterisk (\(*\)) in place of a statistic denotes a mild issue, whereas double asterisks (\(**\)) denote a more severe reliability concern. By following these standards, the report ensures consistency and minimizes the risk of drawing incorrect conclusions, particularly in domain analyses where results are highly disaggregated and more prone to statistical instability.
3.2 Survey Design
The survey employed a multi-stage stratified sampling design to ensure a representative sample. This design was administered to fifteen priority Local Government Areas (LGAs): three sentinel and 12 non-sentinel. These fifteen priority LGAs were identified by the government as priority zero-dose areas in Kano state. The three sentinel LGAs—Gaya, Nassarawa and Gabasawa—each comprise a domain of interest and the 12 non-sentinel LGAs combine to form a single comparison group. In the sentinel LGAs, the goal was to establish an accurate baseline of zero-dose children to allow for post-intervention comparisons.
The stratified multistage probability sample was implemented using a gridded sampling approach. This approach involves creating custom Enumeration Areas (EAs) to be used as Primary Sampling Units (PSUs) rather than using those provided by the official Bureaus of Statistics. These EAs are created using open-source geospatial data generated from satellite images combined with open-source data about administrative boundaries in Nigeria.
To implement the gridded sampling approach, Mindset has relied on guidance from Thomson (2022). The open-source GridEZ algorithm (Dooley 2019) was used to create PSUs with a target of 625 residents, using open-source estimates for 2025 from the Global Human Settlement Layer (GHSL). PSUs were made to follow the administrative boundaries of wards (sub-areas within LGAs). The areas were further constrained to ensure that PSUs were deemed entirely urban or entirely rural, based on GHSL urban coverage data. Due to imperfections in the greedy GridEZ algorithm, Mindset further collapsed small adjacent PSUs together to lower the variance in estimated PSU sizes. Large PSUs created by the GridEZ algorithm were likewise split apart to reduce the variance in PSU sizes.
Biostat Global Consulting provided the sample size calculations for a multistage cluster sample. The requirements for the sample were that it should detect a 10-percentage-point difference in the prevalence of zero-dose children aged 12–23 months with 90% power and 95% confidence. An effective sample size2 of at least 543 infants 12-23 would satisfy these requirements. Based on a presumed DEFF of 3.75, this implied that 291 EAs would be needed in each of the four areas of interest (three sentinel LGAs and one combined non-sentinel comparison group) under a typical multistage design.
3.2.1 Sentinel LGAs
Stage 1: In each sentinel LGA, a stratified simple random sample of building footprints was drawn without replacement within administrative wards, using Google’s open-source building footprint database as the sampling frame (Sirko et al. 2021). The research team chose to sample buildings directly rather than via a multi-stage design based on EAs, because the high sampling fraction in sentinel areas would have reduced the cost-efficiency advantages of cluster sampling. By avoiding first-stage sampling of EAs, the design could justify reducing the assumed design effect from 3.75 to 2.0 while preserving the desired precision with fewer interviews. This relaxation was further supported by empirical benchmarks from previous immunization coverage surveys conducted in Kano, which yielded LGA-level Diphtheria-Pertussis-Tetanus (DPT)-1 design effects of 2.9 in Gaya, 2.2 in Gabasawa, and 1.7 in Nassarawa—each below the 3.75 value assumed for the conventional multistage calculation (Kano State Ministry of Health 2014). This sampling strategy approximates a design in which 100% of EAs are included, though this remains a mathematical simplification that does not hold for variance estimation. Based on (1) a relaxed DEFF of 2.0, (2) a pilot-based estimate of 21 infants per 295 sampled footprints, and (3) a required effective sample size of 543, we sampled 15,256 building footprints per sentinel LGA to meet precision targets comparable to those in the non-sentinel component.
Stage 2: With sampled building footprints at hand, the field team proceeded to identify individual households within each sampled building. If a sampled building contained more than two residential units (e.g., an apartment complex), two units were randomly subsampled. In practice, many enumerators did not follow this protocol closely and only interviewed one household per building. A fuller account of this and other protocol deviations encountered during fieldwork — including geofencing, multiplicity capture, navigation and mapping challenges, and the reasons for non-captured footprints — is provided in Section B.3. Field staff presented neighborhood maps to households and asked respondents how many buildings belonged to their household. This information was used to adjust for multiplicity—i.e., the possibility that a household could enter the sample through multiple buildings.
3.2.2 Non-Sentinel LGAs
Stage 1: For the non-sentinel LGAs, the PSU was a gridded EA, generated using the open-source GridEZ algorithm. EAs were constructed to approximate 625 residents each and aligned with ward boundaries and urban/rural classifications. Stratification was applied at the LGA level, and within each LGA, a sample of EAs was drawn using Probability Proportional-to-Size (PPS) without replacement, using the Tillé method (Tillé 1996).
Stage 2: Within each selected EA, a simple random sample of building footprints was drawn from Google’s open-source dataset. Early planning calculations referenced 39 households per EA and an inflated allowance of 56 sampled buildings to absorb ineligible units and non-response. The realized raw non-sentinel sample file, however, indicates that the operative sample take was closer to 99 sampled building footprints per selected EA. Using that larger denominator implies a planned yield of roughly 0.071 children aged 12-23 months per sampled footprint, similar to the planning assumption used in the sentinel areas.
Stage 3: If a sampled building contained more than two residential units, two units were randomly subsampled.
3.2.3 Achieved Sample
The main survey was conducted between February 22 and Thursday, July 03, 2025, with additional follow-up phone calls for quality control in the weeks that followed. Field teams sampled 72,956 distinct buildings and, including re-visits, made 77,300 survey attempts during this period.3 In one instance, a selected PSU in Tudun Wada LGA was replaced for security reasons with another PSU from the same LGA that had comparable sampling probability. In total, the survey completed 21,048 interviews with caregivers of infants aged 0–23 months, covering 12,004 children aged 12–23 months.
In practice, the survey collected considerably more interviews than the design required. Table 3.2 compares the achieved sample with the original design targets and the key planning inputs that governed field workload and precision. After accounting for expected design effects, the survey planned to reach approximately 5,295 children aged 12–23 months across all four strata.4 The survey ultimately reached 12,004 children in this age band – roughly double the planned figure. Nassarawa roughly met its target, but Gaya, Gabasawa, and the non-sentinel stratum each produced considerable surpluses. The principal cause was the child-yield assumption used to calibrate field workload. The pilot focused on Nassarawa, where eligible children were considerably harder to find per sampled footprint than in the other areas. Because the team applied Nassarawa’s lower yield across all sentinel LGAs, the fixed footprint workload was set higher than Gaya, Gabasawa, and the non-sentinel areas ultimately required. By the time field data revealed the discrepancy, fieldwork was well advanced and the workload could not practicably be reduced.
Stratum | Children 12-23 mo | Yield | DEFF (Penta-1) | Clusters | CVw | HH / elig. child | RR3 | |||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
Plan. | Real. | Plan. | Real. | Plan. | Real. | Plan. | Real. | Plan. | Real. | Plan. | Real. | Plan. | Real. | |
Gaya | 1,086 | 2,912 | 0.071 | 0.191 | 2.00 | 1.51 | — | — | 0.50 | 0.56 | 1.33 | 3.19 | 0.90 | 0.86 |
Nassarawa | 1,086 | 1,283 | 0.071 | 0.084 | 2.00 | 1.10 | — | — | 0.50 | 0.45 | 1.33 | 6.64 | 0.90 | 0.71 |
Gabasawa | 1,086 | 3,092 | 0.071 | 0.203 | 2.00 | 1.48 | — | — | 0.50 | 0.60 | 1.33 | 2.93 | 0.90 | 0.87 |
Non-sentinel (combined) | 2,037 | 3,608 | 0.071 | 0.131 | 3.75 | 3.66 | 291 | 291 | 0.50 | 0.77 | 1.33 | 3.77 | 0.90 | 0.79 |
Total | 5,295 | 10,895 | — | — | — | — | — | — | — | — | — | — | — | — |
Plan. = planned; Real. = realized. Planned children = effective sample size (ESS) of 543 x presumed DEFF (2.0 for sentinel LGAs using stratified simple random sampling of building footprints; 3.75 for non-sentinel using multistage cluster sampling). Realized children = sample size for the Penta-1 analyzability indicator (got_crude_dpt1_to_analyze_01), which may be slightly smaller than the raw 12-23 month interview count. Yield = children aged 12-23 months per sampled building footprint. Clusters shown for non-sentinel only (sentinel sampled buildings directly, not via clusters). CVw = coefficient of variation of trimmed sampling weights. HH / elig. child = households visited per household with >= 1 child 12-23 mo. RR3 = AAPOR Response Rate 3; see Annex for details. Lower-than-planned DEFF generally improves precision, while lower-than-planned response rates or higher-than-planned HH / elig. child generally reduce the achieved sample and work against the original precision target. All realized values exclude the adaptive sampling supplement. | ||||||||||||||
Beyond the child-yield surplus, the remaining planning assumptions present a mixed picture. DEFF values were close to or better than planned – the non-sentinel realized value of 3.66, for example, was near its planned 3.75, and the sentinel values came in well below the assumed 2.0. Other assumptions diverged more substantially: the number of households visited per eligible child was two to five times higher than the planned value of 1.33 across all areas, response rates fell short of the planned 0.90 (particularly in Nassarawa, at 0.71), and the Coefficient of Variation of Sampling Weights (CVw) in the non-sentinel stratum was notably higher than anticipated. On balance, the favorable DEFF values and the higher child yield more than compensated for the shortfalls in response and contact efficiency, producing the net sample overshoot.
While the overshoot increased field effort, it did yield analytical dividends: tighter precision than the original target required, and the ability to conduct analyses – such as detailed geospatial hotspot mapping – that may not have been feasible at the planned scale. As the cost analysis in the companion report shows, field data collection accounts for a relatively modest share of overall study costs, so the additional interviews had a limited effect on the total budget. The realized parameters from this baseline also provide a much stronger empirical foundation for future planning. A follow-up study could draw on updated values for child yield per sampled footprint, response rates, households visited per eligible child, and realized design effects to achieve comparable precision with substantially less field effort. In addition, future pilot exercises should sample a more representative cross-section of study areas rather than concentrating in a single location, so that the planning assumptions reflect the range of conditions likely to be encountered across the full survey domain.
3.2.4 Outcome Rates
An important metric of survey quality are so-called outcome rates. These assess the extent to which our desired sample was achieved, and provide insight into the degree of non-response error that might arise out of uncalibrated analyses. These rates are calculated on the basis of final disposition codes, representing the final categorization after all revisits are conducted for a sampled element:
To probe the nature of the challenges in achieving our desired sample size in each LGA, we can examine American Association of Public Opinion Research (AAPOR) outcome rates. The calculations use the following classification tallies:
code | disposition | n |
|---|---|---|
I | Successful interview with a caregiver of an infant 0-23 months old | 21,048 |
R | Refusal by host or caregiver | 1,936 |
NC | Non-contact (caregiver unavailable) | 420 |
UO | Unknown if household has an infant 0-23 months old | 6,288 |
UH | Unknown if building is residential | 80 |
NE | Known ineligible (non-residential buildings and households without eligible infants) | 44,131 |
Total cases | 73,903 |
The tallies in Table 3.3 are used to calculate the following AAPOR outcome rates:
- Eligibility Rate for Unknown Cases (e): The proportion of cases with unknown eligibility that are estimated to be eligible for the survey. This rate is calculated by decomposing the unknown cases (UH and OH) and separately estimating what proportion of each are expected to yield an eligible household based on the cases with known eligibility status. This is then recombined in a weighted average on the basis of the counts of UO and UH cases.
- Response Rate (RR3): The estimated proportion of eligible households with infants 0-23 that result in a complete interview.
- Contact Rate (CON2): The estimated proportion of eligible households with infants 0-23 that resulted in human contact (i.e., someone answering the door).
- Cooperation Rate (COOP3): The proportion of individuals reached in households with infants 0-23 who agree to participate in the survey process. The denominator excludes cases where contact was unsuccessful.
- Refusal Rate (REF2): The estimated proportion of households with infants 0-23 who refuse to participate in the survey (the denominator includes non-contact cases).
These rates and associated metrics are shown in Table 3.4.
Buildings Attempted | Residential Buildings | Interviews with HH (0-23) | Infants Reached | e | RR3 | CON2 | COOP3 | REF2 | ||
|---|---|---|---|---|---|---|---|---|---|---|
0-23 | 12-23 | |||||||||
State | ||||||||||
Kano | 72,956 | 73% | 21,048 | 24,953 | 12,004 | 48% | 80% | 87% | 92% | 7% |
LGA | ||||||||||
Gabasawa | 15,184 | 75% | 5,912 | 7,302 | 3,410 | 58% | 87% | 92% | 93% | 6% |
Gaya | 15,181 | 74% | 5,299 | 6,543 | 3,146 | 53% | 86% | 92% | 92% | 7% |
Nassarawa | 15,081 | 73% | 2,399 | 2,606 | 1,309 | 29% | 71% | 81% | 88% | 10% |
Tudun Wada | 1,906 | 81% | 763 | 896 | 428 | 56% | 87% | 93% | 94% | 5% |
Takai | 1,625 | 71% | 589 | 664 | 327 | 57% | 87% | 90% | 96% | 4% |
Bebeji | 1,510 | 76% | 526 | 637 | 312 | 53% | 84% | 90% | 94% | 5% |
Sumaila | 2,528 | 73% | 855 | 982 | 461 | 54% | 83% | 93% | 90% | 10% |
Kiru | 2,255 | 72% | 799 | 962 | 518 | 59% | 81% | 91% | 88% | 11% |
Dawakin Tofa | 1,950 | 64% | 516 | 584 | 264 | 50% | 81% | 87% | 92% | 7% |
Dawakin Kudu | 2,194 | 71% | 656 | 735 | 357 | 50% | 80% | 86% | 93% | 6% |
Gezawa | 1,287 | 74% | 423 | 506 | 254 | 54% | 79% | 85% | 92% | 6% |
Dambatta | 1,968 | 73% | 482 | 552 | 256 | 42% | 79% | 87% | 91% | 8% |
Ungogo | 3,147 | 71% | 644 | 717 | 341 | 39% | 71% | 79% | 90% | 8% |
Kumbotso | 5,741 | 67% | 1,015 | 1,088 | 533 | 37% | 68% | 78% | 87% | 10% |
Tarauni | 1,399 | 79% | 170 | 179 | 88 | 29% | 48% | 62% | 77% | 14% |
Urban/Rural | ||||||||||
Rural | 38,264 | 74% | 14,225 | 17,389 | 8,248 | 57% | 85% | 92% | 92% | 7% |
Urban | 34,692 | 72% | 6,823 | 7,564 | 3,756 | 35% | 73% | 81% | 90% | 8% |
e: The proportion of cases of unknown eligibility estimated to be eligible. Because the cases of unknown eligibility are dominated by cases where the eligibility of the household is unknown (rather than where the eligibility of the building is unknown), this rate is virtually identical to the proportion of households with children 0-23 months. | ||||||||||
RR3: The estimated proportion of eligible households with infants 0-23 that result in a complete interview. | ||||||||||
CON2: The estimated proportion of eligible households with infants 0-23 that resulted in human contact (i.e., someone answering the door). | ||||||||||
COOP3: The proportion of individuals reached in households with infants 0-23 who agree to participate in the survey process. The denominator excludes cases where contact was unsuccessful. | ||||||||||
REF2: The estimated proportion of households with infants 0-23 who refuse to participate in the survey (the denominator includes non-contact cases). | ||||||||||
3.3 Survey Weighting
Survey weights have been constructed and adjusted to account for multiple factors potentially affecting representativeness. This can be broken down into seven sequential steps:
- Calculate the building-level base weights
- Post-stratify the building-level base weights
- Adjust building-level weights for unknown building eligibility
- Calculate a calibrated household-in-building weight that accounts for non-response propensity and building multiplicity
- Trim the household-in-building weights to reduce the impact of extreme weights
- Marginalize the weights over distinct household to obtain a calibrated household-level weight
- Join the household level-weight to the child-level data
Each of these steps is described in detail in Appendix D.
3.4 Variance Estimation
Estimates of Standard Errors (SEs) were produced using a generalized bootstrap method (Rao-Wu-Yue-Beaumont) (Beaumont and Émond 2022) as implemented by the {svrep} R package (Schneider 2025). This approach accounts for the complex multi-stage sampling design, which involved differing stratification levels, finite population correction, and sampling stages for sentinel and non-sentinel LGAs.
With the exception of VCQI outputs, all analyses were conducted after applying the replicate design. This ensures that resulting SEs appropriately reflect the uncertainty associated with both the sampling design and the estimation of population sizes.
Further details on variance estimation are available in Section D.2.
When proportions are estimated to be exactly zero or one, the logit interval becomes undefined. In these degenerate cases, a simple Clopper-Pearson confidence interval is returned.↩︎
The effective sample size is the sample size that would be needed in each group if a simple random sample were used. In reality, a simple random sample is seldom a method available to researchers as we do not have sampling frame of all children aged 12-23 months. The effective sample size therefore needs to be inflated account for the clustering effect and other sample losses such as non-response.↩︎
In this report, all tallies and rates pertain to the conventional multi-stage probability sample. Although some head-to-head methods from the companion report share data points with this conventional sample, separate tallies for adaptive sampling, Network Scale-Up Method (NSUM), Lot Quality Assurance Sampling (LQAS), Rapid Convenience Monitoring (RCM), and qualitative fieldwork will be presented in separate reports.↩︎
The planned total of 5,295 children was derived from an Effective Sample Size (ESS) of 543 children per comparison group, inflated by the assumed DEFF for each stratum (2.0 for sentinel LGAs; 3.75 for non-sentinel). The ESS represents the number of children that would be required under Simple Random Sample (SRS) to achieve the target precision; under more complex designs, the nominal sample must be proportionally larger.↩︎