geoestimate

Design-aware estimates for equal-probability observation samples.

PyPI CI Docs Python 3.12+

geoestimate estimates means, population totals, ratios of totals, and means of row-level ratios. It reports iid, cluster-sandwich, or pairs-bootstrap uncertainty. The stable API assumes that rows have equal inclusion probabilities.

The package does not implement unequal-probability weights, PPS, GRTS, stratification, finite population corrections, nonresponse adjustments, annotation-subsampling corrections, or walk-spacing corrections.

Install

pip install geoestimate

Estimate from a sample

import pandas as pd

from geoestimate import Sample

frames = pd.DataFrame(
    {
        "n_women": [3, 4, 2, 5],
        "n_people": [10, 10, 10, 10],
        "itinerary_id": [0, 0, 1, 1],
    }
)

sample = Sample(frames, cluster="itinerary_id")

mean_people = sample.mean("n_people")
total_people = sample.total("n_people", population_size=50_000)
people_share = sample.ratio("n_women", "n_people")
location_share = sample.mean_of_ratios("n_women", "n_people")

print(people_share.summary())

Declare a cluster only when it identifies independent sampling or collection units. A declared cluster makes cluster-sandwich inference with a Student-t interval the default. Without a cluster, the default is an iid standard error with a normal interval.

Choose the estimand

mean("n_people") estimates the average count per population unit. For a binary variable, the mean is a population proportion.

total("n_people", population_size=N) estimates N * mean(n_people). The population size counts the same row-level units represented by the sample. The method does not apply a finite population correction.

ratio("n_women", "n_people") estimates:

sum(n_women) / sum(n_people)

This ratio weights rows by their denominator. Individual denominators may be zero, but they must be nonnegative and their sample total must be positive.

mean_of_ratios("n_women", "n_people") estimates:

mean(n_women / n_people)

This estimand gives every row equal weight. Every denominator must be positive. Filter the DataFrame before constructing Sample when the target population excludes rows with zero denominators.

Select inference

The default inference="design" follows the declared sample design. You can request a method explicitly:

bootstrap = sample.ratio(
    "n_women",
    "n_people",
    inference="bootstrap",
    bootstrap_reps=2_000,
    seed=42,
)

The bootstrap resamples declared clusters, or individual rows when no cluster is declared. It reports the standard deviation of the bootstrap estimates and a percentile interval. Explicit inference="iid" is available as a sensitivity comparison for clustered samples.

Read files and use the command line

Sample.from_file() accepts Parquet, CSV, compressed CSV, and TSV files:

sample = Sample.from_file("frames.parquet", cluster="itinerary_id")
result = sample.ratio("n_women", "n_people")

The command line exposes the same four estimands:

geoestimate mean frames.parquet --variable n_people --cluster itinerary_id
geoestimate total frames.parquet --variable n_people --population-size 50000
geoestimate ratio frames.parquet --numerator n_women --denominator n_people
geoestimate mean-of-ratios frames.parquet --numerator n_women --denominator n_people

Experimental validation tools

geoestimate.spatial, geoestimate.simulate, and geoestimate.pipeline help diagnose dependence and validate collection designs. Their interfaces may change before the stable inference API does.

Install the optional pipeline dependencies to validate a real geo-sampling to allocator geometry:

pip install "geoestimate[pipeline]"
python examples/validate_with_allocator.py

License

MIT