Stable API reference

class geoestimate.Sample(data, *, cluster=None)[source]

Bind observations to an equal-probability sampling design.

Notes

Sample supports equal-probability observations with iid or cluster-sandwich inference. It does not implement weights, strata, finite population corrections, PPS, or GRTS variance estimation.

Parameters:
  • data (DataFrame)

  • cluster (str | None)

classmethod from_file(path, *, cluster=None)[source]

Read a Parquet, CSV, or TSV file and construct a sample.

Parameters:
Return type:

Sample

property cluster: str | None

Return the declared cluster-column name.

property n_observations: int

Return the number of sampled observations.

property n_clusters: int | None

Return the declared cluster count, or None for iid samples.

mean(variable, *, inference='design', confidence_level=0.95, bootstrap_reps=2000, seed=None)[source]

Estimate the population mean of a numeric variable.

Parameters:
  • variable (str)

  • inference (Literal['design', 'iid', 'cluster', 'bootstrap'])

  • confidence_level (float)

  • bootstrap_reps (int)

  • seed (int | None)

Return type:

Estimate

total(variable, *, population_size, inference='design', confidence_level=0.95, bootstrap_reps=2000, seed=None)[source]

Estimate a population total as population size times the sample mean.

Parameters:
  • variable (str)

  • population_size (int)

  • inference (Literal['design', 'iid', 'cluster', 'bootstrap'])

  • confidence_level (float)

  • bootstrap_reps (int)

  • seed (int | None)

Return type:

Estimate

ratio(numerator, denominator, *, inference='design', confidence_level=0.95, bootstrap_reps=2000, seed=None)[source]

Estimate a ratio of population totals.

Parameters:
  • numerator (str)

  • denominator (str)

  • inference (Literal['design', 'iid', 'cluster', 'bootstrap'])

  • confidence_level (float)

  • bootstrap_reps (int)

  • seed (int | None)

Return type:

Estimate

mean_of_ratios(numerator, denominator, *, inference='design', confidence_level=0.95, bootstrap_reps=2000, seed=None)[source]

Estimate the population mean of row-level numerator/denominator ratios.

Parameters:
  • numerator (str)

  • denominator (str)

  • inference (Literal['design', 'iid', 'cluster', 'bootstrap'])

  • confidence_level (float)

  • bootstrap_reps (int)

  • seed (int | None)

Return type:

Estimate

class geoestimate.Estimate(estimand, variables, estimate, standard_error, confidence_interval, confidence_level, inference_method, n_observations, n_clusters, population_size, diagnostics)[source]

A scalar estimate with uncertainty and design diagnostics.

Parameters:
summary()[source]

Return a compact, human-readable summary.

Return type:

str

to_frame()[source]

Return the scalar result as a one-row tidy table.

Return type:

DataFrame

class geoestimate.Diagnostics(iid_standard_error, cluster_standard_error, design_effect, effective_sample_size, cluster_sizes, effective_cluster_count)[source]

Analytic diagnostics for one estimate.

Parameters:
  • iid_standard_error (float)

  • cluster_standard_error (float | None)

  • design_effect (float | None)

  • effective_sample_size (float)

  • cluster_sizes (tuple[int, ...] | None)

  • effective_cluster_count (float | None)

iid_standard_error

Standard error that treats rows as independent.

Type:

float

cluster_standard_error

Cluster-sandwich standard error, when declared.

Type:

float | None

design_effect

Cluster variance divided by iid variance, when defined.

Type:

float | None

effective_sample_size

Observation count divided by the design effect.

Type:

float

cluster_sizes

Number of observations in each declared cluster.

Type:

tuple[int, …] | None

effective_cluster_count

Kish effective count of declared clusters.

Type:

float | None