Getting started¶
Install the package from PyPI:
pip install geoestimate
Create one Sample for a DataFrame and its sampling design. Then request the
estimand you need.
import pandas as pd
from geoestimate import Sample
frames = pd.DataFrame(
{
"n_women": [3, 4, 2, 5],
"n_people": [10, 10, 10, 10],
"itinerary_id": [0, 0, 1, 1],
}
)
sample = Sample(frames, cluster="itinerary_id")
result = sample.ratio("n_women", "n_people")
print(result.estimate)
print(result.standard_error)
print(result.confidence_interval)
Sample takes a snapshot of the DataFrame. Later changes to frames do not
change the sample or its estimates.
The default inference="design" uses cluster-sandwich inference when you
declare a cluster. It uses iid inference otherwise. Read Choosing an estimand before
choosing between ratios and means of ratios.
Work with results¶
Every method returns an immutable Estimate with the same fields. Use
summary() for display or to_frame() to combine results with pandas.
mean_result = sample.mean("n_people")
ratio_result = sample.ratio("n_women", "n_people")
table = pd.concat(
[mean_result.to_frame(), ratio_result.to_frame()],
ignore_index=True,
)