Collection protocol
Saturday pull 2026-09-05
(MFD-2026-09-05-standard-v2 /
MFD-2026-09-05-sample-v2). How the files are made, and
what is in them. Status on the site moves when we refresh; this page
stays pinned to a release.
Abstract
Banlys MFD is a licensed, versioned package of multi-sensor readings from consumer phones. People turn on research sharing in a Banlys mobile sensing app. When they are sensing, the phone sends short batches: magnetometer (used here as an EMF proxy), motion, gyro, light, barometer when the hardware has one, location rounded to about 100 meters, and — after a short baseline calibration — entropy summaries over rolling windows.
A package is a filtered export, a dictionary, a manifest with checksums, and terms.
Why we collect it
Field sessions are more useful when they stack: same rough protocol, many phones, many places, and flags you can filter on. The hardware is already in the pocket. What we add is consent, a cadence, a baseline, and a versioned file.
These channels are what consumer phones will give us on a repeatable interval. Calibrated sessions on the same device family are the ones that compare cleanly.
How a row gets born
Person installs a Banlys sensing app → research sharing stays off until they turn it on → they start sensing → the phone takes a baseline (~45 seconds, hold still) if they go through that step → about every 30 seconds the app may send a batch → Banlys stores the batch → Saturday we pull a full export → a versioned MFD package is built from that export → the licensed copy is that package
Paid features and research sharing are separate. Sharing can be turned off in Settings. A batch is sent when sharing is on and sensing is running.
What we collect
Sensors
Values are device-dependent. Same model compares better than mixing every Android together.
| Channel | What it is | Notes |
|---|---|---|
emf / magnet |
Magnetometer | Same sensor in the current app |
motion |
Accelerometer-derived | Still vs moving |
gyro |
Gyroscope | Common on current phones |
light |
Ambient light | Scale varies |
pressure |
Barometer | About 22% of readings in this pull |
entropy |
Shannon summaries + a composite | After calibration |
Each reading has a device timestamp. Each batch also has a server received time. Keep those two clocks distinct.
Location
Latitude and longitude are rounded to about 100 meters (three decimal degrees). That is enough for regional maps. Some batches have no location; those are not a place.
Calibration and entropy
A ~45 second hold-still baseline lets later readings include
entropy. Those batches are marked
calibrated.
Entropy is Shannon entropy over a rolling window (60 samples in the
live stream, faster than the 30-second research cadence). Values are
scaled to 0–1. composite combines the per-channel
entropies and is shifted so a quiet baseline sits near the middle of
the scale.
If entropy is missing, there was no baseline on that
batch.
Quality and instrument fields (schema v2)
Newer clients send these. Older batches may leave them blank.
| Field | Role |
|---|---|
research_quality |
gold when calibrated, at least 5 readings, and not flagged degenerate; otherwise standard |
entropy_degenerate |
Flat or floor windows. Often blank on older rows |
same_location_as_calibration |
Batch still within ~100 m of the baseline |
device_family |
Coarse model token |
charging / activity_state |
Useful around the magnetometer and still vs moving |
sample_interval_ms |
Nominal 30000 on current mobile |
gold is a starting filter. High-zero sessions can still
wear that label, so look at the values.
What a package is
A release is a folder:
data.json— batches for that SKU- dictionary and terms
manifest.json—release_id, counts, filters, checksumsCHECKSUMS.txt
Sample (this pull: 50 batches) is a small slice of the schema. Standard is the research cut of that Saturday’s export. Annual access covers Standard releases published during the license term.
Snapshot — 2026-09-05
Official pull headlines. Platform counts are schema v2.
| Batches | 5,794 |
| Schema v2 | 5,266 |
| Calibrated (entropy present) | 3,156 |
| Gold (as labeled) | 1,703 |
| Readings | 36,466 |
| Readings with entropy | 21,776 |
| Entropy composite exactly 0 | 5,493 (25.2%) |
| Location cells (~0.001°) | 749 |
| Calendar days (any / entropy) | 78 / 71 |
| v2 Android / iOS | 5,217 / 49 |
About 54% of opt-in batches finished calibration. Most new volume that week was Android.
Suggested cuts
- Use calibrated rows when you need entropy.
- Prefer
same_location_as_calibration = truewhen the baseline should match the site. - Set aside cells that are almost all zero.
- Split by
device_familyor at least by platform. - Use charging and activity when those fields are present.
entropy_degenerate is blank on most older batches. If
the field is missing, work with the numbers you have. A gold /
same-place cut is available as a filter on the same corpus.
Provenance
Banlys MFD packages are derived from opt-in research contributions collected through Banlys mobile sensing applications. Packages include multi-sensor readings, coarse location (~100 m), and quality metadata as documented in the data dictionary. They do not include media files or account identifiers.
Sharing is consumer opt-in. Institutions that need a review for secondary use can treat it as consented consumer sensor data and follow their own process. Location is coarse by design; the license still asks that people not be re-identified from it.
The Banlys sensing app on the stores today is JuJu.
Where to get it
Licenses are on the Access page — sample
and annual seat. Other licenses are quoted. Live counts sit on
Status. If you use these numbers in a
write-up, cite the release_id on the manifest and the
pull date (2026-09-05).
Descriptor dated 2026-09-11. Numbers from the 2026-09-05 pull.