<< All versions

Skill v1.0.0

currentAutomated scan100/100
open-science-pillars/core/data-formats
──Details
PublishedSeptember 28, 2026 at 01:33 PM
Content Hashsha256:e061b9c53bf059aa...
Git SHA
──Files
Files (1 file, 7.4 KB)
SKILL.md7.4 KBactive
SKILL.md · 142 lines · 7.4 KB

version: "1.0.0" name: data-formats description: Open and inspect NetCDF, HDF4/5, GeoTIFF, Zarr, GRIB, CSV earth science files with xarray, rioxarray, cfgrib; fill values, time decoding, CRS, chunking. user-invocable: false


data-formats

Background expertise for opening earth science data files correctly on the first try, and for diagnosing the openings that fail.

Consult the knowledge bundle; do not carry dataset facts here

Fill-value sentinel lists, calendar and season conventions, and the CF metadata contract are dataset knowledge, held once in the core bundle and never memorized in this skill. Consult installed knowledge concepts first, as the consult-knowledge skill sets out; the concepts this skill leans on are conventions/common-fill-values.md, conventions/calendars.md, and conventions/cf-conventions.md. Read them, restate the relevant point, and cite it. When a concept and this skill disagree, the concept wins.

The tables below (magic bytes, openers, failure symptoms) are invariant format method, not dataset facts, and stay here as procedure.

Identify the format first: magic bytes, not extensions

Extensions lie. A .nc file can be classic NetCDF-3 or HDF5-based NetCDF-4 (different engines); a .hdf file can be HDF4 or HDF5 (different libraries entirely). Check the leading bytes before choosing an opener:

python
with open(path, "rb") as f:
head = f.read(8)
Leading bytesFormat
CDF\x01 or CDF\x02NetCDF-3 classic / 64-bit offset
\x89HDF\r\n\x1a\nHDF5 (includes NetCDF-4)
\x0e\x03\x13\x01HDF4 (MODIS-era products)
II*\x00 or MM\x00*TIFF / GeoTIFF (little / big endian)
GRIBGRIB (byte 8 gives the edition)
no magic; a directory (or object store) containing .zgroup, .zarray, or zarr.jsonZarr (v2 / v3)
printable text, delimiter-separatedCSV or similar; sniff, do not assume

A wrong-looking magic number on a downloaded file usually means a truncated or HTML-error-page download; compare file size against the archive listing before debugging anything else.

Openers

FormatOpenerNotes
NetCDF-3/4xr.open_dataset(path)engine netcdf4 or h5netcdf; both read NetCDF-4, only netcdf4 reads classic directly
HDF5 (non-NetCDF layout)xr.open_dataset(path, engine="h5netcdf") or h5pysatellite L1/L2 products often need group selection: group="..."
HDF4xr.open_dataset(path, engine="netcdf4") if libnetcdf has HDF4 support; else pyhdf or rioxarrayconda-forge libnetcdf carries HDF4 support; pip wheels usually do not
GeoTIFFrioxarray.open_rasterio(path)preserves CRS and transform; plain xarray does not
Zarrxr.open_zarr(store)try consolidated=True first; fall back explicitly
GRIBxr.open_dataset(path, engine="cfgrib")heterogeneous files need backend_kwargs={"filter_by_keys": {...}}
CSVpandas.read_csv(...) then .to_xarray()parse_dates=, explicit na_values=; never trust dtype inference on station data

Open lazily by default (chunks={} engages Dask with on-disk chunking). Declare the compute scale rather than assuming a machine: small (laptop), medium (Dask cluster), large (HPC or burst).

Fill values and packing

mask_and_scale=True (the default) applies _FillValue, missing_value, scale_factor, and add_offset from attributes. Two traps:

  1. Unmasked sentinels. Files that omit the _FillValue attribute but

use sentinel values anyway decode as real data and silently bias every statistic. The sentinel list and the detection recipe are dataset knowledge: consult conventions/common-fill-values.md, restate and cite it, and do not reproduce the values here. The refusal that guards this trap is in Hard refusals below.

  1. Packed integers. If a variable that should be continuous arrives as

int16, packing attributes were probably lost or ignored; check var.encoding for scale_factor before trusting values. The CF packed-data rule itself (how scale_factor and add_offset reconstruct physical values) is a metadata convention: consult conventions/cf-conventions.md.

Time decoding

CF time units (days since 1850-01-01) decode automatically. Non-standard calendars (360_day, noleap, all_leap) need cftime objects: pass use_cftime=True and expect cftime datetimes downstream (resampling works; direct numpy datetime comparisons do not). When decoding fails, open with decode_times=False, inspect time.attrs["units"] and time.attrs["calendar"], repair, then xr.decode_cf(ds). The conventions/calendars.md concept owns the CF calendar list, the month-length weighting rule, and the DJF year-boundary trap; consult and cite it rather than restating those here.

CRS and coordinates

  • GeoTIFF via rioxarray: CRS lives at da.rio.crs; reproject with

da.rio.reproject(...). Never assume EPSG:4326.

  • NetCDF: CF encodes projection in a grid_mapping variable (its contract

lives in conventions/cf-conventions.md); projected data has x/y coordinates in meters, not degrees. Check before treating coordinates as lat/lon.

  • Longitude arrives as either 0..360 or -180..180; normalize deliberately

(ds.assign_coords(lon=(ds.lon + 180) % 360 - 180).sortby("lon")) and say which convention the output uses.

  • Latitude may be descending; sortby("lat") before slicing with ranges.

Per-format failure guidance

SymptomLikely causeFix
[Errno -51] NetCDF: Unknown file formatfile is not what the extension claims, or truncated downloadcheck magic bytes and file size against the archive
unable to decode time unitsnon-CF units string or exotic calendardecode_times=False, inspect attrs, use_cftime=True, then xr.decode_cf
cfgrib: multiple values for unique keyGRIB file mixes levels or stepsbackend_kwargs={"filter_by_keys": {"typeOfLevel": "isobaricInhPa"}} (or the relevant key)
Zarr: KeyError: '.zgroup' or metadata not foundunconsolidated v2 store, v3 store, or wrong path/prefixconsolidated=False; confirm store layout and zarr-python version
GeoTIFF opens with row/col integer coordsopened without rioxarray, or georeferencing absentreopen with rioxarray.open_rasterio; if CRS is still None the file lacks it, ask the producer
HDF4 open fails with netcdf4 enginelibnetcdf built without HDF4conda-forge libnetcdf, or read via pyhdf
variable is all NaN after openbad _FillValue/valid_range attrs masking everythingmask_and_scale=False, inspect raw values and attrs, mask manually
times are object dtype strings (CSV)dates not parsedparse_dates= in read_csv, then .to_xarray()

Hard refusals (stop; never proceed past these)

  • Never trust a file extension over magic bytes.
  • Never report a statistic before confirming fill values are masked.
  • Never assume EPSG:4326, a longitude convention, or an ascending latitude

axis without checking.

  • Never load a multi-GB dataset eagerly; open lazily and state the compute

scale.

  • Never paper over a time-decoding failure by dropping the time axis.

What "open and summarize" means here

A complete summary of an opened dataset states: dimensions and sizes; coordinate ranges (with longitude convention and latitude order); variables with units and dtypes; time span, cadence, and calendar; CRS or grid mapping if any; fill-value status (masked, or sentinels detected); and in-memory vs on-disk size with the chunking in effect.

All versions