Session Overview
A short closing lecture that explains what HDF5 (Hierarchical Data Format v5) is, its advantages for large array-based genomic datasets, and specifically why SNP-Seek uses it to store variant call matrices. Sets up the Day 2 pipeline run lab.
Topics Covered
What is HDF5?
HDF5 is a binary container format designed for massive datasets — think of it as a filesystem inside a single file. Four key properties make it suited to genomic data:
- Binary Container — stores massive datasets in one file
- Fast Random Access — read any slice without loading the whole file into memory
- Chunked Storage — data divided into chunks for efficient compression and retrieval
- Self-Describing — contains metadata about dimensions, types, and structure
HDF5 by the Numbers — RICE-RP Dataset
Why HDF5 over Text Files?
| Text Matrix | HDF5 | |
|---|---|---|
| File Size | Very large (uncompressed) | Compressed (much smaller) |
| Read Speed | Must scan entire file | Jump to any location |
| Partial Query | Not possible | Read only what you need |
| Human Readable | Yes (good for checking) | No (binary format) |
| Web Portal Use | Too slow for queries | Powers SNP-Seek |
How HDF5 Chunking Works
Chunk shape determines query speed — HDF5 always reads whole chunks. A B-tree index maps each chunk to its byte position on disk, so no scanning is needed.
- Small chunks (e.g.
-n 20): 16 chunks touched, 230 chunk rows to stitch together - Ideal chunks (e.g.
-n 4591, all samples): 2 chunks touched, no stitching needed
The Loading Pipeline
How raw genotype files become an HDF5 file for SNP-Seek:
Input Formats: VCF & PLINK
- VCF — merged multisample file; SNPs only (no indels); contains genotype calls per sample; most common sequencing output
- PLINK — binary PED format (
.bed/.bim/.fam); SNPs only (no structural variants); needs REF allele verification; common in GWAS studies
The Text Matrix — Intermediate Step
Each row is one SNP position; each column after CHR/POS/REF/ALT is a sample's diploid genotype call (e.g. CC, AG, TT). Missing calls are encoded as 00 or ...
Three Output Files
The pipeline produces three files that must stay in the same order:
e.g. ricerp.h5
e.g. sample_list.txt
e.g. pos.txt
How SNP-Seek Uses HDF5
When a user queries SNPs on Chr 1 from 1000–5000 for 10 samples, SNP-Seek:
- Finds matching SNP rows in the position file and sample column indices
- Reads only those rows and columns from the HDF5 file — not the whole file
- Returns the genotype table in the browser in under a second
Key Takeaways
- HDF5 is what makes SNP-Seek fast — random access to billions of data points
- The text matrix is the human-readable checkpoint for verification
- VCF/PLINK files are converted to a text matrix as an intermediate step
- Three files work together: HDF5 + Sample IDs + SNP Positions
- Order consistency across all three files is critical