🆕 New session for 2026. This brief closing lecture frames the HDF5 format before participants work directly with it in the Day 2 pipeline lab.

Session Overview

A short closing lecture that explains what HDF5 (Hierarchical Data Format v5) is, its advantages for large array-based genomic datasets, and specifically why SNP-Seek uses it to store variant call matrices. Sets up the Day 2 pipeline run lab.

Topics Covered

What is HDF5?

HDF5 is a binary container format designed for massive datasets — think of it as a filesystem inside a single file. Four key properties make it suited to genomic data:

HDF5 by the Numbers — RICE-RP Dataset

5,231,433
SNP positions
4,591
rice samples
~24 Billion
genotype data points
< 1 sec
query response time

Why HDF5 over Text Files?

Text MatrixHDF5
File SizeVery large (uncompressed)Compressed (much smaller)
Read SpeedMust scan entire fileJump to any location
Partial QueryNot possibleRead only what you need
Human ReadableYes (good for checking)No (binary format)
Web Portal UseToo slow for queriesPowers SNP-Seek
⚡ Key Insight: The text matrix is the human-friendly checkpoint — you can open it and verify the data. HDF5 is the machine-optimized version that makes SNP-Seek fast enough for a web browser.

How HDF5 Chunking Works

Chunk shape determines query speed — HDF5 always reads whole chunks. A B-tree index maps each chunk to its byte position on disk, so no scanning is needed.

The Loading Pipeline

How raw genotype files become an HDF5 file for SNP-Seek:

1 · Input
VCF or PLINK
Standard genotype formats from sequencing pipelines
→
2 · Convert
Text Matrix
Human-readable table: CHR | POS | REF | ALT | Samples
→
3 · Output
HDF5 File
Binary, indexed, compressed — optimized for fast queries

Input Formats: VCF & PLINK

⚠ Important: Only SNP genotypes — indels must be processed separately.

The Text Matrix — Intermediate Step

Each row is one SNP position; each column after CHR/POS/REF/ALT is a sample's diploid genotype call (e.g. CC, AG, TT). Missing calls are encoded as 00 or ...

Three Output Files

The pipeline produces three files that must stay in the same order:

HDF5 File
Genotype matrix in binary format — optimized for random access queries
e.g. ricerp.h5
Sample IDs
Sample names in the same column order as the HDF5 file
e.g. sample_list.txt
SNP Positions
CHR, POS, REF, ALT — same row order as HDF5
e.g. pos.txt
⚠ Order matters — all three files must be in the same order (same SNP rows, same sample columns).

How SNP-Seek Uses HDF5

When a user queries SNPs on Chr 1 from 1000–5000 for 10 samples, SNP-Seek:

  1. Finds matching SNP rows in the position file and sample column indices
  2. Reads only those rows and columns from the HDF5 file — not the whole file
  3. Returns the genotype table in the browser in under a second

Key Takeaways

💾 Slide Deck

📂 Open slides in new tab