Session Overview
Continues from Jeffrey Detras' Ontology Concepts for Chado lecture. Where that session covered the vocabulary (cv, cvterm, dbxref, relationships), this session covers the plumbing — the actual tables you load into and query for SNP-Seek, and how every one of them hangs off the cvterm bridge.
Topics Covered
What is Chado?
- A modular, ontology-driven, extensible schema — the foundation SNP-Seek is built on
- Hard-coded vs. ontology-driven design: why Chado uses one
featuretable typed bycvterminstead of separate tables per biological type
The Core Pattern — type_id → cvterm
feature.type_id— is this a gene, a SNP, a chromosome?featureprop.type_id— what kind of property is this annotation?- Relationship tables —
is_a,part_of,derives_from synonym.type_id— which naming system this name belongs to- New feature type tomorrow? Add a
cvterm, not a table
Commonly Used SNP-Seek Tables
db— external databases / sources (GenBank, Ensembl)cv/cvterm— controlled vocabularies and ontology terms (SO, GO)organism— taxonomy for species & subspeciesfeature— central table: genes, SNPs, chromosomes, contigssnp_feature/snp_featureloc— SNP-specific features and their genomic locationsstock/stock_sample— biological accessions and their sequencing-run linksvariantset/variant_variantset— named variant groupings and their join table
Schema ERDs
- The
cvtermcluster:cvterm ↔ dbxref ↔ db, withcvdefining the vocabulary - Full modular ERD: organism, feature, stock, variantset — each stored separately, linked by foreign keys
The feature Table and Location
- Genes, SNPs, contigs, chromosomes — all rows in one table, distinguished by
type_id featureloc:srcfeature_id(chromosome),fmin/fmax(coordinates),strandfeatureprop: notes, descriptions, scores — also typed bycvterm- Loading gotcha:
fminis zero-based and interbase — not the 1-based number in a genome browser - Tracing one SNP through the schema:
feature→cvterm(type),organism,featureloc,srcfeature(chromosome)
Stocks, Variant Sets & the PostgreSQL / HDF5 Boundary
organism— reference genomes & rice lines;stock/stock_sample— the 3,000+ accessionsvariantset/variant_variantset— named groupings of variants- The boundary: PostgreSQL/Chado holds WHAT and WHERE; the dense genotype matrix (~60 billion data points) lives in HDF5 files — not in PostgreSQL
Synonyms — One Locus, Many Names
- Naming systems in SNP-Seek: MSU7/LOC_, IRIC, RAP representative/predicted, FGENESH, per-genome (IR64, MH63, N22, Azucena)
synonym+feature_synonymtyped bycvterm— no wide locus-mapping table needed- New naming system = add a cvterm, not a column
Why This Design Wins
- Ontology-driven typing — meaning lives in
cvterm, referenced bytype_ideverywhere - Extensibility — new types and new names are new cvterms, not schema changes
- Integrity — foreign keys hold features, locations, stocks, and sets together
- Most loader failures are term-resolution problems — a name in the file that doesn't match a term