cartVersion cartVersion cartVersion cartVersion 0 -10 0 0 0 0 0 0 0 0 0 cartVersion cartVersion cartVersion 0 cartVersion 10 gnomad3MeanCoverage Mean Coverage bigWig gnomAD Mean Genome Sample Coverage v3.0.1 2 0.1 255 0 0 255 127 127 0 0 0 varRep 0 alwaysZero on\ autoScale on\ bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.mean.bw\ color 255,0,0\ longLabel gnomAD Mean Genome Sample Coverage v3.0.1\ parent gnomad3Coverage on\ priority 0.1\ shortLabel Mean Coverage\ track gnomad3MeanCoverage\ gnomad4ExomeMeanCoverage Mean Coverage bigWig gnomAD Mean Exome Sample Coverage v4.0 2 0.1 255 0 0 255 127 127 0 0 0 varRep 0 alwaysZero on\ autoScale on\ bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.mean.bw\ color 255,0,0\ longLabel gnomAD Mean Exome Sample Coverage v4.0\ parent gnomad4ExomeCoverage on\ priority 0.1\ shortLabel Mean Coverage\ track gnomad4ExomeMeanCoverage\ varFreqsBackground Population reference bigBed 9 + SNV Frequencies: variants in ~1.5 million individuals from population cohorts and unaffected or control arms 3 0.1 0 0 0 127 127 127 0 0 0
\ This track shows small variants (SNVs and short indels) seen in population reference\ cohorts and in unaffected or control individuals of disease-study cohorts, annotated\ with their predicted protein consequence and colored by severity. It is the background half\ of a matched pair: the companion\ Disease cohorts track shows the same\ kind of variants seen in affected or case individuals. Displaying the two together lets you\ see how common a variant is in the general/unaffected population compared with affected\ individuals. For the full list of contributing projects, see the\ SNV Frequencies collection page.\
\\ The background combines two kinds of data: the population/biobank reference cohorts (such as\ gnomAD HGDP+1kG, TOPMed, ALFA, HRC and the many national WGS projects), and the\ unaffected/control or unknown-phenotype arms of the disease-study cohorts (non-ASD family\ members in SFARI SPARK WES/WGS, SCHEMA controls, and GREGoR unaffected/unknown\ participants). Genotyping-array cohorts are not included. A variant that also appears in\ affected individuals is shown in both this track and the\ Disease cohorts track.\
\ \Variants are colored by their most severe predicted consequence:
\| Color | Consequence class | Examples |
|---|---|---|
| \ | Protein-truncating / loss-of-function | \stop_gained, frameshift, splice_donor, splice_acceptor, stop_lost, start_lost |
| \ | Missense / in-frame | \missense, inframe_insertion, inframe_deletion, protein_altering |
| \ | Synonymous | \synonymous, stop_retained |
| \ | Non-coding / intergenic | \intron, non_coding, intergenic, UTR |
\ The score (used for shading) is the pooled background allele frequency times 1000.\
\ \\
Background AF is the pooled rate across contributing population cohorts and\
unaffected/control arms: backgroundAF = sum(AC) / sum(AN), where\
backgroundAC sums the allele counts and backgroundAN sums the allele\
numbers across each cohort/arm that provides both AC and AF (the per-arm AN is derived as\
round(AC / AF)). Two cohorts that publish only AF (ABraOM, ALFA) are still\
pooled by assigning them an assumed allele number, set as a default_an in the\
build configuration; their per-arm AC is then derived as round(AF × default_an).\
Cohorts that publish\
only AC with no default_an set (currently MGRB and the GREGoR unaffected and\
unknown arms), and cohorts that contribute only through per-population AC/AF (currently\
AllOfUs), are listed in backgroundSources but do not contribute to the pool\
numerator or denominator; their data remain visible in the per-database and per-population\
AC/AF columns. The pooled rate is preferred over a max-across-cohorts statistic so a small\
cohort with a high local AF (for example AllOfUs Oceanian) cannot dominate the displayed\
frequency.\
\ The pooled rate also inherits a shared-sample bias: several source cohorts overlap\ in the individuals they include. For example, 1000 Genomes samples appear in both gnomAD\ HGDP+1kG and HRC; HGDP and SGDP overlap; AllOfUs and TOPMed share participants; ALFA\ aggregates dbGaP studies used elsewhere in the pool. Where a variant sits in a shared\ sample, both its AC and AN are counted more than once, so pooled AN is inflated and\ pooled AF is skewed toward the frequency in the shared subset. Treat the pooled rate as\ a cross-cohort summary rather than an unbiased population estimate; the per-cohort\ AC/AF/AN fields on each variant give the single-cohort numbers.\
\ \\
Alongside the pooled rate, the mouseover lists the top 3 contributing\
background sources ranked by their own per-source AF, formatted as\
Source (AF). This surfaces population cohorts where a variant\
is specifically enriched, even when the pooled rate is small; the\
East-Asian founder allele\
rs4986893,\
for example, ranks ToMMo Japan and KOVA Korea at the top while the pooled\
rate across all contributing sources sits much lower. For disease cohorts\
that ship a phenotype split (SPARK, SFARI WGS, SCHEMA, GREGoR), the\
displayed AF is the unaffected-arm AF and the label includes the arm (for\
example SPARK non-ASD, SCHEMA ctrl); for\
population cohorts, the label is the cohort name and the AF is the unified\
cohort AF. Per-population sub-ancestries of a cohort (such as gnomAD\
HGDP+1kG continental groups) are deliberately excluded from this ranking so\
sub-population frequencies do not crowd out actual project-level signals.\
\
Two source cohorts are also excluded from the Top-3 ranking: SGDP\
and SVatalog. Their VCFs encode allele counts per genotyped site\
rather than per population (each variant in a single individual produces\
AC=1, AN=2, AF=0.5), so the per-source AF is not a population\
frequency and would always sit near the top of the ranking with a\
meaningless value. Both cohorts still appear in backgroundSources\
and still contribute their (small) AC and AN to the pooled\
backgroundAF; they are only suppressed from the Top-3 list.\
\
Variant-frequency VCFs from the contributing cohorts were stripped of unneeded INFO fields,\
normalized with bcftools norm (splitting multi-allelic sites), and merged with\
bcftools merge. The merged callset was annotated with predicted protein\
consequences using bcftools csq against the\
Ensembl\
GRCh38 release 115 gene models.\
\
A custom Python script (vcfToBigBed.py) then read the per-cohort allele\
counts and frequencies and, for each variant, pooled the allele counts and allele numbers\
across the population cohorts and unaffected/control subgroups to produce this track, and\
across the affected arms to produce the companion\
Disease cohorts track. A variant seen\
in both groups appears in both tracks. The build is documented in the\
makeDoc, and the scripts are on\
GitHub.\
\ Because the merged callset combines cohorts whose redistribution licenses differ, this\ track is not available for download and is not in the Table Browser. It can be\ reconstructed from the individual source VCFs using the\ conversion scripts and the\ build documentation. The per-project subtracks on the\ SNV Frequencies collection page document how to obtain\ each source dataset.\
\ \\ This track is only possible thanks to the data from millions of volunteers around the world\ who contributed to the population reference projects and to the unaffected/control arms of\ the disease cohorts. Click the individual project subtracks on the\ SNV Frequencies collection page for the specific credits\ and citations of each cohort. Thanks to Alex Ioannidis, UCSC, for the inspiration for this\ track and to Andreas Lahner, MGZ, for feedback.\
\ \\ For the primary citation of each source cohort, see the References section on the\ SNV Frequencies collection page. The merged-track build\ uses the following tools:\
\\ Danecek P, McCarthy SA.\ \ BCFtools/csq: haplotype-aware variant consequences.\ Bioinformatics. 2017 Jul 1;33(13):2037-2039.\ PMID: 28205675;\ PMC: PMC5870570\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_background/varFreqsBackground.bb\ filter.affectedAC 0:500000\ filter.affectedAF 0:1\ filter.affectedAN 0:500000\ filter.altLen 1:6294\ filter.backgroundAC 0:5000000\ filter.backgroundAF 0:1\ filter.backgroundAN 0:5000000\ filter.inAffected 0:1\ filter.refLen 1:28037\ filter.varLen -28036:6293\ filterByRange.affectedAC on\ filterByRange.affectedAF on\ filterByRange.affectedAN on\ filterByRange.altLen on\ filterByRange.backgroundAC on\ filterByRange.backgroundAF on\ filterByRange.backgroundAN on\ filterByRange.inAffected on\ filterByRange.refLen on\ filterByRange.varLen on\ filterLabel.affectedAC Affected/case AC\ filterLabel.affectedAF Affected/case AF (pooled)\ filterLabel.affectedAN Affected/case AN (pool denominator)\ filterLabel.affectedCohorts Affected/case cohort\ filterLabel.altLen Alternate Length\ filterLabel.backgroundAC Background AC (population + unaffected)\ filterLabel.backgroundAF Background AF (pooled)\ filterLabel.backgroundAN Background AN (pool denominator)\ filterLabel.backgroundSources Background source (population or unaffected)\ filterLabel.consequence Consequence\ filterLabel.inAffected Seen in an affected/case arm (1=yes, 0=no)\ filterLabel.refLen Reference Length\ filterLabel.varLen Length Change\ filterLabel.varType Variant Type\ filterLimits.affectedAC 0:500000\ filterLimits.affectedAF 0:1\ filterLimits.affectedAN 0:500000\ filterLimits.altLen 1:6294\ filterLimits.backgroundAC 0:5000000\ filterLimits.backgroundAF 0:1\ filterLimits.backgroundAN 0:5000000\ filterLimits.inAffected 0:1\ filterLimits.refLen 1:28037\ filterLimits.varLen -28036:6293\ filterType.affectedCohorts multipleListOr\ filterType.backgroundSources multipleListOr\ filterType.consequence multipleListOr\ filterValues.affectedCohorts SPARK|SFARI SPARK WES,SFARI_WGS|SFARI SPARK WGS,GREGoR|GREGoR,SCHEMA|SCHEMA,GA4K|GA4K PacBio LR\ filterValues.backgroundSources AllOfUs|AllOfUs,SPARK|SFARI SPARK WES,SFARI_WGS|SFARI SPARK WGS,GenomeAsia|GenomeAsia SNVs,GenomeAsiaIndel|GenomeAsia Indels,NPM|NPM Singapore,KOVA|KOVA Korea,ToMMo|ToMMo Japan,FinnGen|FinnGen Finland,Saudi|Saudi,SweGen|SweGen Sweden,TOPMed|TOPMed,ABraOM|ABraOM Brazil,ALFA|ALFA,MGRB|MGRB Australia,HRC|HRC,SGDP|SGDP,HGDP1kG|gnomAD HGDP+1kG,GREGoR|GREGoR,SCHEMA|SCHEMA,CoLoRSdb|CoLoRSdb PacBio LR,SVatalog|SVatalog 101 10XG SR,Tishkoff180|Tishkoff 180 African WGS,WBBC|WBBC China,ChinaMAP|China ChinaMAP,GenomeIndia|GenomeIndia 9.7k WGS,GoNL|GoNL Netherlands ~13x SR\ filterValues.consequence missense|Missense,synonymous|Synonymous,stop_gained|Stop Gained,frameshift|Frameshift,splice_donor|Splice Donor,splice_acceptor|Splice Acceptor,intron|Intron,3_prime_utr|3' UTR,5_prime_utr|5' UTR,non_coding|Non-coding,.|Intergenic,others|Other\ filterValues.varType SNV|SNV,INS|Insertion,DEL|Deletion,MNV|MNV\ itemRgb on\ longLabel SNV Frequencies: variants in ~1.5 million individuals from population cohorts and unaffected or control arms\ maxWindowToDraw 5000000\ mouseOver Var: ${name}\ This track shows small variants (SNVs and short indels) that were observed in\ affected or case individuals of disease-study cohorts, annotated with their\ predicted protein consequence and colored by severity. It is one half of a matched pair:\ the companion\ Population reference track shows the same\ kind of variants seen in population reference cohorts and in unaffected relatives or\ controls. Displaying the two together lets you compare, for example, how often a\ loss-of-function variant in a gene of interest is seen in affected individuals versus the\ general/unaffected background. For the full list of contributing projects, see the\ SNV Frequencies collection page.\
\\ The affected counts are drawn from the affected or case arm of five disease-study cohorts:\ SFARI SPARK WES and SFARI SPARK WGS (autism spectrum disorder probands), SCHEMA\ (schizophrenia cases), GREGoR (affected rare-disease participants), and GA4K (a pediatric\ rare-disease cohort). For SPARK, SFARI WGS, SCHEMA, and GREGoR, the source data carries an\ explicit affected/unaffected (or case/control) label, and only the affected arm feeds this\ track. GA4K reports a single cohort-wide frequency with no per-individual label; because it\ is a rare-disease cohort, it is counted as affected here, with the caveat that it enrolls\ parent-child trios, so a minority of its carriers are unaffected parents. Genotyping-array\ cohorts are not included in either track.\
\ \Variants are colored by their most severe predicted consequence:
\| Color | Consequence class | Examples |
|---|---|---|
| \ | Protein-truncating / loss-of-function | \stop_gained, frameshift, splice_donor, splice_acceptor, stop_lost, start_lost |
| \ | Missense / in-frame | \missense, inframe_insertion, inframe_deletion, protein_altering |
| \ | Synonymous | \synonymous, stop_retained |
| \ | Non-coding / intergenic | \intron, non_coding, intergenic, UTR |
\ The score (used for shading) is the pooled affected/case allele frequency times 1000.\
\ \\
Affected AF is the pooled rate across contributing affected arms:\
affectedAF = sum(AC) / sum(AN), where affectedAC sums the allele counts\
and affectedAN sums the allele numbers across each cohort/arm that provides both AC and\
AF (the per-arm AN is derived as round(AC / AF)). Cohorts that publish only AF\
(with no AC or AN of their own) are still pooled by assigning them an assumed allele number,\
set as a default_an in the build configuration; their per-arm AC is then derived\
as round(AF × default_an). Cohorts\
that publish only AC and have no default_an set (currently GREGoR's per-arm\
AC_AFFECTED/UNAFFECTED/UNKNOWN) are listed in affectedCohorts but do not contribute\
to the pool numerator or denominator; their carriers are visible in the per-database AC\
column instead. The pooled rate is preferred over a max-across-cohorts statistic so a\
small cohort with a high local AF cannot dominate the displayed frequency.\
\ The pooled rate also inherits a shared-sample bias: the SFARI SPARK WGS cohort\ (~12k probands) is a subset of the larger SFARI SPARK WES cohort (~155k probands), so\ probands sequenced in both contribute their AC and AN twice to the affected pool. Where\ this happens, pooled AN is inflated and pooled AF is skewed toward the frequency in the\ shared subset. Treat the pooled rate as a cross-cohort summary rather than an unbiased\ population estimate; the per-cohort AC/AF/AN fields on each variant give the\ single-cohort numbers.\
\ \\
Alongside the pooled rate, the mouseover lists the top 3 contributing\
affected arms ranked by their own per-source AF, formatted as\
Source (AF). This surfaces case cohorts where the variant is\
specifically enriched, even when the pooled rate across all arms is small.\
For disease cohorts that ship a phenotype split (SPARK, SFARI WGS, SCHEMA,\
GREGoR), the displayed AF is the affected-arm AF and the label includes the\
arm (for example SPARK ASD, SCHEMA case); for\
cohorts with no split (GA4K) the label is just the cohort name and the AF\
is the whole-cohort AF. Arms that ship only AC and no AF (currently GREGoR\
per-arm) are not included in this ranking because no AF is available;\
they still appear in affectedCohorts.\
\ To look for protein-truncating variants that are common in affected individuals but rare\ in the background, set the Consequence filter to Stop Gained, Frameshift, Splice Donor and\ Splice Acceptor (these appear red), then add an upper limit on the\ Background AF filter. Each variant here carries both its affected frequency and its\ background frequency, so this isolates variants seen in cases with little or no presence in\ the population/unaffected set. Comparing visually against the\ Population reference track shows the same\ contrast across a whole gene.\
\ \\
Variant-frequency VCFs from the contributing cohorts were stripped of unneeded INFO fields,\
normalized with bcftools norm (splitting multi-allelic sites), and merged with\
bcftools merge. The merged callset was annotated with predicted protein\
consequences using bcftools csq against the\
Ensembl\
GRCh38 release 115 gene models.\
\
A custom Python script (vcfToBigBed.py) then read the per-cohort allele\
counts and frequencies and, for each variant, pooled the allele counts and allele numbers\
across the affected arms (case/proband subgroups, plus GA4K whole-cohort) to produce this\
track, and across the population cohorts and unaffected/control subgroups to produce the\
companion Population reference track. A variant\
seen in both groups appears in both tracks. The build is documented in the\
makeDoc, and the scripts are on\
GitHub.\
\ Because the merged callset combines cohorts whose redistribution licenses differ, this\ track is not available for download and is not in the Table Browser. It can be\ reconstructed from the individual source VCFs using the\ conversion scripts and the\ build documentation. The per-project subtracks on the\ SNV Frequencies collection page document how to obtain\ each source dataset.\
\ \\ This track is only possible thanks to the data from the participants and families of the\ SFARI SPARK, SCHEMA, GREGoR and GA4K studies. Click the individual project subtracks on the\ SNV Frequencies collection page for the specific credits\ and citations of each cohort. Thanks to Alex Ioannidis, UCSC, for the inspiration for this\ track and to Andreas Lahner, MGZ, for feedback.\
\ \\ For the primary citation of each source cohort, see the References section on the\ SNV Frequencies collection page. The merged-track build\ uses the following tools:\
\\ Danecek P, McCarthy SA.\ \ BCFtools/csq: haplotype-aware variant consequences.\ Bioinformatics. 2017 Jul 1;33(13):2037-2039.\ PMID: 28205675;\ PMC: PMC5870570\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_affected/varFreqsAffected.bb\ filter.affectedAC 0:500000\ filter.affectedAF 0:1\ filter.affectedAN 0:500000\ filter.altLen 1:6294\ filter.backgroundAC 0:5000000\ filter.backgroundAF 0:1\ filter.backgroundAN 0:5000000\ filter.inAffected 0:1\ filter.refLen 1:28037\ filter.varLen -28036:6293\ filterByRange.affectedAC on\ filterByRange.affectedAF on\ filterByRange.affectedAN on\ filterByRange.altLen on\ filterByRange.backgroundAC on\ filterByRange.backgroundAF on\ filterByRange.backgroundAN on\ filterByRange.inAffected on\ filterByRange.refLen on\ filterByRange.varLen on\ filterLabel.affectedAC Affected/case AC\ filterLabel.affectedAF Affected/case AF (pooled)\ filterLabel.affectedAN Affected/case AN (pool denominator)\ filterLabel.affectedCohorts Affected/case cohort\ filterLabel.altLen Alternate Length\ filterLabel.backgroundAC Background AC (population + unaffected)\ filterLabel.backgroundAF Background AF (pooled)\ filterLabel.backgroundAN Background AN (pool denominator)\ filterLabel.backgroundSources Background source (population or unaffected)\ filterLabel.consequence Consequence\ filterLabel.inAffected Seen in an affected/case arm (1=yes, 0=no)\ filterLabel.refLen Reference Length\ filterLabel.varLen Length Change\ filterLabel.varType Variant Type\ filterLimits.affectedAC 0:500000\ filterLimits.affectedAF 0:1\ filterLimits.affectedAN 0:500000\ filterLimits.altLen 1:6294\ filterLimits.backgroundAC 0:5000000\ filterLimits.backgroundAF 0:1\ filterLimits.backgroundAN 0:5000000\ filterLimits.inAffected 0:1\ filterLimits.refLen 1:28037\ filterLimits.varLen -28036:6293\ filterType.affectedCohorts multipleListOr\ filterType.backgroundSources multipleListOr\ filterType.consequence multipleListOr\ filterValues.affectedCohorts SPARK|SFARI SPARK WES,SFARI_WGS|SFARI SPARK WGS,GREGoR|GREGoR,SCHEMA|SCHEMA,GA4K|GA4K PacBio LR\ filterValues.backgroundSources AllOfUs|AllOfUs,SPARK|SFARI SPARK WES,SFARI_WGS|SFARI SPARK WGS,GenomeAsia|GenomeAsia SNVs,GenomeAsiaIndel|GenomeAsia Indels,NPM|NPM Singapore,KOVA|KOVA Korea,ToMMo|ToMMo Japan,FinnGen|FinnGen Finland,Saudi|Saudi,SweGen|SweGen Sweden,TOPMed|TOPMed,ABraOM|ABraOM Brazil,ALFA|ALFA,MGRB|MGRB Australia,HRC|HRC,SGDP|SGDP,HGDP1kG|gnomAD HGDP+1kG,GREGoR|GREGoR,SCHEMA|SCHEMA,CoLoRSdb|CoLoRSdb PacBio LR,SVatalog|SVatalog 101 10XG SR,Tishkoff180|Tishkoff 180 African WGS,WBBC|WBBC China,ChinaMAP|China ChinaMAP,GenomeIndia|GenomeIndia 9.7k WGS,GoNL|GoNL Netherlands ~13x SR\ filterValues.consequence missense|Missense,synonymous|Synonymous,stop_gained|Stop Gained,frameshift|Frameshift,splice_donor|Splice Donor,splice_acceptor|Splice Acceptor,intron|Intron,3_prime_utr|3' UTR,5_prime_utr|5' UTR,non_coding|Non-coding,.|Intergenic,others|Other\ filterValues.varType SNV|SNV,INS|Insertion,DEL|Deletion,MNV|MNV\ itemRgb on\ longLabel SNV Frequencies: variants in ~130,000 affected or case individuals (autism, schizophrenia, rare disease cohorts)\ maxWindowToDraw 5000000\ mouseOver Var: ${name}\ This track merges variants from three genotyping-array cohorts into a single bigBed file\ with predicted protein consequences and cross-database filtering. It contains 14.7 million\ variants from the Taiwan Precision Medicine Initiative (TPMI Axiom TPM1 chip,\ ~1 million Han Chinese), the Mexico Biobank (MexBB, 6,011 individuals), and the UK Biobank\ (361k unrelated white British, imputed from the Neale Lab Round 2 release).\
\ \\ The array track is kept separate from the sequencing-based combined tracks\ (Disease cohorts and\ Population reference) so that\ sequencing-based and array-based frequencies can be inspected independently. For a summary\ of all available variant frequency databases, see the\ SNV Frequencies supertrack page.\
\ \Variants are colored by their most severe predicted consequence:
\| Color | Consequence class | Examples |
|---|---|---|
| \ | Protein-truncating / loss-of-function | \stop_gained, frameshift, splice_donor, splice_acceptor, stop_lost, start_lost |
| \ | Missense / in-frame | \missense, inframe_insertion, inframe_deletion, protein_altering |
| \ | Synonymous | \synonymous, stop_retained |
| \ | Non-coding / intergenic | \intron, non_coding, intergenic, UTR |
\ The "AA change" field uses bcftools csq notation: 23I>23V means position\ 23 changed from Isoleucine (I) to Valine (V) (missense). 23I alone (no arrow)\ means position 23 is Isoleucine and unchanged (synonymous). A "*" indicates a\ stop codon (e.g. 45R>45* is a stop_gained).\
\ \\ Allele frequencies from genotyping arrays are not directly comparable to those from\ whole-genome or whole-exome sequencing. Two limitations to keep in mind:\
\NGS_concordance value (chip-vs-sequencing concordance from\
its own validation) in the source VCF; high-AF claims with low concordance are\
common. MexBB provides only AN/AF/AC with no FILTER column and no per-site QC.\
For both arrays, high-AF rare-disease candidates should be cross-checked against the\
sequencing-based\
Population reference track before\
drawing conclusions.\ This track supports filtering via the track settings page. Click the track title or use the\ "Configure" button to access filters.\
\ \\ The Source Database filter restricts the display to variants present in specific\ databases. It uses OR logic.\
\ \\
The same merge-and-annotate pipeline used for the sequencing-based combined tracks\
(Disease cohorts and\
Population reference) was run on the\
array-cohort subset of source VCFs. Each VCF was stripped of its INFO fields, normalized\
with bcftools norm (splitting multi-allelic sites), and merged with\
bcftools merge. The merged VCF was then annotated with predicted protein\
consequences using bcftools csq with the\
Ensembl\
GRCh38 release 115 gene annotation (GFF3).\
\ The track's\ makeDoc file documents how each source VCF was converted. Scripts are\ available from\ Github.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator. For programmatic access, our\ REST API can be used; the track\ name is varFreqsArray.\
\\ Because the merged callset includes data from multiple sources whose redistribution\ licenses differ, the combined bigBed is not available for download from our download\ server. The combined track can be reconstructed from the individual source VCFs using the\ conversion scripts on GitHub together with the\ build documentation.\
\ \\ This track is only possible thanks to the participants in TPMI, the Mexico Biobank, and UK\ Biobank, who donated samples and provided health information. Click on the individual\ TPMI, MexBB, or UK Biobank subtracks in the\ SNV Frequencies supertrack for full project credits.\ Thanks to Alex Ioannidis, UCSC, for the motivation for this track family and to Andreas\ Lahner, MGZ, for feedback.\
\ \\ For primary citations of each source dataset, see the References section on the\ SNV Frequencies supertrack page. The merged-track\ build itself uses the following tools:\
\\ Danecek P, McCarthy SA.\ \ BCFtools/csq: haplotype-aware variant consequences.\ Bioinformatics. 2017 Jul 1;33(13):2037-2039.\ PMID: 28205675; PMC: PMC5870570\
\\ McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GR, Thormann A, Flicek P, Cunningham F.\ \ The Ensembl Variant Effect Predictor.\ Genome Biol. 2016 Jun 6;17(1):122.\ PMID: 27268795; PMC: PMC4893825\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_array/varFreqsArray.bb\ filterByRange.MexBBAC on\ filterByRange.MexBBAF on\ filterByRange.TPMIAC on\ filterByRange.TPMIAF on\ filterByRange.UKBBAC on\ filterByRange.UKBBAF on\ filterByRange.altLen on\ filterByRange.maxAF on\ filterByRange.refLen on\ filterByRange.totalAC on\ filterByRange.varLen on\ filterLabel.MexBBAC Mexico Biobank AC\ filterLabel.MexBBAF Mexico Biobank AF\ filterLabel.TPMIAC TPMI Taiwan AC\ filterLabel.TPMIAF TPMI Taiwan AF\ filterLabel.UKBBAC UK Biobank imputed AC\ filterLabel.UKBBAF UK Biobank imputed AF\ filterLabel.altLen Alternate Length\ filterLabel.consequence Consequence\ filterLabel.maxAF Max Allele Frequency\ filterLabel.refLen Reference Length\ filterLabel.sources Source Database\ filterLabel.totalAC Total Allele Count (all databases)\ filterLabel.varLen Length Change\ filterLabel.varType Variant Type\ filterLimits.maxAF 0:1\ filterType.consequence multipleListOr\ filterType.sources multipleListOr\ filterValues.consequence missense|Missense,synonymous|Synonymous,stop_gained|Stop Gained,frameshift|Frameshift,splice_donor|Splice Donor,splice_acceptor|Splice Acceptor,intron|Intron,3_prime_utr|3' UTR,5_prime_utr|5' UTR,non_coding|Non-coding,.|Intergenic,others|Other\ filterValues.sources TPMI|TPMI Taiwan,MexBB|Mexico Biobank,UKBB|UK Biobank imputed\ filterValues.varType SNV|SNV,INS|Insertion,DEL|Deletion,MNV|MNV\ itemRgb on\ longLabel SNV Frequencies: Genotyping-array cohorts combined (TPMI, Mexico Biobank, UK Biobank imputed)\ maxWindowToDraw 5000000\ mouseOver Var: $name\ This track shows short genetic variants\ (up to approximately 50 base pairs) from\ dbSNP\ build 155:\ single-nucleotide variants (SNVs),\ small insertions, deletions, and complex deletion/insertions (indels),\ relative to the reference genome assembly.\ Most variants in dbSNP are rare, not true polymorphisms,\ and some variants are known to be pathogenic.\
\ For hg38 (GRCh38), approximately 998 million distinct variants\ (RefSNP clusters with rs# ids)\ have been mapped to more than 1.06 billion genomic locations\ including alternate haplotype and fix patch sequences.\ dbSNP remapped variants from hg38 to hg19 (GRCh37);\ approximately 981 million distinct variants were mapped to\ more than 1.02 billion genomic locations\ including alternate haplotype and fix patch sequences (not\ all of which are included in UCSC's hg19).\
\\ This track includes four subtracks of variants:\
\ A fifth subtrack highlights coordinate ranges to which dbSNP mapped a variant but with genomic\ coordinates that are not internally consistent, i.e. different coordinate ranges were provided\ when describing different alleles. This can occur due to a bug with mapping variants from one\ assembly sequence to another when there is an indel difference between the assembly sequences:\
\ SNVs and pure deletions are displayed as boxes covering the affected base(s).\ Pure insertions are drawn as single-pixel tickmarks between\ the base before and the base after the insertion.\
\ Insertions and/or deletions in repetitive regions may be represented by a half-height box\ showing uncertainty in placement, followed by a full-height box showing the number of deleted\ bases, or a full-height tickmark to indicate an insertion.\ When an insertion or deletion falls in a repetitive region, the placement may be ambiguous.\ For example, if the reference genome contains "TAAAG" but some\ individuals have "TAAG" at the same location, then the variant is a deletion of a single\ A relative to the reference genome.\ However, which A was deleted? There is no way to tell whether the first, second or third A\ was removed.\ Different variant mapping tools may place the deletion at different bases in the reference genome.\ To reduce errors in merging variant calls made with different left vs. right biases,\ dbSNP made a major change in its representation of deletion/insertion variants in build 152.\ Now, instead of assigning a single-base genomic location at one of the A's,\ dbSNP expands the coordinates to encompass the whole repetitive region,\ so the variant is represented as a deletion of 3 A's combined with an insertion of 2 A's.\ In the track display, there will be a half-height box covering the first two A's,\ followed by a full-height box covering the third A, to show a net loss of one base\ but an uncertain placement within the three A's.\
\\ When a variant has both insertion and deletion alternate alleles, the full-height box for the\ deletion(s) is drawn in a lighter shade so that the insertion tickmark is still visible.\
\\ Variants are colored according to functional effect on genes annotated by dbSNP:\
\ \Protein-altering variants and splice site variants are\
red.\
Synonymous codon variants are\
green.\
\
Non-coding transcript or Untranslated Region (UTR) variants are\
blue.\
\ On the track controls page, several variant properties can be included or excluded from\ the item labels:\ rs# identifier assigned by dbSNP,\ reference/alternate alleles,\ major/minor alleles (when available) and\ minor allele frequency (when available).\ Allele frequencies are reported independently by the project\ (some of which may have overlapping sets of samples):\
\ Using the track controls, variants can be filtered by\ \
\ While processing the information downloaded from dbSNP,\ UCSC annotates some properties of interest.\ These are noted on the item details page,\ and may be useful to include or exclude affected variants.\ \
\ Some are purely informational:\
\| keyword in data file (dbSnp155.bb) | \# in hg19 | # in hg38 | description |
|---|---|---|---|
| clinvar | \627817 | \630503 | \Variant is in ClinVar.\ | \
| clinvarBenign | \275541 | \276409 | \Variant is in ClinVar with clinical significance of benign and/or likely benign.\ | \
| clinvarConflicting | \16925 | \16834 | \Variant is in ClinVar with reports of both benign and pathogenic significance.\ | \
| clinvarPathogenic | \56373 | \56475 | \Variant is in ClinVar with clinical significance of pathogenic and/or likely pathogenic.\ | \
| commonAll | \14904503 | \15862783 | \Variant is "common", i.e. has a Minor Allele Frequency of at least 1% in all projects reporting frequencies.\ | \
| commonSome | \59633864 | \62095091 | \Variant is "common", i.e. has a Minor Allele Frequency of at least 1% in some, but not all, projects reporting frequencies.\ | \
| diffMajor | \12748733 | \13073288 | \Different frequency sources have different major alleles.\ | \
| overlapDiffClass | \198945442 | \207101421 | \This variant overlaps another variant with a different type/class.\ | \
| overlapSameClass | \29281958 | \30301090 | \This variant overlaps another with the same type/class but different start/end.\ | \
| rareAll | \906113910 | \938985356 | \Variant is "rare", i.e. has a Minor Allele Frequency of less than 1% in all projects reporting frequencies, or has no frequency data.\ | \
| rareSome | \950843271 | \985217664 | \Variant is "rare", i.e. has a Minor Allele Frequency of less than 1% in some, but not all, projects reporting frequencies, or has no frequency data.\ | \
| revStrand | \5540864 | \6770772 | \Alleles are displayed on the + strand at the current position. dbSNP's alleles are displayed on the + strand of a different assembly sequence, so dbSNP's variant page shows alleles that are reverse-complemented with respect to the alleles displayed above.\ | \
\ while others may indicate that the reference genome contains a rare variant or sequencing issue:\
\| keyword in data file (dbSnp155.bb) | \# in hg19 | # in hg38 | description |
|---|---|---|---|
| refIsAmbiguous | \19 | \41 | \The reference genome allele contains an IUPAC ambiguous base (e.g. 'R' for 'A or G', or 'N' for 'any base').\ | \
| refIsMinor | \14950212 | \15386394 | \The reference genome allele is not the major allele in at least one project.\ | \
| refIsRare | \793081 | \822757 | \The reference genome allele is rare (i.e. allele frequency < 1%).\ | \
| refIsSingleton | \694310 | \712794 | \The reference genome allele has never been observed in a population sequencing project reporting frequencies.\ | \
| refMismatch | \1 | \18 | \The reference genome allele reported by dbSNP differs from the GenBank assembly sequence. This is very rare and in all cases observed so far, the GenBank assembly has an 'N' while the RefSeq assembly used by dbSNP has a less ambiguous character such as 'R'.\ | \
\ and others may indicate an anomaly or problem with the variant data:\
\| keyword in data file (dbSnp155.bb) | \# in hg19 | # in hg38 | description |
|---|---|---|---|
| altIsAmbiguous | \5294 | \5361 | \At least one alternate allele contains an IUPAC ambiguous base (e.g. 'R' for 'A or G'). For alleles containing more than one ambiguous base, this may create a combinatoric explosion of possible alleles.\ | \
| classMismatch | \13289 | \18475 | \Variation class/type is inconsistent with alleles mapped to this genome assembly.\ | \
| clusterError | \373258 | \459130 | \This variant has the same start, end and class as another variant; they probably should have been merged into one variant.\ | \
| freqIncomplete | \0 | \0 | \At least one project reported counts for only one allele which implies that at least one allele is missing from the report; that project's frequency data are ignored.\ | \
| freqIsAmbiguous | \4332 | \4399 | \At least one allele reported by at least one project that reports frequencies contains an IUPAC ambiguous base.\ | \
| freqNotMapped | \1149972 | \1141935 | \At least one project reported allele frequencies relative to a different assembly; However, dbSNP does not include a mapping of this variant to that assembly, which implies a problem with mapping the variant across assemblies. The mapping on this assembly may have an issue; evaluate carefully vs. original submissions, which you can view by clicking through to dbSNP above.\ | \
| freqNotRefAlt | \74139 | \110646 | \At least one allele reported by at least one project that reports frequencies does not match any of the reference or alternate alleles listed by dbSNP.\ | \
| multiMap | \799777 | \286666 | \This variant has been mapped to more than one distinct genomic location.\ | \
| otherMapErr | \91260 | \195051 | \At least one other mapping of this variant has erroneous coordinates. The mapping(s) with erroneous coordinates are excluded from this track and are included in the Map Err subtrack. Sometimes despite this mapping having legal coordinates, there may still be an issue with this mapping's coordinates and alleles; you may want to click through to dbSNP to compare the initial submission's coordinates and alleles. In hg19, 55454 distinct rsIDs are affected; in hg38, 86636. \ | \
\ dbSNP has collected genetic variant reports from researchers worldwide for \ more than 20 years.\ Since the advent of next-generation sequencing methods and the population sequencing efforts\ that they enable, dbSNP has grown exponentially, requiring a new data schema, computational pipeline,\ web infrastructure, and download files.\ (Holmes et al.)\ The same challenges of exponential growth affected UCSC's presentation of dbSNP variants,\ so we have taken the opportunity to change our internal representation and import pipeline.\ Most notably, flanking sequences are no longer provided by dbSNP,\ because most submissions have been genomic variant calls in VCF format as opposed to\ independent sequences.\
\\ We downloaded JSON files available from dbSNP at\ https://ftp.ncbi.nlm.nih.gov/snp/archive/b155/JSON/,\ extracted a subset of the information about each variant, and collated\ it into a bigBed file using the\ bigDbSnp.as schema with the information\ necessary for filtering and displaying the variants,\ as well as a separate file containing more detailed information to be\ displayed on each variant's details page\ (dbSnpDetails.as schema).\ \
\ Note: It is not recommended to use LiftOver to convert SNPs between assemblies,\ and more information about how to convert SNPs between assemblies can be found on the following\ FAQ entry.
\\ Since dbSNP has grown to include over 1 billion variants, the size of the All dbSNP (155)\ subtrack can cause the\ Table Browser and\ Data Integrator\ to time out, leading to a blank page or truncated output,\ unless queries are restricted to a chromosomal region, to particular defined regions, to a specific set \ of rs# IDs (which can be pasted/uploaded into the Table Browser),\ or to one of the subset tracks such as Common (~15 million variants) or ClinVar (~0.8M variants).\
\ For automated analysis, the track data files can be downloaded from the downloads server for\ hg19 and\ hg38.\
| file | \format | \subtrack | \||
|---|---|---|---|---|
| dbSnp155.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \All dbSNP (155) | \
| dbSnp155ClinVar.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \ClinVar dbSNP (155) | \
| dbSnp155Common.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \Common dbSNP (155) | \
| dbSnp155Mult.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \Mult. dbSNP (155) | \
| dbSnp155BadCoords.bb | \hg19 | \hg38 | \bigBed4 | \Map Err (155) | \
| \ dbSnp155Details.tab.gz\ | \gzip-compressed tab-separated text | \Detailed variant properties, independent of genome assembly version | \||
\ Several utilities for working with bigBed-formatted binary files can be downloaded\ here.\ Run a utility with no arguments to see a brief description of the utility and its options.\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/snp/dbSnp155.bb -chrom=chr1 -start=200000 -end=200400 stdout\ \
bigBedNamedItems dbSnp155.bb rs6657048 stdout\ \
bigBedNamedItems -nameFile dbSnp155.bb myIds.txt dbSnp155.myIds.bed\ \
\ The columns in the bigDbSnp/bigBed files and dbSnp155Details.tab.gz file are described in\ bigDbSnp.as and\ dbSnpDetails.as respectively.\ \ For columns that contain lists of allele frequency data, the order of projects\ providing the data listed is as follows:\
\ UCSC also has an\ API\ that can be used to retrieve values from a particular chromosome range.\
\ A list of rs# IDs can be pasted/uploaded in the\ Variant Annotation Integrator\ tool to find out which genes (if any) the variants are located in,\ as well as functional effect such as intron, coding-synonymous, missense, frameshift, etc.\
\ Please refer to our searchable\ mailing list archives\ for more questions and example queries, or our\ Data Access FAQ\ for more information.\
\ \\ Holmes JB, Moyer E, Phan L, Maglott D, Kattman B.\ \ SPDI: Data Model for Variants and Applications at NCBI.\ Bioinformatics. 2019 Nov 18;.\ PMID: 31738401\
\\ Sayers EW, Agarwala R, Bolton EE, Brister JR, Canese K, Clark K, Connor R, Fiorini N, Funk K,\ Hefferon T et al.\ \ Database resources of the National Center for Biotechnology Information.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D23-D28.\ PMID: 30395293; PMC: PMC6323993\
\\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122;\ PMC: PMC29783\
\ \ varRep 1 compositeTrack on\ group varRep\ longLabel Short Genetic Variants from dbSNP release 155\ maxWindowCoverage 4000000\ priority 0.8\ shortLabel dbSNP 155\ subGroup1 view Views variants=Variants errs=Mapping_Errors\ track dbSnp155Composite\ type bed 3\ url https://www.ncbi.nlm.nih.gov/snp/$$\ urlLabel dbSNP:\ visibility pack\ cCREs ENCODE cCREs ENCODE Registry of cCREs (candidate Cis-Regulatory Elements) 0 0.8 0 0 0 127 127 127 0 0 0\ This track collection displays candidate Cis-Regulatory Elements (cCREs) generated by the \ ENCODE Consortium during Phase 4 (ENCODE4) and Phase 3 (ENCODE3), with the ENCODE3 track \ retained for archival purposes. The tracks include both integrated (biosample-agnostic) and \ biosample-specific annotations derived from core epigenomic assays.
\ \\ For information on track configuration, data description, data access, methods, and data provenance, \ see the individual track description pages via their links above
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J, Kawli T,\ Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N, Fu Y et\ al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 group regulation\ html cCREsSuper.html\ longLabel ENCODE Registry of cCREs (candidate Cis-Regulatory Elements)\ priority 0.8\ shortLabel ENCODE cCREs\ superTrack on show\ track cCREs\ dbSnp155ViewErrs Mapping Errors bed 3 Short Genetic Variants from dbSNP release 155 1 0.8 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$ varRep 1 longLabel Short Genetic Variants from dbSNP release 155\ parent dbSnp155Composite\ shortLabel Mapping Errors\ track dbSnp155ViewErrs\ view errs\ visibility dense\ dbSnp155ViewVariants Variants bigDbSnp Short Genetic Variants from dbSNP release 155 1 0.8 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$ varRep 1 classFilterType multipleListOr\ classFilterValues snv,mnv,ins,del,delins,identity\ detailsTabUrls _dataOffset=/gbdb/hgFixed/dbSnp/dbSnp155Details.tab.gz\ freqSourceOrder 1000Genomes,dbGaP_PopFreq,TOPMED,KOREAN,SGDP_PRJ,Qatari,NorthernSweden,Siberian,TWINSUK,TOMMO,ALSPAC,GENOME_DK,GnomAD,GoNL,Estonian,Vietnamese,Korea1K,HapMap,PRJEB36033,HGDP_Stanford,Daghestan,PAGE_STUDY,Chileans,MGP,PRJEB37584,GoESP,ExAC,GnomAD_exomes,FINRISK,PharmGKB,PRJEB37766\ longLabel Short Genetic Variants from dbSNP release 155\ maxFuncImpactFilterLabel Greatest functional impact on gene\ maxFuncImpactFilterType multipleListOr\ maxFuncImpactFilterValues 0|(not annotated),1589|frameshift,1587|stop_gained,1574|splice_acceptor_variant,1575|splice_donor_variant,1821|inframe_insertion,1583|missense_variant,1590|terminator_codon_variant,1819|synonymous_variant,1580|coding_sequence_variant,1623|5_prime_UTR_variant,1624|3_prime_UTR_variant,1619|nc_transcript_variant,2|genic_upstream_transcript_variant,1986|upstream_transcript_variant,2152|genic_downstream_transcript_variant,1987|downstream_transcript_variant,1627|intron_variant\ parent dbSnp155Composite\ shortLabel Variants\ track dbSnp155ViewVariants\ type bigDbSnp\ ucscNotesFilterType multipleListOr\ ucscNotesFilterValues altIsAmbiguous|Alternate allele contains IUPAC ambiguous base(s),classMismatch|Variant class/type is inconsistent with allele sizes,clinvar|Present in ClinVar,clinvarBenign|ClinVar significance of benign and/or likely benign,clinvarConflicting|ClinVar includes both benign and pathogenic reports,clinvarPathogenic|ClinVar significance of pathogenic and/or likely pathogenic,clusterError|Overlaps a variant with the same type/class and position,commonAll|MAF >= 1% in all projects that report frequencies,commonSome|MAF >= 1% in at least one project that reports frequencies,diffMajor|Different projects report different major alleles,freqIncomplete|Frequency reported with incomplete allele data,freqIsAmbiguous|Frequency reported for allele with IUPAC ambiguous base(s),freqNotMapped|Frequency reported on different assembly but not mapped by dbSNP,freqNotRefAlt|Reference genome allele is not major allele in at least one project,multiMap|Variant is placed in more than one genomic position,otherMapErr|Another mapping of this variant has illegal coords (indel mapping error?),overlapDiffClass|Variant overlaps other variant(s) of different type/class,overlapSameClass|Variant overlaps other variant(s) of same type/class but different position,rareAll|MAF < 1% in all projects that report frequencies (or no frequency data),rareSome|MAF < 1% in at least one project that reports frequencies,refIsAmbiguous|Reference genome allele contains IUPAC ambiguous base(s),refIsMinor|Reference genome allele is minor allele in at least one project that reports frequencies,refIsRare|Reference genome allele frequency is <1% in at least one project,refIsSingleton|Reference genome frequency is 0 in all projects that report frequencies,refMismatch|Reference allele mismatches reference genome sequence,revStrand|Variant maps to opposite strand relative to dbSNP's preferred top-level placement\ view variants\ visibility dense\ wgEncodeReg4 ENCODE4 Regulation Integrated Regulation from ENCODE 4 0 0.9 0 0 0 127 127 127 0 0 0\ This collection of tracks offers an integrated view of genomic annotations and experimental\ data from all phases of the\ ENCODE Project,\ with a focus on transcriptional regulation. It includes averaged and representative signals\ from assays that measure chromatin accessibility (DNase-seq and ATAC-seq), transcription\ factor (TF) binding (ChIP-seq for individual TFs), histone modifications (ChIP-seq for\ H3K4me3 and H3K27ac), CTCF binding, and transcription (RNA-seq).
\ \Tracks labeled (Layered) show organ-averaged signals as a transparent\ overlay of multiple organs within a single track. Tracks labeled (Indiv.)\ show signals from individual experiments in specific biosamples.
\ \\ These tracks complement one another and collectively provide a resource for\ interpreting regulatory DNA. Histone marks are broadly informative but have limited resolution\ (~200 bp) and relatively low functional specificity. DNase-seq assays offer higher resolution\ and scalability across many cell types, and they reliably indicate regulatory potential, though\ they lack detailed functional context. ATAC-seq serves a similar role to DNase-seq, with\ comparable resolution and limitations. Transcription factor ChIP-seq has high positional\ resolution and, due to the specificity of TFs, often provides more direct functional insight.\ However, because each TF must be assayed individually, the data are limited in biosample\ coverage. Despite the individual strengths and limitations of these assays, their independence\ from one another increases confidence when multiple assays suggest a regulatory function for\ the same genomic region.
\ \\ For additional information, click on the hyperlinks for the individual tracks above.\ Additional histone marks and transcription data are available in other ENCODE tracks. This\ integrative supertrack presents a curated selection of the most informative and broadly\ relevant datasets. Further functional annotations of individual regulatory elements are\ available at SCREEN.
\ \\ By default, the DNase (Layered), ATAC (Layered),\ H3K4me3 (Layered), H3K27ac (Layered), CTCF (Layered), and\ Transcription (Layered) tracks use a transparent overlay to visualize signals from\ multiple organs or tissues within a single track. For each organ or tissue, signal values from\ all associated experiments are averaged. Each organ or tissue is assigned a distinct color,\ selected to be light and saturated to maintain clarity when overlaid. Initially, each layered\ track displays an overlay of representative organs: blood, brain, kidney, liver, and\ muscle (the ATAC track has no kidney data). Clicking on the track opens a details page where you can view and select organs or\ tissues.
\ \\ For the TF rPeaks track, each rPeak (representative peak) is colored in\ grayscale by the maximum ChIP-seq signal for the corresponding TF across all contributing\ biosamples (darker = higher signal, score 0 to 1,000). The HGNC gene symbol of the TF is\ displayed to the left when viewed in pack display mode. If the rPeak overlaps a\ cognate TF motif from a previously curated collection (Andrews et al., 2023), the motif\ site is colored green using decorators.
\ \\ The TF ChIP-seq (Indiv.), DNase/ATAC/Histone/CTCF (Indiv.), and\ RNA-seq (Indiv.) tracks are hidden by default. Clicking on any of these tracks opens\ a details page where you can select specific biosample-level experiments to display.
\ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the track data files can be downloaded from\ our download server or queried using the\ REST API.\ Individual regions or the whole genome annotation can be accessed as text using\ our utilities bigWigToWig and bigBedToBed. Instructions for\ downloading source code and binaries can be found\ here.\ The original data files are also available from the\ ENCODE portal.
\ \\ Data were generated by the ENCODE Consortium. The data were further processed for visualization\ through a collaborative effort between the\ Weng lab and the\ Moore lab\ at UMass Chan Medical School (funded by NIH grant HG012343). Integration and visualization\ were developed by Drs. Mingshi Gao, Greg Andrews, Jill Moore, and Zhiping Weng at UMass Chan\ Medical School, who were part of the ENCODE Data Analysis Center.
\ \\ Users may freely download, analyze, and publish results based on any ENCODE data without\ restrictions.\ Researchers using unpublished ENCODE data are encouraged to contact the data producers to\ discuss possible coordinated publications; however, this is optional.
\\ Users of ENCODE datasets are requested to cite the ENCODE Consortium and ENCODE\ production laboratory(s) that generated the datasets used, as described in\ Citing\ ENCODE.
\ \\ Andrews G, Fan K, Pratt HE, Phalke N, Zoonomia Consortium, Karlsson EK, Lindblad-Toh K,\ Weng Z.\ \ Mammalian evolution of human cis-regulatory elements and transcription factor binding\ sites.\ Science. 2023;380(6643):eabn7930.\ PMID: 37104580\
\\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 group regulation\ html wgEncodeReg4.html\ longLabel Integrated Regulation from ENCODE 4\ pennantIcon New red\ priority 0.9\ shortLabel ENCODE4 Regulation\ superTrack on hide\ track wgEncodeReg4\ dbSnp153Composite dbSNP 153 bed 6 + Short Genetic Variants from dbSNP release 153 3 0.908 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$\ This track shows short genetic variants\ (up to approximately 50 base pairs) from\ dbSNP\ build 153:\ single-nucleotide variants (SNVs),\ small insertions, deletions, and complex deletion/insertions (indels),\ relative to the reference genome assembly.\ Most variants in dbSNP are rare, not true polymorphisms,\ and some variants are known to be pathogenic.\
\ For hg38 (GRCh38), approximately 667 million distinct variants\ (RefSNP clusters with rs# ids)\ have been mapped to more than 702 million genomic locations\ including alternate haplotype and fix patch sequences.\ dbSNP remapped variants from hg38 to hg19 (GRCh37);\ approximately 658 million distinct variants were mapped to\ more than 683 million genomic locations\ including alternate haplotype and fix patch sequences (not\ all of which are included in UCSC's hg19).\
\\ This track includes four subtracks of variants:\
\ A fifth subtrack highlights coordinate ranges to which dbSNP mapped a variant but with genomic\ coordinates that are not internally consistent, i.e. different coordinate ranges were provided\ when describing different alleles. This can occur due to a bug with mapping variants from one\ assembly sequence to another when there is an indel difference between the assembly sequences:\
\ SNVs and pure deletions are displayed as boxes covering the affected base(s).\ Pure insertions are drawn as single-pixel tickmarks between\ the base before and the base after the insertion.\
\ Insertions and/or deletions in repetitive regions may be represented by a half-height box\ showing uncertainty in placement, followed by a full-height box showing the number of deleted\ bases, or a full-height tickmark to indicate an insertion.\ When an insertion or deletion falls in a repetitive region, the placement may be ambiguous.\ For example, if the reference genome contains "TAAAG" but some\ individuals have "TAAG" at the same location, then the variant is a deletion of a single\ A relative to the reference genome.\ However, which A was deleted? There is no way to tell whether the first, second or third A\ was removed.\ Different variant mapping tools may place the deletion at different bases in the reference genome.\ To reduce errors in merging variant calls made with different left vs. right biases,\ dbSNP made a major change in its representation of deletion/insertion variants in build 152.\ Now, instead of assigning a single-base genomic location at one of the A's,\ dbSNP expands the coordinates to encompass the whole repetitive region,\ so the variant is represented as a deletion of 3 A's combined with an insertion of 2 A's.\ In the track display, there will be a half-height box covering the first two A's,\ followed by a full-height box covering the third A, to show a net loss of one base\ but an uncertain placement within the three A's.\
\\ Variants are colored according to functional effect on genes annotated by dbSNP:\
\ \Protein-altering variants and splice site variants are\
red.\
Synonymous codon variants are\
green.\
\
Non-coding transcript or Untranslated Region (UTR) variants are\
blue.\
\ On the track controls page, several variant properties can be included or excluded from\ the item labels:\ rs# identifier assigned by dbSNP,\ reference/alternate alleles,\ major/minor alleles (when available) and\ minor allele frequency (when available).\ Allele frequencies are reported independently by twelve projects\ (some of which may have overlapping sets of samples):\
\ Using the track controls, variants can be filtered by\ \
\ While processing the information downloaded from dbSNP,\ UCSC annotates some properties of interest.\ These are noted on the item details page,\ and may be useful to include or exclude affected variants.\
\ Some are purely informational:\
\| keyword in data file (dbSnp153.bb) | \# in hg19 | # in hg38 | description |
|---|---|---|---|
| clinvar | \454678 | \453996 | \Variant is in ClinVar. | \
| clinvarBenign | \143864 | \143736 | \Variant is in ClinVar with clinical significance of benign and/or likely benign. | \
| clinvarConflicting | \7932 | \7950 | \Variant is in ClinVar with reports of both benign and pathogenic significance. | \
| clinvarPathogenic | \96242 | \95262 | \Variant is in ClinVar with clinical significance of pathogenic and/or likely pathogenic. | \
| commonAll | \12184521 | \12438655 | \Variant is "common", i.e. has a Minor Allele Frequency of at least 1% in all\ projects reporting frequencies. | \
| commonSome | \20541190 | \20902944 | \Variant is "common", i.e. has a Minor Allele Frequency of at least 1% in some, but not all,\ projects reporting frequencies. | \
| diffMajor | \1377831 | \1399109 | \Different frequency sources have different major alleles. | \
| overlapDiffClass | \107015341 | \110007682 | \This variant overlaps another variant with a different type/class. | \
| overlapSameClass | \16915239 | \17291289 | \This variant overlaps another with the same type/class but different start/end. | \
| rareAll | \662601770 | \681696398 | \Variant is "rare", i.e. has a Minor Allele Frequency of less than 1%\ in all projects reporting frequencies, or has no frequency data. | \
| rareSome | \670958439 | \690160687 | \Variant is "rare", i.e. has a Minor Allele Frequency of less than 1%\ in some, but not all, projects reporting frequencies, or has no frequency data. | \
| revStrand | \3813702 | \4532511 | \Alleles are displayed on the + strand at the current position.\ dbSNP's alleles are displayed on the + strand of a different assembly sequence,\ so dbSNP's variant page shows alleles that are reverse-complemented with respect to\ the alleles displayed above. | \
\ while others may indicate that the reference genome contains a rare variant or sequencing issue:\
\| keyword in data file (dbSnp153.bb) | \# in hg19 | # in hg38 | description |
|---|---|---|---|
| refIsAmbiguous | \101 | \111 | \The reference genome allele contains an IUPAC ambiguous base\ (e.g. 'R' for 'A or G', or 'N' for 'any base'). | \
| refIsMinor | \3272116 | \3360435 | \The reference genome allele is not the major allele in at least one project. | \
| refIsRare | \136547 | \160827 | \The reference genome allele is rare (i.e. allele frequency < 1%). | \
| refIsSingleton | \37832 | \50927 | \The reference genome allele has never been observed in a population sequencing project\ reporting frequencies. | \
| refMismatch | \4 | \33 | \The reference genome allele reported by dbSNP differs from the GenBank assembly sequence.\ This is very rare and in all cases observed so far, the GenBank assembly has an 'N'\ while the RefSeq assembly used by dbSNP has a less ambiguous character such as 'R'. | \
\ and others may indicate an anomaly or problem with the variant data:\
\| keyword in data file (dbSnp153.bb) | \# in hg19 | # in hg38 | description |
|---|---|---|---|
| altIsAmbiguous | \10755 | \10888 | \At least one alternate allele contains an IUPAC ambiguous base (e.g. 'R' for 'A or G').\ For alleles containing more than one ambiguous base, this may create a\ combinatoric explosion of possible alleles. | \
| classMismatch | \5998 | \6216 | \Variation class/type is inconsistent with alleles mapped to this genome assembly. | \
| clusterError | \114826 | \128306 | \This variant has the same start, end and class as another variant;\ they probably should have been merged into one variant. | \
| freqIncomplete | \3922 | \4673 | \At least one project reported counts for only one allele which implies that at\ least one allele is missing from the report;\ that project's frequency data are ignored. | \
| freqIsAmbiguous | \7656 | \7756 | \At least one allele reported by at least one project that reports frequencies\ contains an IUPAC ambiguous base. | \
| freqNotMapped | \2685 | \6590 | \At least one project reported allele frequencies relative to a different assembly;\ However, dbSNP does not include a mapping of this variant to that assembly, which\ implies a problem with mapping the variant across assemblies. The mapping on this\ assembly may have an issue; evaluate carefully vs. original submissions, which you\ can view by clicking through to dbSNP above. | \
| freqNotRefAlt | \17694 | \32170 | \At least one allele reported by at least one project that reports frequencies\ does not match any of the reference or alternate alleles listed by dbSNP. | \
| multiMap | \562180 | \132123 | \This variant has been mapped to more than one distinct genomic location. | \
| otherMapErr | \114095 | \204219 | \At least one other mapping of this variant has erroneous coordinates.\ The mapping(s) with erroneous coordinates are excluded from this track\ and are included in the Map Err subtrack. Sometimes despite this mapping\ having legal coordinates, there may still be an issue with this mapping's\ coordinates and alleles; you may want to click through to dbSNP to compare\ the initial submission's coordinates and alleles.\ In hg19, 55454 distinct rsIDs are affected; in hg38, 86636.\ |
\ dbSNP has collected genetic variant reports from researchers worldwide for \ more than 20 years.\ Since the advent of next-generation sequencing methods and the population sequencing efforts\ that they enable, dbSNP has grown exponentially, requiring a new data schema, computational pipeline,\ web infrastructure, and download files.\ (Holmes et al.)\ The same challenges of exponential growth affected UCSC's presentation of dbSNP variants,\ so we have taken the opportunity to change our internal representation and import pipeline.\ Most notably, flanking sequences are no longer provided by dbSNP,\ because most submissions have been genomic variant calls in VCF format as opposed to\ independent sequences.\
\\ We downloaded JSON files available from dbSNP at\ ftp://ftp.ncbi.nlm.nih.gov/snp/archive/b153/JSON/,\ extracted a subset of the information about each variant, and collated\ it into a bigBed file using the\ bigDbSnp.as schema with the information\ necessary for filtering and displaying the variants,\ as well as a separate file containing more detailed information to be\ displayed on each variant's details page\ (dbSnpDetails.as schema).\ \
\ Note: It is not recommended to use LiftOver to convert SNPs between assemblies,\ and more information about how to convert SNPs between assemblies can be found on the following\ FAQ entry.
\\ Since dbSNP has grown to include approximately 700 million variants, the size of the All dbSNP (153)\ subtrack can cause the\ Table Browser and\ Data Integrator\ to time out, leading to a blank page or truncated output,\ unless queries are restricted to a chromosomal region, to particular defined regions, to a specific set \ of rs# IDs (which can be pasted/uploaded into the Table Browser),\ or to one of the subset tracks such as Common (~15 million variants) or ClinVar (~0.5M variants).\
\ For automated analysis, the track data files can be downloaded from the downloads server for\ hg19 and\ hg38.\
| file | \format | \subtrack | \||
|---|---|---|---|---|
| dbSnp153.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \All dbSNP (153) | \
| dbSnp153ClinVar.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \ClinVar dbSNP (153) | \
| dbSnp153Common.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \Common dbSNP (153) | \
| dbSnp153Mult.bb | \hg19 | \hg38 | \bigDbSnp (bigBed4+13) | \Mult. dbSNP (153) | \
| dbSnp153BadCoords.bb | \hg19 | \hg38 | \bigBed4 | \Map Err (153) | \
| \ dbSnp153Details.tab.gz\ | \gzip-compressed tab-separated text | \Detailed variant properties, independent of genome assembly version | \||
\ Several utilities for working with bigBed-formatted binary files can be downloaded\ here.\ Run a utility with no arguments to see a brief description of the utility and its options.\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/snp/dbSnp153.bb -chrom=chr1 -start=200000 -end=200400 stdout\ \
bigBedNamedItems dbSnp153.bb rs6657048 stdout\ \
bigBedNamedItems -nameFile dbSnp153.bb myIds.txt dbSnp153.myIds.bed\ \
\ The columns in the bigDbSnp/bigBed files and dbSnp153Details.tab.gz file are described in\ bigDbSnp.as and\ dbSnpDetails.as respectively.\ For columns that contain lists of allele frequency data, the order of projects\ providing the data listed is as follows:\
\ UCSC also has an\ API\ that can be used to retrieve values from a particular chromosome range.\
\ A list of rs# IDs can be pasted/uploaded in the\ Variant Annotation Integrator\ tool to find out which genes (if any) the variants are located in,\ as well as functional effect such as intron, coding-synonymous, missense, frameshift, etc.\
\ Please refer to our searchable\ mailing list archives\ for more questions and example queries, or our\ Data Access FAQ\ for more information.\
\ \\ Holmes JB, Moyer E, Phan L, Maglott D, Kattman B.\ \ SPDI: Data Model for Variants and Applications at NCBI.\ Bioinformatics. 2019 Nov 18;.\ PMID: 31738401\
\\ Sayers EW, Agarwala R, Bolton EE, Brister JR, Canese K, Clark K, Connor R, Fiorini N, Funk K,\ Hefferon T et al.\ \ Database resources of the National Center for Biotechnology Information.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D23-D28.\ PMID: 30395293; PMC: PMC6323993\
\\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122;\ PMC: PMC29783\
\ \ varRep 1 compositeTrack on\ group varRep\ html ../dbSnp153Composite\ longLabel Short Genetic Variants from dbSNP release 153\ maxWindowCoverage 4000000\ parent dbSnpArchive on\ priority 0.908\ shortLabel dbSNP 153\ subGroup1 view Views variants=Variants errs=Mapping_Errors\ track dbSnp153Composite\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/snp/$$\ urlLabel dbSNP:\ visibility pack\ dbSnp153ViewErrs Mapping Errors bed 6 + Short Genetic Variants from dbSNP release 153 1 0.908 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$ varRep 1 longLabel Short Genetic Variants from dbSNP release 153\ parent dbSnp153Composite\ shortLabel Mapping Errors\ track dbSnp153ViewErrs\ view errs\ visibility dense\ dbSnp153ViewVariants Variants bigDbSnp Short Genetic Variants from dbSNP release 153 1 0.908 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$ varRep 1 classFilterType multipleListOr\ classFilterValues snv,mnv,ins,del,delins,identity\ detailsTabUrls _dataOffset=/gbdb/hgFixed/dbSnp/dbSnp153Details.tab.gz\ freqSourceOrder 1000Genomes,GnomAD_exomes,TOPMED,ExAC,PAGE_STUDY,GnomAD,GoESP,Estonian,ALSPAC,TWINSUK,NorthernSweden,Vietnamese\ longLabel Short Genetic Variants from dbSNP release 153\ maxFuncImpactFilterLabel Greatest functional impact on gene\ maxFuncImpactFilterType multipleListOr\ maxFuncImpactFilterValues 0|(not annotated),865|frameshift,1587|stop_gained,1574|splice_acceptor_variant,1575|splice_donor_variant,1821|inframe_insertion,1583|missense_variant,1590|terminator_codon_variant,1819|synonymous_variant,1580|coding_sequence_variant,1623|5_prime_UTR_variant,1624|3_prime_UTR_variant,1619|nc_transcript_variant,2153|genic_upstream_transcript_variant,1986|upstream_transcript_variant,2152|genic_downstream_transcript_variant,1987|downstream_transcript_variant,1627|intron_variant\ parent dbSnp153Composite\ shortLabel Variants\ showCfg on\ track dbSnp153ViewVariants\ type bigDbSnp\ ucscNotesFilterType multipleListOr\ ucscNotesFilterValues altIsAmbiguous|Alternate allele contains IUPAC ambiguous base(s),classMismatch|Variant class/type is inconsistent with allele sizes,clinvar|Present in ClinVar,clinvarBenign|ClinVar significance of benign and/or likely benign,clinvarConflicting|ClinVar includes both benign and pathogenic reports,clinvarPathogenic|ClinVar significance of pathogenic and/or likely pathogenic,clusterError|Overlaps a variant with the same type/class and position,commonAll|MAF >= 1% in all projects that report frequencies,commonSome|MAF >= 1% in at least one project that reports frequencies,diffMajor|Different projects report different major alleles,freqIncomplete|Frequency reported with incomplete allele data,freqIsAmbiguous|Frequency reported for allele with IUPAC ambiguous base(s),freqNotMapped|Frequency reported on different assembly but not mapped by dbSNP,freqNotRefAlt|Reference genome allele is not major allele in at least one project,multiMap|Variant is placed in more than one genomic position,otherMapErr|Another mapping of this variant has illegal coords (indel mapping error?),overlapDiffClass|Variant overlaps other variant(s) of different type/class,overlapSameClass|Variant overlaps other variant(s) of same type/class but different position,rareAll|MAF < 1% in all projects that report frequencies (or no frequency data),rareSome|MAF < 1% in at least one project that reports frequencies,refIsAmbiguous|Reference genome allele contains IUPAC ambiguous base(s),refIsMinor|Reference genome allele is minor allele in at least one project that reports frequencies,refIsRare|Reference genome allele frequency is <1% in at least one project,refIsSingleton|Reference genome frequency is 0 in all projects that report frequencies,refMismatch|Reference allele mismatches reference genome sequence,revStrand|Variant maps to opposite strand relative to dbSNP's preferred top-level placement\ view variants\ visibility dense\ snp151Common Common SNPs(151) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 151) Found in >= 1% of Samples 0 0.909 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 151, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs that have a minor allele frequency (MAF) of at least 1% and\ are mapped to a single location in the reference genome assembly are\ included in this subset. Frequency data are not available for all SNPs,\ so this subset is incomplete.\ Allele counts from all submissions that include frequency data are combined\ when determining MAF, so for example the allele counts from\ the 1000 Genomes Project and an independent submitter may be combined for the\ same variant.\
\\ dbSNP provides\ download files\ in the\ Variant Call Format (VCF)\ that include a "COMMON" flag in the INFO column. That is determined by a different method,\ and is generally a superset of the UCSC Common set.\ dbSNP uses frequency data from the\ 1000 Genomes Project\ only, and considers a variant COMMON if it has a MAF of at least 0.01 in any of the five\ super-populations:\
\ The selection of SNPs with a minor allele frequency of 1% or greater\ is an attempt to identify variants that appear to be reasonably common\ in the general population. Taken as a set, common variants should be\ less likely to be associated with severe genetic diseases due to the\ effects of natural selection,\ following the view that deleterious variants are not likely to become\ common in the population.\ However, the significance of any particular variant should be interpreted\ only by a trained medical geneticist using all available information.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp151*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp151OrthoPt5Pa2Rm8\ codingAnnotations snp151CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp151Common\ longLabel Simple Nucleotide Polymorphisms (dbSNP 151) Found in >= 1% of Samples\ macaqueDb rheMac8\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.909\ shortLabel Common SNPs(151)\ snpExceptionDesc snp151ExceptionDesc\ snpSeq snp151Seq\ track snp151Common\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp151 All SNPs(151) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 151) 0 0.91 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 151, available from\ ftp.ncbi.nlm.nih.gov/snp.\
\\ Three tracks contain subsets of the items in this track:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic\ locations will be omitted from display. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp151*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp151OrthoPt5Pa2Rm8\ codingAnnotations snp151CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp151\ longLabel Simple Nucleotide Polymorphisms (dbSNP 151)\ macaqueDb rheMac8\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.910\ shortLabel All SNPs(151)\ tableBrowser noGenome\ track snp151\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp151Flagged Flagged SNPs(151) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 151) Flagged by dbSNP as Clinically Assoc 0 0.911 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 151, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs flagged as clinically associated by dbSNP,\ mapped to a single location in the reference genome assembly, and\ not known to have a minor allele frequency of at\ least 1%, are included in this subset.\ Frequency data are not available for all SNPs, so this subset probably\ includes some SNPs whose true minor allele frequency is 1% or greater.\
\\ The significance of any particular variant in this track should be\ interpreted only by a trained medical geneticist using all available\ information. For example, some variants are included in this track\ because of their inclusion in a Locus-Specific Database (LSDB) or\ mention in OMIM, but are not thought to be disease-causing, so\ inclusion of a variant in this track is not necessarily an indicator\ of risk. Again, all available information must be carefully considered\ by a qualified professional.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b151_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp151*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp151OrthoPt5Pa2Rm8\ codingAnnotations snp151CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html snp151Flagged\ longLabel Simple Nucleotide Polymorphisms (dbSNP 151) Flagged by dbSNP as Clinically Assoc\ macaqueDb rheMac8\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.911\ shortLabel Flagged SNPs(151)\ snpExceptionDesc snp151ExceptionDesc\ snpSeq snp151Seq\ track snp151Flagged\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp151Mult Mult. SNPs(151) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 151) That Map to Multiple Genomic Loci 0 0.912 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 150, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs that have been mapped to multiple locations in the reference\ genome assembly are included in this subset. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\\ Since build 149, dbSNP has been filtering out almost all such "SNPs" so\ there are very few items in this track.\
\\ The default maximum weight for this track is 3,\ unlike the other dbSNP build 150 tracks which have a maximum weight of 1.\ That enables these multiply-mapped SNPs to appear in the display, while\ by default they will not appear in the All SNPs(150) track because of its\ maximum weight filter.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp150*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp151OrthoPt5Pa2Rm8\ codingAnnotations snp151CodingDbSnp,\ defaultGeneTracks knownGene\ defaultMaxWeight 3\ group varRep\ hapmapPhase III\ html ../snp150Mult\ longLabel Simple Nucleotide Polymorphisms (dbSNP 151) That Map to Multiple Genomic Loci\ macaqueDb rheMac8\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.912\ shortLabel Mult. SNPs(151)\ snpExceptionDesc snp151ExceptionDesc\ snpSeq snp151Seq\ track snp151Mult\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp150Mult Mult. SNPs(150) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 150) That Map to Multiple Genomic Loci 0 0.913 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 150, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs that have been mapped to multiple locations in the reference\ genome assembly are included in this subset. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\\ Since build 149, dbSNP has been filtering out almost all such "SNPs" so\ there are very few items in this track.\
\\ The default maximum weight for this track is 3,\ unlike the other dbSNP build 150 tracks which have a maximum weight of 1.\ That enables these multiply-mapped SNPs to appear in the display, while\ by default they will not appear in the All SNPs(150) track because of its\ maximum weight filter.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp150*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp150OrthoPt5Pa2Rm8\ codingAnnotations snp150CodingDbSnp,\ defaultGeneTracks knownGene\ defaultMaxWeight 3\ group varRep\ hapmapPhase III\ html ../snp150Mult\ longLabel Simple Nucleotide Polymorphisms (dbSNP 150) That Map to Multiple Genomic Loci\ macaqueDb rheMac8\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.913\ shortLabel Mult. SNPs(150)\ snpExceptionDesc snp150ExceptionDesc\ snpSeq snp150Seq\ track snp150Mult\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp150 All SNPs(150) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 150) 0 0.914 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 150, available from\ ftp.ncbi.nlm.nih.gov/snp.\
\\ Three tracks contain subsets of the items in this track:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic\ locations will be omitted from display. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp150*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp150OrthoPt5Pa2Rm8\ codingAnnotations snp150CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp150\ longLabel Simple Nucleotide Polymorphisms (dbSNP 150)\ macaqueDb rheMac8\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.914\ shortLabel All SNPs(150)\ tableBrowser noGenome\ track snp150\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp150Common Common SNPs(150) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 150) Found in >= 1% of Samples 0 0.915 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 150, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs that have a minor allele frequency (MAF) of at least 1% and\ are mapped to a single location in the reference genome assembly are\ included in this subset. Frequency data are not available for all SNPs,\ so this subset is incomplete.\ Allele counts from all submissions that include frequency data are combined\ when determining MAF, so for example the allele counts from\ the 1000 Genomes Project and an independent submitter may be combined for the\ same variant.\
\\ dbSNP provides\ download files\ in the\ Variant Call Format (VCF)\ that include a "COMMON" flag in the INFO column. That is determined by a different method,\ and is generally a superset of the UCSC Common set.\ dbSNP uses frequency data from the\ 1000 Genomes Project\ only, and considers a variant COMMON if it has a MAF of at least 0.01 in any of the five\ super-populations:\
\ The selection of SNPs with a minor allele frequency of 1% or greater\ is an attempt to identify variants that appear to be reasonably common\ in the general population. Taken as a set, common variants should be\ less likely to be associated with severe genetic diseases due to the\ effects of natural selection,\ following the view that deleterious variants are not likely to become\ common in the population.\ However, the significance of any particular variant should be interpreted\ only by a trained medical geneticist using all available information.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp150*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp150OrthoPt5Pa2Rm8\ codingAnnotations snp150CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp150Common\ longLabel Simple Nucleotide Polymorphisms (dbSNP 150) Found in >= 1% of Samples\ macaqueDb rheMac8\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.915\ shortLabel Common SNPs(150)\ snpExceptionDesc snp150ExceptionDesc\ snpSeq snp150Seq\ track snp150Common\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp150Flagged Flagged SNPs(150) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 150) Flagged by dbSNP as Clinically Assoc 0 0.916 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 150, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs flagged as clinically associated by dbSNP,\ mapped to a single location in the reference genome assembly, and\ not known to have a minor allele frequency of at\ least 1%, are included in this subset.\ Frequency data are not available for all SNPs, so this subset probably\ includes some SNPs whose true minor allele frequency is 1% or greater.\
\\ The significance of any particular variant in this track should be\ interpreted only by a trained medical geneticist using all available\ information. For example, some variants are included in this track\ because of their inclusion in a Locus-Specific Database (LSDB) or\ mention in OMIM, but are not thought to be disease-causing, so\ inclusion of a variant in this track is not necessarily an indicator\ of risk. Again, all available information must be carefully considered\ by a qualified professional.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/database/data/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/database/data/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b150_GRCh38p7/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp150*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro5\ chimpOrangMacOrthoTable snp150OrthoPt5Pa2Rm8\ codingAnnotations snp150CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp150Flagged\ longLabel Simple Nucleotide Polymorphisms (dbSNP 150) Flagged by dbSNP as Clinically Assoc\ macaqueDb rheMac8\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.916\ shortLabel Flagged SNPs(150)\ snpExceptionDesc snp150ExceptionDesc\ snpSeq snp150Seq\ track snp150Flagged\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp147Mult Mult. SNPs(147) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 147) That Map to Multiple Genomic Loci 0 0.921 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 147, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs that have been mapped to multiple locations in the reference\ genome assembly are included in this subset. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\\ The default maximum weight for this track is 3,\ unlike the other dbSNP build 147 tracks which have a maximum weight of 1.\ That enables these multiply-mapped SNPs to appear in the display, while\ by default they will not appear in the All SNPs(147) track because of its\ maximum weight filter.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp147*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp147OrthoPt4Pa2Rm3\ codingAnnotations snp147CodingDbSnp,\ defaultGeneTracks knownGene\ defaultMaxWeight 3\ group varRep\ hapmapPhase III\ html ../snp147Mult\ longLabel Simple Nucleotide Polymorphisms (dbSNP 147) That Map to Multiple Genomic Loci\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.921\ shortLabel Mult. SNPs(147)\ snpExceptionDesc snp147ExceptionDesc\ snpSeq snp147Seq\ track snp147Mult\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp147Flagged Flagged SNPs(147) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 147) Flagged by dbSNP as Clinically Assoc 0 0.922 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 147, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs flagged as clinically associated by dbSNP,\ mapped to a single location in the reference genome assembly, and\ not known to have a minor allele frequency of at\ least 1%, are included in this subset.\ Frequency data are not available for all SNPs, so this subset probably\ includes some SNPs whose true minor allele frequency is 1% or greater.\
\\ The significance of any particular variant in this track should be\ interpreted only by a trained medical geneticist using all available\ information. For example, some variants are included in this track\ because of their inclusion in a Locus-Specific Database (LSDB) or\ mention in OMIM, but are not thought to be disease-causing, so\ inclusion of a variant in this track is not necessarily an indicator\ of risk. Again, all available information must be carefully considered\ by a qualified professional.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp147*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp147OrthoPt4Pa2Rm3\ codingAnnotations snp147CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp147Flagged\ longLabel Simple Nucleotide Polymorphisms (dbSNP 147) Flagged by dbSNP as Clinically Assoc\ macaqueDb rheMac3\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.922\ shortLabel Flagged SNPs(147)\ snpExceptionDesc snp147ExceptionDesc\ snpSeq snp147Seq\ track snp147Flagged\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp147Common Common SNPs(147) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 147) Found in >= 1% of Samples 0 0.923 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 147, available from\ ftp.ncbi.nlm.nih.gov/snp.\ Only SNPs that have a minor allele frequency of at least 1% and\ are mapped to a single location in the reference genome assembly are\ included in this subset. Frequency data are not available for all SNPs,\ so this subset is incomplete.\
\\ The selection of SNPs with a minor allele frequency of 1% or greater\ is an attempt to identify variants that appear to be reasonably common\ in the general population. Taken as a set, common variants should be\ less likely to be associated with severe genetic diseases due to the\ effects of natural selection,\ following the view that deleterious variants are not likely to become\ common in the population.\ However, the significance of any particular variant should be interpreted\ only by a trained medical geneticist using all available information.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp147*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp147OrthoPt4Pa2Rm3\ codingAnnotations snp147CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp147Common\ longLabel Simple Nucleotide Polymorphisms (dbSNP 147) Found in >= 1% of Samples\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.923\ shortLabel Common SNPs(147)\ snpExceptionDesc snp147ExceptionDesc\ snpSeq snp147Seq\ track snp147Common\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp147 All SNPs(147) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 147) 0 0.924 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 147, available from\ ftp.ncbi.nlm.nih.gov/snp.\
\\ Three tracks contain subsets of the items in this track:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic\ locations will be omitted from display. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period >= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b147_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp147*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp147OrthoPt4Pa2Rm3\ codingAnnotations snp147CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp147\ longLabel Simple Nucleotide Polymorphisms (dbSNP 147)\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.924\ shortLabel All SNPs(147)\ tableBrowser noGenome\ track snp147\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp146Mult Mult. SNPs(146) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 146) That Map to Multiple Genomic Loci 0 0.925 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 146, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have been mapped to multiple locations in the reference\ genome assembly are included in this subset. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\\ The default maximum weight for this track is 3,\ unlike the other dbSNP build 146 tracks which have a maximum weight of 1.\ That enables these multiply-mapped SNPs to appear in the display, while\ by default they will not appear in the All SNPs(146) track because of its\ maximum weight filter.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp146*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp146OrthoPt4Pa2Rm3\ codingAnnotations snp146CodingDbSnp,\ defaultGeneTracks knownGene\ defaultMaxWeight 3\ group varRep\ hapmapPhase III\ html ../snp146Mult\ longLabel Simple Nucleotide Polymorphisms (dbSNP 146) That Map to Multiple Genomic Loci\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.925\ shortLabel Mult. SNPs(146)\ snpExceptionDesc snp146ExceptionDesc\ snpSeq snp146Seq\ track snp146Mult\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp146Flagged Flagged SNPs(146) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 146) Flagged by dbSNP as Clinically Assoc 0 0.926 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 146, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs flagged as clinically associated by dbSNP,\ mapped to a single location in the reference genome assembly, and\ not known to have a minor allele frequency of at\ least 1%, are included in this subset.\ Frequency data are not available for all SNPs, so this subset probably\ includes some SNPs whose true minor allele frequency is 1% or greater.\
\\ The significance of any particular variant in this track should be\ interpreted only by a trained medical geneticist using all available\ information. For example, some variants are included in this track\ because of their inclusion in a Locus-Specific Database (LSDB) or\ mention in OMIM, but are not thought to be disease-causing, so\ inclusion of a variant in this track is not necessarily an indicator\ of risk. Again, all available information must be carefully considered\ by a qualified professional.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp146*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp146OrthoPt4Pa2Rm3\ codingAnnotations snp146CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp146Flagged\ longLabel Simple Nucleotide Polymorphisms (dbSNP 146) Flagged by dbSNP as Clinically Assoc\ macaqueDb rheMac3\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.926\ shortLabel Flagged SNPs(146)\ snpExceptionDesc snp146ExceptionDesc\ snpSeq snp146Seq\ track snp146Flagged\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp146Common Common SNPs(146) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 146) Found in >= 1% of Samples 0 0.927 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 146, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have a minor allele frequency of at least 1% and\ are mapped to a single location in the reference genome assembly are\ included in this subset. Frequency data are not available for all SNPs,\ so this subset is incomplete.\
\\ The selection of SNPs with a minor allele frequency of 1% or greater\ is an attempt to identify variants that appear to be reasonably common\ in the general population. Taken as a set, common variants should be\ less likely to be associated with severe genetic diseases due to the\ effects of natural selection,\ following the view that deleterious variants are not likely to become\ common in the population.\ However, the significance of any particular variant should be interpreted\ only by a trained medical geneticist using all available information.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp146*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp146OrthoPt4Pa2Rm3\ codingAnnotations snp146CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp146Common\ longLabel Simple Nucleotide Polymorphisms (dbSNP 146) Found in >= 1% of Samples\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.927\ shortLabel Common SNPs(146)\ snpExceptionDesc snp146ExceptionDesc\ snpSeq snp146Seq\ track snp146Common\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp146 All SNPs(146) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 146) 0 0.928 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 146, available from\ ftp.ncbi.nih.gov/snp.\
\\ Three tracks contain subsets of the items in this track:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic\ locations will be omitted from display. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b146_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp146*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp146OrthoPt4Pa2Rm3\ codingAnnotations snp146CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp146\ longLabel Simple Nucleotide Polymorphisms (dbSNP 146)\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.928\ shortLabel All SNPs(146)\ tableBrowser noGenome\ track snp146\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp144Mult Mult. SNPs(144) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 144) That Map to Multiple Genomic Loci 0 0.929 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 144, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have been mapped to multiple locations in the reference\ genome assembly are included in this subset. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\\ The default maximum weight for this track is 3,\ unlike the other dbSNP build 144 tracks which have a maximum weight of 1.\ That enables these multiply-mapped SNPs to appear in the display, while\ by default they will not appear in the All SNPs(144) track because of its\ maximum weight filter.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp144*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp144OrthoPt4Pa2Rm3\ codingAnnotations snp144CodingDbSnp,\ defaultGeneTracks knownGene\ defaultMaxWeight 3\ group varRep\ hapmapPhase III\ html ../snp144Mult\ longLabel Simple Nucleotide Polymorphisms (dbSNP 144) That Map to Multiple Genomic Loci\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.929\ shortLabel Mult. SNPs(144)\ snpExceptionDesc snp144ExceptionDesc\ snpSeq snp144Seq\ snpSeqFile /gbdb/hg38/snp/snp144.fa\ track snp144Mult\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp144Flagged Flagged SNPs(144) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 144) Flagged by dbSNP as Clinically Assoc 0 0.93 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 144, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs flagged as clinically associated by dbSNP,\ mapped to a single location in the reference genome assembly, and\ not known to have a minor allele frequency of at\ least 1%, are included in this subset.\ Frequency data are not available for all SNPs, so this subset probably\ includes some SNPs whose true minor allele frequency is 1% or greater.\
\\ The significance of any particular variant in this track should be\ interpreted only by a trained medical geneticist using all available\ information. For example, some variants are included in this track\ because of their inclusion in a Locus-Specific Database (LSDB) or\ mention in OMIM, but are not thought to be disease-causing, so\ inclusion of a variant in this track is not necessarily an indicator\ of risk. Again, all available information must be carefully considered\ by a qualified professional.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp144*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp144OrthoPt4Pa2Rm3\ codingAnnotations snp144CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp144Flagged\ longLabel Simple Nucleotide Polymorphisms (dbSNP 144) Flagged by dbSNP as Clinically Assoc\ macaqueDb rheMac3\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.93\ shortLabel Flagged SNPs(144)\ snpExceptionDesc snp144ExceptionDesc\ snpSeq snp144Seq\ snpSeqFile /gbdb/hg38/snp/snp144.fa\ track snp144Flagged\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp144Common Common SNPs(144) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 144) Found in >= 1% of Samples 0 0.931 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the\ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 144, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have a minor allele frequency of at least 1% and\ are mapped to a single location in the reference genome assembly are\ included in this subset. Frequency data are not available for all SNPs,\ so this subset is incomplete.\
\\ The selection of SNPs with a minor allele frequency of 1% or greater\ is an attempt to identify variants that appear to be reasonably common\ in the general population. Taken as a set, common variants should be\ less likely to be associated with severe genetic diseases due to the\ effects of natural selection,\ following the view that deleterious variants are not likely to become\ common in the population.\ However, the significance of any particular variant should be interpreted\ only by a trained medical geneticist using all available information.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp144*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp144OrthoPt4Pa2Rm3\ codingAnnotations snp144CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp144Common\ longLabel Simple Nucleotide Polymorphisms (dbSNP 144) Found in >= 1% of Samples\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.931\ shortLabel Common SNPs(144)\ snpExceptionDesc snp144ExceptionDesc\ snpSeq snp144Seq\ snpSeqFile /gbdb/hg38/snp/snp144.fa\ track snp144Common\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp144 All SNPs(144) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 144) 0 0.932 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 144, available from\ ftp.ncbi.nih.gov/snp.\
\\ Three tracks contain subsets of the items in this track:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic\ locations will be omitted from display. When a SNP's flanking sequences\ map to multiple locations in the reference genome, it calls into question\ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width\ of a single base, and multiple nucleotide variants are represented by a\ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the\ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to\ particular gene sets. Choose the gene sets from the list on the SNP\ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to.\ When one or more gene tracks are selected, the SNP details page\ lists all genes that the SNP hits (or is close to), with the same keywords\ used in the function category. The function usually\ agrees with NCBI's function, except when NCBI's functional annotation is\ relative to an XM_* predicted RefSeq (not included in the UCSC Genome\ Browser's RefSeq Genes track) and/or UCSC's functional annotation is\ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking\ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences\ to the neighboring genomic sequence for display on SNP details pages.\ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking\ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files\ and headers of fasta files downloaded from NCBI.\ The database dump files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/database/organism_data/\ for hg38.\ The fasta files were downloaded from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606_b144_GRCh38p2/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp144*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies.\ We use our liftOver utility to identify the orthologous alleles.\ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K.\ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp144OrthoPt4Pa2Rm3\ codingAnnotations snp144CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp144\ longLabel Simple Nucleotide Polymorphisms (dbSNP 144)\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.932\ shortLabel All SNPs(144)\ snpSeqFile /gbdb/hg38/snp/snp144.fa\ tableBrowser noGenome\ track snp144\ trackHandler snp125\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp142Mult Mult. SNPs(142) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 142) That Map to Multiple Genomic Loci 0 0.933 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the \ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 142, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have been mapped to multiple locations in the reference\ genome assembly are included in this subset. When a SNP's flanking sequences \ map to multiple locations in the reference genome, it calls into question \ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\\ The default maximum weight for this track is 3,\ unlike the other dbSNP build 142 tracks which have a maximum weight of 1. \ That enables these multiply-mapped SNPs to appear in the display, while \ by default they will not appear in the All SNPs(142) track because of its \ maximum weight filter.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/database/organism_data/\ for hg38.\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp142*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp142OrthoPt4Pa2Rm3\ codingAnnotations snp142CodingDbSnp,\ defaultGeneTracks knownGene\ defaultMaxWeight 3\ group varRep\ hapmapPhase III\ html ../snp142Mult\ longLabel Simple Nucleotide Polymorphisms (dbSNP 142) That Map to Multiple Genomic Loci\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.933\ shortLabel Mult. SNPs(142)\ snpExceptionDesc snp142ExceptionDesc\ snpSeq snp142Seq\ snpSeqFile /gbdb/hg38/snp/snp142.fa\ track snp142Mult\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp142Flagged Flagged SNPs(142) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 142) Flagged by dbSNP as Clinically Assoc 0 0.934 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the \ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 142, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs flagged as clinically associated by dbSNP, \ mapped to a single location in the reference genome assembly, and \ not known to have a minor allele frequency of at \ least 1%, are included in this subset.\ Frequency data are not available for all SNPs, so this subset probably\ includes some SNPs whose true minor allele frequency is 1% or greater.\
\\ The significance of any particular variant in this track should be\ interpreted only by a trained medical geneticist using all available\ information. For example, some variants are included in this track\ because of their inclusion in a Locus-Specific Database (LSDB) or\ mention in OMIM, but are not thought to be disease-causing, so\ inclusion of a variant in this track is not necessarily an indicator\ of risk. Again, all available information must be carefully considered\ by a qualified professional.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/database/organism_data/\ for hg38.\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp142*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp142OrthoPt4Pa2Rm3\ codingAnnotations snp142CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp142Flagged\ longLabel Simple Nucleotide Polymorphisms (dbSNP 142) Flagged by dbSNP as Clinically Assoc\ macaqueDb rheMac3\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.934\ shortLabel Flagged SNPs(142)\ snpExceptionDesc snp142ExceptionDesc\ snpSeq snp142Seq\ snpSeqFile /gbdb/hg38/snp/snp142.fa\ track snp142Flagged\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp142Common Common SNPs(142) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 142) Found in >= 1% of Samples 0 0.935 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the \ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 142, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have a minor allele frequency of at least 1% and\ are mapped to a single location in the reference genome assembly are\ included in this subset. Frequency data are not available for all SNPs,\ so this subset is incomplete.\
\\ The selection of SNPs with a minor allele frequency of 1% or greater\ is an attempt to identify variants that appear to be reasonably common\ in the general population. Taken as a set, common variants should be\ less likely to be associated with severe genetic diseases due to the\ effects of natural selection,\ following the view that deleterious variants are not likely to become\ common in the population.\ However, the significance of any particular variant should be interpreted\ only by a trained medical geneticist using all available information.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/database/organism_data/\ for hg38.\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp142*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp142OrthoPt4Pa2Rm3\ codingAnnotations snp142CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp142Common\ longLabel Simple Nucleotide Polymorphisms (dbSNP 142) Found in >= 1% of Samples\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.935\ shortLabel Common SNPs(142)\ snpExceptionDesc snp142ExceptionDesc\ snpSeq snp142Seq\ snpSeqFile /gbdb/hg38/snp/snp142.fa\ track snp142Common\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp142 All SNPs(142) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 142) 0 0.936 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 142, available from\ ftp.ncbi.nih.gov/snp.\
\\ Three tracks contain subsets of the items in this track:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic \ locations will be omitted from display. When a SNP's flanking sequences \ map to multiple locations in the reference genome, it calls into question \ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/database/organism_data/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/database/organism_data/\ for hg38.\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh37p13/rs_fasta/\ for hg19 and from\ ftp://ftp.ncbi.nih.gov/snp/organisms/human_9606_b142_GRCh38/rs_fasta/\ for hg38.\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp142*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp142OrthoPt4Pa2Rm3\ codingAnnotations snp142CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp142\ longLabel Simple Nucleotide Polymorphisms (dbSNP 142)\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.936\ shortLabel All SNPs(142)\ snpSeqFile /gbdb/hg38/snp/snp142.fa\ tableBrowser noGenome\ track snp142\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp141Mult Mult. SNPs(141) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 141) That Map to Multiple Genomic Loci 0 0.937 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the \ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 141, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have been mapped to multiple locations in the reference\ genome assembly are included in this subset. When a SNP's flanking sequences \ map to multiple locations in the reference genome, it calls into question \ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\\ The default maximum weight for this track is 3,\ unlike the other dbSNP build 141 tracks which have a maximum weight of 1. \ That enables these multiply-mapped SNPs to appear in the display, while \ by default they will not appear in the All SNPs(141) track because of its \ maximum weight filter.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/database/\ (for human, organism_tax_id = human_9606;\ for mouse, organism_tax_id = mouse_10090).\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/rs_fasta/\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp141*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp141OrthoPt4Pa2Rm3\ codingAnnotations snp141CodingDbSnp,\ defaultGeneTracks knownGene\ defaultMaxWeight 3\ group varRep\ hapmapPhase III\ html ../snp141Mult\ longLabel Simple Nucleotide Polymorphisms (dbSNP 141) That Map to Multiple Genomic Loci\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.937\ shortLabel Mult. SNPs(141)\ snpExceptionDesc snp141ExceptionDesc\ snpSeq snp141Seq\ snpSeqFile /gbdb/hg38/snp/snp141.fa\ track snp141Mult\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp141Flagged Flagged SNPs(141) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 141) Flagged by dbSNP as Clinically Assoc 0 0.938 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the \ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 141, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs flagged as clinically associated by dbSNP, \ mapped to a single location in the reference genome assembly, and \ not known to have a minor allele frequency of at \ least 1%, are included in this subset.\ Frequency data are not available for all SNPs, so this subset probably\ includes some SNPs whose true minor allele frequency is 1% or greater.\
\\ The significance of any particular variant in this track should be\ interpreted only by a trained medical geneticist using all available\ information. For example, some variants are included in this track\ because of their inclusion in a Locus-Specific Database (LSDB) or\ mention in OMIM, but are not thought to be disease-causing, so\ inclusion of a variant in this track is not necessarily an indicator\ of risk. Again, all available information must be carefully considered\ by a qualified professional.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/database/\ (for human, organism_tax_id = human_9606;\ for mouse, organism_tax_id = mouse_10090).\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/rs_fasta/\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp141*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp141OrthoPt4Pa2Rm3\ codingAnnotations snp141CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp141Flagged\ longLabel Simple Nucleotide Polymorphisms (dbSNP 141) Flagged by dbSNP as Clinically Assoc\ macaqueDb rheMac3\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.938\ shortLabel Flagged SNPs(141)\ snpExceptionDesc snp141ExceptionDesc\ snpSeq snp141Seq\ snpSeqFile /gbdb/hg38/snp/snp141.fa\ track snp141Flagged\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp141Common Common SNPs(141) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 141) Found in >= 1% of Samples 0 0.939 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about a subset of the \ single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 141, available from\ ftp.ncbi.nih.gov/snp.\ Only SNPs that have a minor allele frequency of at least 1% and\ are mapped to a single location in the reference genome assembly are\ included in this subset. Frequency data are not available for all SNPs,\ so this subset is incomplete.\
\\ The selection of SNPs with a minor allele frequency of 1% or greater\ is an attempt to identify variants that appear to be reasonably common\ in the general population. Taken as a set, common variants should be\ less likely to be associated with severe genetic diseases due to the\ effects of natural selection,\ following the view that deleterious variants are not likely to become\ common in the population.\ However, the significance of any particular variant should be interpreted\ only by a trained medical geneticist using all available information.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/database/\ (for human, organism_tax_id = human_9606;\ for mouse, organism_tax_id = mouse_10090).\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/rs_fasta/\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp141*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp141OrthoPt4Pa2Rm3\ codingAnnotations snp141CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp141Common\ longLabel Simple Nucleotide Polymorphisms (dbSNP 141) Found in >= 1% of Samples\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.939\ shortLabel Common SNPs(141)\ snpExceptionDesc snp141ExceptionDesc\ snpSeq snp141Seq\ snpSeqFile /gbdb/hg38/snp/snp141.fa\ track snp141Common\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ snp141 All SNPs(141) bed 6 + Simple Nucleotide Polymorphisms (dbSNP 141) 0 0.94 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track contains information about single nucleotide polymorphisms\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP\ build 141, available from\ ftp.ncbi.nih.gov/snp.\
\\ Three tracks contain subsets of the items in this track:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic \ locations will be omitted from display. When a SNP's flanking sequences \ map to multiple locations in the reference genome, it calls into question \ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \ The remainder of this page is identical on the following tracks:\\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/database/\ (for human, organism_tax_id = human_9606;\ for mouse, organism_tax_id = mouse_10090).\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/rs_fasta/\
\\ The raw data can be explored interactively with the Table Browser,\ Data Integrator, or Variant Annotation Integrator.\ For automated analysis, the genome annotation can be downloaded from the downloads server for hg38 and\ hg19 (snp141*.txt.gz) or the public MySQL server.\ Please refer to our mailing list archives\ for questions and example queries, or our Data Access FAQ for more information.\
\ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ \ \ varRep 1 chimpDb panTro4\ chimpOrangMacOrthoTable snp141OrthoPt4Pa2Rm3\ codingAnnotations snp141CodingDbSnp,\ defaultGeneTracks knownGene\ group varRep\ hapmapPhase III\ html ../snp141\ longLabel Simple Nucleotide Polymorphisms (dbSNP 141)\ macaqueDb rheMac3\ maxWindowToDraw 10000000\ orangDb ponAbe2\ parent dbSnpArchive\ priority 0.94\ shortLabel All SNPs(141)\ snpSeqFile /gbdb/hg38/snp/snp141.fa\ tableBrowser noGenome\ track snp141\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ chainCriGriChoV2 Chinese hamster Chain chain criGriChoV2 Chinese hamster (Jun. 2017 (CHOK1S_HZDv1/criGriChoV2)) Chained Alignments 3 1 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Chinese hamster (Jun. 2017 (CHOK1S_HZDv1/criGriChoV2)) Chained Alignments\ otherDb criGriChoV2\ parent placentalChainNetViewchain off\ shortLabel Chinese hamster Chain\ subGroups view=chain species=s004b clade=c00\ track chainCriGriChoV2\ type chain criGriChoV2\ chainMonDom5 Opossum Chain chain monDom5 Opossum (Oct. 2006 (Broad/monDom5)) Chained Alignments 3 1 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Opossum (Oct. 2006 (Broad/monDom5)) Chained Alignments\ otherDb monDom5\ parent vertebrateChainNetViewchain off\ shortLabel Opossum Chain\ subGroups view=chain species=s003 clade=c00\ track chainMonDom5\ type chain monDom5\ chainPanTro6 Chimp Chain chain panTro6 Chimp (Jan. 2018 (Clint_PTRv2/panTro6)) Chained Alignments 3 1 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Chimp (Jan. 2018 (Clint_PTRv2/panTro6)) Chained Alignments\ otherDb panTro6\ parent primateChainNetViewchain off\ shortLabel Chimp Chain\ subGroups view=chain species=s0025 clade=c00\ track chainPanTro6\ type chain panTro6\ tishkoff180 12 Afr Pops 180 WGS vcfTabix SNV Frequencies: 180 WGS from 12 Indigenous African Populations (Fan 2023) 0 1 0 0 0 127 127 127 0 0 0\ This track shows allele frequencies from high-coverage whole-genome sequencing of\ 180 individuals (15 per population) from 12 indigenous African populations that cover\ all four major African language phyla (Khoesan, Niger-Congo, Nilo-Saharan, Afroasiatic).\ The cohort, generated by the Tishkoff lab and collaborators (Fan et al., Cell 2023),\ spans the Amhara, Dizi, Chabu and Mursi from Ethiopia; the Hadza and Sandawe from Tanzania;\ the Central African rainforest hunter-gatherers (Baka and Bagyeli, merged), Fulani and Tikari\ from Cameroon; and the Herero, Ju|'hoansi and !Xoo (the latter two collectively the "San")\ from Botswana. The dataset was generated to capture demographic history and signatures of\ local adaptation in African populations that are poorly represented in other reference panels.\
\ \\ Only aggregate allele frequencies (AC, AF, AN summed over all 180 individuals) are\ shown for each variant; per-population frequencies are not provided in the released\ sites VCF. The original variant calls were on the GRCh37/hs37d5 reference and were\ lifted to hg38 at UCSC.\
\ \\ Variants display as standard VCF allele frequency tracks. On mouseover and click,\ the allele count (AC), total allele number (AN) and allele frequency (AF) are shown.\ When zoomed in, alleles are colored by base. Multi-allelic records were split into\ biallelic rows during normalization upstream.\
\ \\ Whole genome sequencing of 180 individuals (15 unrelated samples per population)\ was performed at >30× average coverage on the Illumina HiSeq X Ten platform\ using PCR-free library preparation with paired-end 150 bp reads and a 350 bp\ insert size. Adapters were trimmed with trimadap, optical duplicates were marked with\ SAMBLASTER (v0.1.22), and reads were aligned to the hs37d5 decoy version of GRCh37\ with BWA-MEM (v0.7.10). Reads with mapping quality < 20 were filtered. Per-sample\ short variants were called with GATK HaplotypeCaller (nightly-2016-09-26-gfade77f) in\ gVCF mode with a custom genotype prior (0.4995, 0.001, 0.4995) to reduce reference\ bias, as recommended by SGDP. Joint genotyping was performed with GATK\ GenotypeGVCFs. Variants were filtered with GATK VQSR using 1000 Genomes Phase 3,\ Illumina Omni 5M and HapMap as SNP truth sets and Mills indels as the indel truth set.\ Variants overlapping potential duplications detected by Delly (v0.7.6) and low-complexity\ regions were excluded. After QC the cohort yielded 32.4 M SNPs and 2.8 M small\ indels. The publicly released SNP-only sites VCF used here contains 33.6 M\ biallelic SNPs with aggregate AC/AF/AN summaries. See Fan et al. (2023) for full\ methods.\
\ \\ The hg19 SNPs sites VCF was provided directly by Matthew Hansen at the Tishkoff lab\ (University of Pennsylvania) via a Box link\ (180wgs.SNPs.sites.AF.vcf.gz). Bare chromosome names (1-22) were converted\ to UCSC-style names with bcftools annotate --rename-chrs, the VCF was lifted\ from hg19 to hg38 with CrossMap.py vcf using the UCSC\ hg19ToHg38.over.chain.gz chain, then sorted, bgzip-compressed and tabix-indexed\ with bcftools sort and tabix. Step-by-step processing instructions are in\ the\ makeDoc file; the supporting scripts live under\ kent/src/hg/makeDb/scripts/varFreqs.\
\ \\ The original (hg19) variant calls and supplementary data accompany the publication;\ see the "Data and code availability" section of Fan et al. (2023). The dataset is\ not available for redistribution from our website, so the Table Browser, Data\ Integrator and download server are disabled for this track. The hg19 sites VCF can\ be requested from the Tishkoff lab at the University of Pennsylvania.\
\ \\ Thanks to Matthew Hansen and Sarah Tishkoff (University of Pennsylvania) for sharing\ the sites-only allele-frequency VCF, and to all participating individuals and field\ collaborators in Ethiopia, Tanzania, Cameroon and Botswana whose contributions made\ this dataset possible.\
\ \\ Fan S, Spence JP, Feng Y, Hansen MEB, Terhorst J, Beltrame MH, Ranciaro A, Hirbo J, Beggs W, Thomas\ N et al.\ \ Whole-genome sequencing reveals a complex African population demographic history and signatures of\ local adaptation.\ Cell. 2023 Mar 2;186(5):923-939.e14.\ PMID: 36868214; PMC: PMC10568978\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_tishkoff/tishkoff180.vcf.gz\ dataVersion Cell 2023 (hg19 lift)\ longLabel SNV Frequencies: 180 WGS from 12 Indigenous African Populations (Fan 2023)\ parent varFreqs on\ priority 1\ shortLabel 12 Afr Pops 180 WGS\ tableBrowser off\ track tishkoff180\ type vcfTabix\ visibility hide\ tgpNA12878_1463_CEU 1463 CEU Trio vcfPhasedTrio 1000 Genomes Utah CEPH Trio 2 1 0 0 0 127 127 127 0 0 23 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX, varRep 0 longLabel 1000 Genomes Utah CEPH Trio\ parent tgpTrios\ shortLabel 1463 CEU Trio\ track tgpNA12878_1463_CEU\ type vcfPhasedTrio\ vcfChildSample NA12878|child\ vcfParentSamples NA12892|mother,NA12891|father\ visibility full\ phyloP447wayBW 447 phyloP REV bigWig -20 11.936 447 mammals / 233 primates Basewise Conservation by PhyloP phyloFit REV model 2 1 60 60 140 140 60 60 0 0 0 compGeno 0 altColor 140,60,60\ autoScale off\ bigDataUrl https://hgdownload.soe.ucsc.edu/goldenPath/hg38/phyloP447way/hg38.phyloP447way.bw\ color 60,60,140\ configurable on\ logo on\ longLabel 447 mammals / 233 primates Basewise Conservation by PhyloP phyloFit REV model\ maxHeightPixels 100:50:11\ noInherit on\ parent cons447wayViewphyloP\ priority 1\ shortLabel 447 phyloP REV\ spanList 1\ subGroups view=phyloP\ track phyloP447wayBW\ type bigWig -20 11.936\ viewLimits -4.5:7.5\ windowingFunction mean\ encTfChipPkENCFF851UTY A549 ATF3 narrowPeak Transcription Factor ChIP-seq Peaks of ATF3 in A549 from ENCODE 3 (ENCFF851UTY) 0 1 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of ATF3 in A549 from ENCODE 3 (ENCFF851UTY)\ parent encTfChipPk off\ shortLabel A549 ATF3\ subGroups cellType=A549 factor=ATF3\ track encTfChipPkENCFF851UTY\ cloneEndABC10 ABC10 bed 12 Agencourt fosmid library 10 3 1 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 10\ parent cloneEndSuper off\ priority 1\ shortLabel ABC10\ subGroups source=agencourt\ track cloneEndABC10\ type bed 12\ visibility pack\ gtexCovAdiposeSubcutaneous Adip Subcut bigWig Adipose Subcutaneous 0 1 255 165 79 255 210 167 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-NFK9-0326-SM-3MJGV.Adipose_Subcutaneous.RNAseq.bw\ color 255,165,79\ longLabel Adipose Subcutaneous\ parent gtexCov\ shortLabel Adip Subcut\ track gtexCovAdiposeSubcutaneous\ lincRNAsCTAdipose Adipose bed 5 + lincRNAs from adipose 1 1 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from adipose\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Adipose\ subGroups view=lincRNAsRefseqExp tissueType=adipose\ track lincRNAsCTAdipose\ wgEncodeReg4TxnAdiposePlus Adipose + bigWig Avg. + strand total RNA-seq level of 9 adipose experiments (tissues and primary cells only) 0 1 255 119 39 255 187 147 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adiposePlus.bw\ color 255,119,39\ longLabel Avg. + strand total RNA-seq level of 9 adipose experiments (tissues and primary cells only)\ parent wgEncodeReg4Txn off\ priority 1\ shortLabel Adipose +\ track wgEncodeReg4TxnAdiposePlus\ type bigWig\ genetiSureCytoCghSnp Agilent GenetiSure Cyto CGH+SNP bigBed 4 Agilent GenetiSure Cyto CGH+SNP 4x180K 085591 20200302 3 1 0 0 0 127 127 127 0 0 0 varRep 1 bigDataUrl /gbdb/hg38/snpCnvArrays/agilent/hg38.GenetiSure_Cyto_CGH+SNP_Microarray_4x180K_085591_D_BED_20200302.bb\ longLabel Agilent GenetiSure Cyto CGH+SNP 4x180K 085591 20200302\ parent genotypeArrays on\ priority 1\ shortLabel Agilent GenetiSure Cyto CGH+SNP\ track genetiSureCytoCghSnp\ type bigBed 4\ visibility pack\ allCancer All Cancers bigLolly 12 + All TCGA Pan-Cancer mutations: 33 TCGA Cancer Projects Summary (Pan-Can 33) 0 1 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/gdcCancer.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel All TCGA Pan-Cancer mutations: 33 TCGA Cancer Projects Summary (Pan-Can 33)\ parent gdcCancer on\ priority 1\ shortLabel All Cancers\ track allCancer\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/$$\ alldifficultregions All difficult regions bigBed 3 Genome In a Bottle: all difficult regions 1 1 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/GIAB/alldifficultregions.bb\ longLabel Genome In a Bottle: all difficult regions\ parent problematicGIAB on\ shortLabel All difficult regions\ track alldifficultregions\ type bigBed 3\ visibility dense\ hmaSummaryUnmethylated All Unmeth Regions bigBed 9 . Methylation Atlas: All unmethylated regions 1 1 0 0 0 127 127 127 0 0 0 regulation 1 bigDataUrl /gbdb/hg38/dnaMethylationAtlas/hmaSummaryUnmethylated.bb\ filterLabel.name Cell/Tissue Type\ filterValues.name Adipocytes,Bladder-Ep,Blood-B,Blood-Granul,Blood-Mono+Macro,Blood-NK,Blood-T,Bone-Osteob,Breast-Basal-Ep,Breast-Luminal-Ep,Colon-Ep,Colon-Fibro,Dermal-Fibro,Endothel,Epid-Kerat,Eryth-prog,Fallopian-Ep,Gallbladder,Gastric-Ep,Head-Neck-Ep,Heart-Cardio,Heart-Fibro,Kidney-Ep,Liver-Hep,Lung-Ep-Alveo,Lung-Ep-Bron,Neuron,Oligodend,Ovary-Ep,Pancreas-Acinar,Pancreas-Alpha,Pancreas-Beta,Pancreas-Delta,Pancreas-Duct,Prostate-Ep,Skeletal-Musc,Small-Int-Ep,Smooth-Musc,Thyroid-Ep\ itemRgb on\ longLabel Methylation Atlas: All unmethylated regions\ parent humanMethylationAtlasSummary on\ priority 1\ shortLabel All Unmeth Regions\ track hmaSummaryUnmethylated\ type bigBed 9 .\ visibility dense\ AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_ctss_fwd AorticSmsToFgf2_00hr00minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_forward 0 1 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr00min%2c%20biol_rep1%20%28LK1%29.CNhs13339.12642-134G5.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12642-134G5 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr00minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_tpm_fwd AorticSmsToFgf2_00hr00minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_forward 1 1 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr00min%2c%20biol_rep1%20%28LK1%29.CNhs13339.12642-134G5.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12642-134G5 sequence_tech=hCAGE\ parent TSS_activity_TPM on\ shortLabel AorticSmsToFgf2_00hr00minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5\ urlLabel FANTOM5 Details:\ ashkenazimTrio Ashkenazim Trio vcfPhasedTrio Genome In a Bottle Ashkenazim Trio 0 1 0 0 0 127 127 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/giab/AshkenazimTrio/merged.vcf.gz\ longLabel Genome In a Bottle Ashkenazim Trio\ maxWindowToDraw 5000000\ parent triosView\ shortLabel Ashkenazim Trio\ subGroups view=trios\ track ashkenazimTrio\ type vcfPhasedTrio\ vcfChildSample HG002|son\ vcfDoFilter off\ vcfDoMaf off\ vcfDoQual off\ vcfParentSamples HG003|father,HG004|mother\ vcfUseAltSampleNames on\ cons100wayViewphyloP Basewise Conservation (phyloP) bed 4 UCSC 100 Vertebrates - 100 vertebrate genomes aligned with MultiZ by the UCSC Browser Group 2 1 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel UCSC 100 Vertebrates - 100 vertebrate genomes aligned with MultiZ by the UCSC Browser Group\ parent cons100way\ shortLabel Basewise Conservation (phyloP)\ track cons100wayViewphyloP\ view phyloP\ viewLimits -20.0:9.869\ viewLimitsMax -20:0.869\ visibility full\ wgEncodeGencodeBasicV20 Basic genePred Basic Gene Annotation Set from GENCODE Version 20 (Ensembl 76) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 20 (Ensembl 76)\ parent wgEncodeGencodeV20ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV20\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV22 Basic genePred Basic Gene Annotation Set from GENCODE Version 22 (Ensembl 79) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 22 (Ensembl 79)\ parent wgEncodeGencodeV22ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV22\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV23 Basic genePred Basic Gene Annotation Set from GENCODE Version 23 (Ensembl 81) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 23 (Ensembl 81)\ parent wgEncodeGencodeV23ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV23\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV24 Basic genePred Basic Gene Annotation Set from GENCODE Version 24 (Ensembl 83) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 24 (Ensembl 83)\ parent wgEncodeGencodeV24ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV24\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV25 Basic genePred Basic Gene Annotation Set from GENCODE Version 25 (Ensembl 85) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 25 (Ensembl 85)\ parent wgEncodeGencodeV25ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV25\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV26 Basic genePred Basic Gene Annotation Set from GENCODE Version 26 (Ensembl 88) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 26 (Ensembl 88)\ parent wgEncodeGencodeV26ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV26\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV27 Basic genePred Basic Gene Annotation Set from GENCODE Version 27 (Ensembl 90) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 27 (Ensembl 90)\ parent wgEncodeGencodeV27ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV27\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV28 Basic genePred Basic Gene Annotation Set from GENCODE Version 28 (Ensembl 92) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 28 (Ensembl 92)\ parent wgEncodeGencodeV28ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV28\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV29 Basic genePred Basic Gene Annotation Set from GENCODE Version 29 (Ensembl 94) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 29 (Ensembl 94)\ parent wgEncodeGencodeV29ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV29\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV30 Basic genePred Basic Gene Annotation Set from GENCODE Version 30 (Ensembl 96) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 30 (Ensembl 96)\ parent wgEncodeGencodeV30ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV30\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV31 Basic genePred Basic Gene Annotation Set from GENCODE Version 31 (Ensembl 97) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 31 (Ensembl 97)\ parent wgEncodeGencodeV31ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV31\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV32 Basic genePred Basic Gene Annotation Set from GENCODE Version 32 (Ensembl 98) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 32 (Ensembl 98)\ parent wgEncodeGencodeV32ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV32\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV33 Basic genePred Basic Gene Annotation Set from GENCODE Version 33 (Ensembl 99) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 33 (Ensembl 99)\ parent wgEncodeGencodeV33ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV33\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV34 Basic genePred Basic Gene Annotation Set from GENCODE Version 34 (Ensembl 100) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 34 (Ensembl 100)\ parent wgEncodeGencodeV34ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV34\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV35 Basic genePred Basic Gene Annotation Set from GENCODE Version 35 (Ensembl 101) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 35 (Ensembl 101)\ parent wgEncodeGencodeV35ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV35\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV36 Basic genePred Basic Gene Annotation Set from GENCODE Version 36 (Ensembl 102) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 36 (Ensembl 102)\ parent wgEncodeGencodeV36ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV36\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV37 Basic genePred Basic Gene Annotation Set from GENCODE Version 37 (Ensembl 103) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 37 (Ensembl 103)\ parent wgEncodeGencodeV37ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV37\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV38 Basic genePred Basic Gene Annotation Set from GENCODE Version 38 (Ensembl 104) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 38 (Ensembl 104)\ parent wgEncodeGencodeV38ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV38\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV39 Basic genePred Basic Gene Annotation Set from GENCODE Version 39 (Ensembl 105) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 39 (Ensembl 105)\ parent wgEncodeGencodeV39ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV39\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV40 Basic genePred Basic Gene Annotation Set from GENCODE Version 40 (Ensembl 106) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 40 (Ensembl 106)\ parent wgEncodeGencodeV40ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV40\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV41 Basic genePred Basic Gene Annotation Set from GENCODE Version 41 (Ensembl 107) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 41 (Ensembl 107)\ parent wgEncodeGencodeV41ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV41\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV42 Basic genePred Basic Gene Annotation Set from GENCODE Version 42 (Ensembl 108) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 42 (Ensembl 108)\ parent wgEncodeGencodeV42ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV42\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV43 Basic genePred Basic Gene Annotation Set from GENCODE Version 43 (Ensembl 109) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 43 (Ensembl 109)\ parent wgEncodeGencodeV43ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV43\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV44 Basic genePred Basic Gene Annotation Set from GENCODE Version 44 (Ensembl 110) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 44 (Ensembl 110)\ parent wgEncodeGencodeV44ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV44\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV45 Basic genePred Basic Gene Annotation Set from GENCODE Version 45 (Ensembl 111) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 45 (Ensembl 111)\ parent wgEncodeGencodeV45ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV45\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV46 Basic genePred Basic Gene Annotation Set from GENCODE Version 46 (Ensembl 112) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 46 (Ensembl 112)\ parent wgEncodeGencodeV46ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV46\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV47 Basic genePred Basic Gene Annotation Set from GENCODE Version 47 (Ensembl 113) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 47 (Ensembl 113)\ parent wgEncodeGencodeV47ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV47\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV48 Basic genePred Basic Gene Annotation Set from GENCODE Version 48 (Ensembl 114) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 48 (Ensembl 114)\ parent wgEncodeGencodeV48ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV48\ trackHandler wgEncodeGencode\ type genePred\ wgEncodeGencodeBasicV49 Basic genePred Basic Gene Annotation Set from GENCODE Version 49 (Ensembl 115) 3 1 0 0 0 127 127 127 0 0 0 genes 1 longLabel Basic Gene Annotation Set from GENCODE Version 49 (Ensembl 115)\ parent wgEncodeGencodeV49ViewGenes on\ priority 1\ shortLabel Basic\ subGroups view=aGenes name=Basic\ track wgEncodeGencodeBasicV49\ trackHandler wgEncodeGencode\ type genePred\ bismap24Pos Bismap S24 + bigBed 6 Single-read mappability with 24-mers after bisulfite conversion (forward strand) 1 1 240 20 80 247 137 167 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k24.C2T-Converted.bb\ color 240,20,80\ longLabel Single-read mappability with 24-mers after bisulfite conversion (forward strand)\ parent bismapBigBed on\ priority 1\ shortLabel Bismap S24 +\ subGroups view=SR\ track bismap24Pos\ visibility dense\ wgEncodeReg4MarkCtcfBlood Blood bigWig Avg. CTCF level of 3 blood experiments (tissues and primary cells only) 0 1 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBloodCTCF.bw\ color 254,75,173\ longLabel Avg. CTCF level of 3 blood experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf\ priority 1\ shortLabel Blood\ track wgEncodeReg4MarkCtcfBlood\ type bigWig\ wgEncodeReg4AtacBlood Blood bigWig Avg. ATAC level of 48 blood experiments (tissues and primary cells only) 0 1 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBloodATAC.bw\ color 254,75,173\ longLabel Avg. ATAC level of 48 blood experiments (tissues and primary cells only)\ parent wgEncodeReg4Atac\ priority 1\ shortLabel Blood\ track wgEncodeReg4AtacBlood\ type bigWig\ wgEncodeReg4DnaseBlood Blood bigWig Avg. DNase level of 359 blood experiments (tissues and primary cells only) 0 1 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBloodDNase.bw\ color 254,75,173\ longLabel Avg. DNase level of 359 blood experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase\ priority 1\ shortLabel Blood\ track wgEncodeReg4DnaseBlood\ type bigWig\ wgEncodeReg4MarkH3k27acBlood Blood bigWig Avg. H3K27ac level of 142 blood experiments (tissues and primary cells only) 2 1 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBloodH3K27ac.bw\ color 254,75,173\ longLabel Avg. H3K27ac level of 142 blood experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac\ priority 1\ shortLabel Blood\ track wgEncodeReg4MarkH3k27acBlood\ type bigWig\ wgEncodeReg4MarkH3k4me3Blood Blood bigWig Avg. H3K4me3 level of 146 blood experiments (tissues and primary cells only) 0 1 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBloodH3K4me3.bw\ color 254,75,173\ longLabel Avg. H3K4me3 level of 146 blood experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3\ priority 1\ shortLabel Blood\ track wgEncodeReg4MarkH3k4me3Blood\ type bigWig\ cnvDevDelayCase Case gvf Copy Number Variation Morbidity Map of Developmental Delay - Case 3 1 0 0 0 127 127 127 0 0 0 phenDis 1 longLabel Copy Number Variation Morbidity Map of Developmental Delay - Case\ parent cnvDevDelay on\ priority 1\ shortLabel Case\ track cnvDevDelayCase\ type gvf\ visibility pack\ clinGenHaplo ClinGen Haploinsufficiency bigBed 9 + ClinGen Dosage Sensitivity Map - Haploinsufficiency 3 1 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/bbi/clinGen/clinGenHaplo.bb\ dataVersion /gbdb/$D/bbi/clinGen/clinGenDosageVersion.txt\ filterLabel.haploScore Dosage Sensitivity Score\ filterValues.haploScore 0|No evidence available,1|Little evidence for dosage pathogenicity,2|Some evidence for dosage pathogenicity,3|Sufficient evidence for dosage pathogenicity,30|Gene associated with autosomal recessive phenotype,40|Dosage sensitivity unlikely\ longLabel ClinGen Dosage Sensitivity Map - Haploinsufficiency\ mouseOver Gene/ISCA ID: $name\ This track shows structural variants (SVs) from the \ Consortium of Long-Read Sequencing database (CoLoRSdb).\ The sequencing data was contributed by labs and research groups around the world and covers 1,427 individuals in total, all sequenced with PacBio HiFi.\ The track contains 426,239 SVs: 232,973 insertions,\ 192,534 deletions and 732 inversions, with per-site allele frequencies,\ genotype counts and Hardy-Weinberg statistics across the cohort.\
\\ Note that CoLoRSdb also published short variants, in the Genome Browser,\ these can be found in the Variants Frequencies track.\
\ \\ Items are colored by SV type:\
\ Insertions are placed at the insertion site; deletions and inversions span\ the affected reference interval. Filters are available for SV type, SV\ length and alternate allele count. Mousing over an item shows the SV type,\ length, allele frequency, allele counts (homozygous / heterozygous /\ hemizygous) and the number of carrier samples.\
\\ The detail page additionally shows the total allele number (AN), the\ Hardy-Weinberg equilibrium p-value (HWE), the excess-heterozygosity p-value\ (ExcHet) and the REF / ALT allele sequences.\
\ \\ SVs were called on each sample's long-read alignments with\ pbsv\ and then merged across the CoLoRSdb cohort with\ Jasmine\ to produce a site-level joint callset. Per-site allele counts, allele\ frequencies, HWE and ExcHet p-values were computed from the joint VCF. The\ VCF was converted to a bigBed for display in the Genome Browser.\
\\ The step-by-step build commands are recorded in the UCSC makeDoc,\ doc/hg38/lrSv.txt;\ the conversion scripts and autoSql schemas live in\ makeDb/scripts/lrSv,\ and the track configuration is in\ trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator, and accessed\ programmatically through our API,\ track=colorsDbSv.\
\\ The bigBed is available from\ our\ download server as sv.hg38.bb. Example:\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/colorsDb/sv.hg38.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\\ The original VCF files and full release documentation are available from\ the CoLoRSdb v1.2.0 dataset on Zenodo:\ zenodo.org/records/14814308.\
\ \\ Thanks to Mike Schatz, Evan Eichler, and all\ CoLoRSdb investigators\ for generating and making the data publicly available.\
\ \\ Lake, J. A., & Consortium of Long-Read Sequencing (CoLoRS).\ Consortium of Long-Read\ Sequencing Database (CoLoRSdb) (v1.2.0) [Data set].\ Zenodo. 2025 Feb 5.\ DOI: 10.5281/zenodo.14814308\
\ \\ Kirsche M, Prabhu G, Sherman R, Ni B, Battle A, Aganezov S, Schatz MC.\ \ Jasmine and Iris: population-scale structural variant comparison and analysis.\ Nat Methods. 2023 Mar;20(3):408-417.\ PMID: 36658279;\ PMC: PMC10006329\
\ \\ Eisfeldt J, Ameur A, Lenner F, Ten Berk de Boer E, Ek M, Wincent J, Vaz R, Ottosson J, Jonson T,\ Ivarsson S et al.\ \ A national long-read sequencing study on chromosomal rearrangements uncovers hidden complexities.\ Genome Res. 2024 Nov 20;34(11):1774-1784.\ PMID: 39472022;\ PMC: PMC11610602\
\ varRep 1 bigDataUrl /gbdb/hg38/lrSv/colorsDb/sv.hg38.bb\ dataVersion v1.2.0\ filter.AC 0:2854\ filter.AF 0:1\ filter.insLen 0:18724\ filter.svLen 0:101381\ filterByRange.AC on\ filterByRange.AF on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Alt Allele Count (AC)\ filterLabel.AF Allele Frequency (AF)\ filterLabel.insLen Insertion Length (bp)\ filterLabel.svLen SV Length (bp)\ filterLabel.svType SV Type\ filterLimits.AF 0:1\ filterType.svType multipleListOr\ filterValues.svType DEL,INS,INV,DUP\ itemRgb on\ longLabel Structural Variants from CoLoRSdb (Consortium of Long-Read Sequencing, 1,427 Samples)\ mouseOver Var: $name ($svType)CpG islands are associated with genes, particularly housekeeping\ genes, in vertebrates. CpG islands are typically common near\ transcription start sites and may be associated with promoter\ regions. Normally a C (cytosine) base followed immediately by a \ G (guanine) base (a CpG) is rare in\ vertebrate DNA because the Cs in such an arrangement tend to be\ methylated. This methylation helps distinguish the newly synthesized\ DNA strand from the parent strand, which aids in the final stages of\ DNA proofreading after duplication. However, over evolutionary time,\ methylated Cs tend to turn into Ts because of spontaneous\ deamination. The result is that CpGs are relatively rare unless\ there is selective pressure to keep them or a region is not methylated\ for some other reason, perhaps having to do with the regulation of gene\ expression. CpG islands are regions where CpGs are present at\ significantly higher levels than is typical for the genome as a whole.
\ \\ The unmasked version of the track displays potential CpG islands\ that exist in repeat regions and would otherwise not be visible\ in the repeat masked version.\
\ \\ By default, only the masked version of the track is displayed. To view the\ unmasked version, change the visibility settings in the track controls at\ the top of this page.\
\ \CpG islands were predicted by searching the sequence one base at a\ time, scoring each dinucleotide (+17 for CG and -1 for others) and\ identifying maximally scoring segments. Each segment was then\ evaluated for the following criteria:\ \
\ The entire genome sequence, masking areas included, was\ used for the construction of the track Unmasked CpG.\ The track CpG Islands is constructed on the sequence after\ all masked sequence is removed.\
\ \The CpG count is the number of CG dinucleotides in the island. \ The Percentage CpG is the ratio of CpG nucleotide bases\ (twice the CpG count) to the length. The ratio of observed to expected \ CpG is calculated according to the formula (cited in \ Gardiner-Garden et al. (1987)):\ \
Obs/Exp CpG = Number of CpG * N / (Number of C * Number of G)\ \ where N = length of sequence.\
\ The calculation of the track data is performed by the following command sequence:\
\
twoBitToFa assembly.2bit stdout | maskOutFa stdin hard stdout \\\
| cpg_lh /dev/stdin 2> cpg_lh.err \\\
| awk '{$2 = $2 - 1; width = $3 - $2; printf("%s\\t%d\\t%s\\t%s %s\\t%s\\t%s\\t%0.0f\\t%0.1f\\t%s\\t%s\\n", $1, $2, $3, $5, $6, width, $6, width*$7*0.01, 100.0*2*$6/width, $7, $9);}' \\\
| sort -k1,1 -k2,2n > cpgIsland.bed\
\
The unmasked track data is constructed from\
twoBitToFa -noMask output for the twoBitToFa command.\
\
\
\ CpG islands and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator.\ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\\ The source for the cpg_lh program can be obtained from\ src/utils/cpgIslandExt/.\ The cpg_lh program binary can be obtained from: http://hgdownload.soe.ucsc.edu/admin/exe/linux.x86_64/cpg_lh (choose "save file")\
\ \This track was generated using a modification of a program developed by G. Micklem and L. Hillier \ (unpublished).
\ \\ Gardiner-Garden M, Frommer M.\ \ CpG islands in vertebrate genomes.\ J Mol Biol. 1987 Jul 20;196(2):261-82.\ PMID: 3656447\
\ regulation 1 html cpgIslandSuper\ longLabel CpG Islands (Islands < 300 Bases are Light Green)\ parent cpgIslandSuper pack\ priority 1\ shortLabel CpG Islands\ track cpgIslandExt\ cq56Vcf CQ-56 Variants vcfTabix CQ-56 Variants 0 1 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/highRepro/CQ-56.sort.vcf.gz\ longLabel CQ-56 Variants\ parent highReproVcfs\ shortLabel CQ-56 Variants\ subGroups view=vcfs\ track cq56Vcf\ type vcfTabix\ crossTissueMapsTissueCellType Cross Tissue Nuclei bigBarChart Cross tissue nuclei RNA by tissue and cell type 3 1 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$\
This track collection shows data from \
Single-nucleus cross-tissue molecular reference maps toward\
understanding disease gene function. The dataset covers ~200,000 single nuclei\
from a total of 16 human donors across 25 samples, using 4 different sample preparation\
protocols followed by droplet based single-cell RNA-seq. The samples were obtained from\
frozen tissue as part of the Genotype-Tissue Expression (GTEx) project.\
Samples were taken from the esophagus, skeletal muscle, heart, lung, prostate, breast,\
and skin. The dataset includes 43 broad cell classes, some specific to certain tissues\
and some shared across all tissue types.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ This track collection contains three bar chart tracks of RNA expression. The first track,\ Cross Tissue Nuclei, allows\ cells to be grouped together and faceted on up to 4 categories: tissue, cell class, cell subclass,\ and cell type. The second track,\ Cross Tissue Details, allows\ cells to be grouped together and faceted on up to 7 categories: tissue, cell class, cell subclass,\ cell type, granular cell type, sex, and donor. The third track,\ GTEx Immune Atlas,\ allows cells to be grouped together and faceted on up to 5 categories: tissue, cell type, cell\ class, sex, and donor.\
\ \\ Please see the\ GTEx portal\ for further interactive displays and additional data.
\ \\ Tissue-cell type combinations in the Full and Combined tracks are\ colored by which cell type they belong to in the below table:\
\
| Color | \Cell Type | \
|---|---|
| Endothelial | |
| Epithelial | |
| Glia | |
| Immune | |
| Neuron | |
| Stromal | |
| Other |
\ Tissue-cell type combinations in the Immune Atlas track are shaded according\ to the below table:\
| Color | \Cell Type | \
|---|---|
| Inflammatory Macrophage | |
| Lung Macrophage | |
| Monocyte/Macrophage FCGR3A High | |
| Monocyte/Macrophage FCGR3A Low | |
| Macrophage HLAII High | |
| Macrophage LYVE1 High | |
| Proliferating Macrophage | |
| Dendritic Cell 1 | |
| Dendritic Cell 2 | |
| Mature Dendritic Cell | |
| Langerhans | |
| CD14+ Monocyte | |
| CD16+ Monocyte | |
| LAM-like | |
| Other |
\ Using the previously collected tissue samples from the Genotype-Tissue Expression\ project, nuclei were isolated using four different protocols and sequenced\ using droplet based single cell RNA-seq. CellBender v2.1 and other standard quality\ control techniques were applied, resulting in 209,126 nuclei profiles across eight\ tissues, with a mean of 918 genes and 1519 transcripts per profile.\
\ \\ Data from all samples was integrated with a conditional variation autoencoder\ in order to correct for multiple sources of variation like sex, and protocol\ while preserving tissue and cell type specific effects.\
\ \\ For detailed methods, please refer to Eraslan et al, or the\ \ GTEx portal website.\
\ \\
The gene expression files were downloaded from the\
\
GTEx portal. The UCSC command line utilities matrixClusterColumns,\
matrixToBarChartBed, and bedToBigBed were used to transform\
these into a bar chart format bigBed file that can be visualized.\
The UCSC utilities can be found on\
our download server.\
\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions or our Data Access FAQ for more\ information.
\ \\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the GTEx Consortium for creating and analyzing these data.
\ \\ Eraslan G, Drokhlyansky E, Anand S, Fiskin E, Subramanian A, Slyper M, Wang J, Van Wittenberghe N,\ Rouhana JM, Waldman J et al.\ \ Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function.\ Science. 2022 May 13;376(6594):eabl4290.\ PMID: 35549429; PMC: PMC9383269\
\ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/crossTissueMaps/tissue_cell_type.categories\ barChartFacets tissue,cell_class,cell_subclass,cell_type\ barChartMerge on\ barChartMetric gene/genome\ barChartStatsUrl /gbdb/hg38/bbi/crossTissueMaps/tissue_cell_type.facets\ barChartStretchToItem on\ barChartUnit parts per million\ bigDataUrl /gbdb/hg38/bbi/crossTissueMaps/tissue_cell_type.bb\ defaultLabelFields name\ html crossTissueMaps\ labelFields name,name2\ longLabel Cross tissue nuclei RNA by tissue and cell type\ parent crossTissueMaps\ priority 1\ shortLabel Cross Tissue Nuclei\ track crossTissueMapsTissueCellType\ type bigBarChart\ url https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$NOTE:
\
While the DECIPHER database is \
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions.\
Because the UCSC Genes mappings for CNVs are based on associations from\ RefSeq and UniProt, they are dependent on any interpretations from those\ sources. Furthermore, because many DECIPHER records refer to multiple gene\ names, or syndromes not tightly mapped to individual genes, the associations\ in this track should be treated with skepticism and any conclusions\ based on them should be carefully scrutinized using independent\ resources.\
\Data Display Agreement Notice
\
The CNV/SNV data are only available for display in the Browser, and not for bulk\
download. Access to bulk data may be obtained directly from DECIPHER\
(https://www.deciphergenomics.org/about/data-sharing) and is subject to a\
Data Access Agreement, in which the user certifies that no attempt to\
identify individual patients will be undertaken. The same restrictions\
apply to the public data displayed at UCSC in the UCSC Genome Browser;\
no one is authorized to attempt to identify patients by any means.\
These data are made available as soon as possible and may be a\ pre-publication release. For information on the proper use of DECIPHER\ data, please see https://www.deciphergenomics.org/about/data-sharing.\
\The DECIPHER consortium provides these data in good faith as a research\ tool, but without verifying the accuracy, clinical validity, or utility of\ the data. The DECIPHER consortium makes no warranty, express or implied,\ nor assumes any legal liability or responsibility for any purpose for\ which the data are used.\
\\ The \ DECIPHER\ database of submicroscopic chromosomal imbalance \ collects clinical information about chromosomal \ microdeletions/duplications/insertions, translocations and inversions, \ and displays this information on the human genome map.\
\ The CNVs and SNVs tracks show genomic regions of reported cases and their \ associated phenotype information. All data have passed the strict\ consent requirements of the DECIPHER project and are approved for\ unrestricted public release. Clicking the Patient View ID link\ brings up a more detailed informational page on the patient at the \ DECIPHER web site.
\ \\ The Population CNVs track shows common copy-number variants (CNVs) and their\ population frequencies, lifted over from the hg19 assembly.
\ \\ The genomic locations of DECIPHER variants are labeled with the DECIPHER variant descriptions. \ Mouseover on items shows variant details, clinical interpretation, and associated conditions. \ Further information on each variant is displayed on the details page by a click onto any variant. \
\ \\ For the CNVs track, the entries are colored by the type of variant:\
\ A light-to-dark color gradient indicates the clinical significance of each variant, with \ the lightest shade being benign, to the darkest shade being pathogenic. Detailed information on the \ CNV color code is described here.\ Items can be filtered according to the size of the variant, variant type, and clinical significance \ using the track Configure options.\
\ \\ For the SNVs track, the entries are colored according to the estimated clinical significance \ of the variant:\
\ For the Population CNVs track, genomic variants are visually differentiated to facilitate quick and\ clear identification. Variants are colored according to their clinical significance and type:\
\\ The Population CNVs track's mouseover tooltip provides the following information\ about the data:\
\\ Data provided by the DECIPHER project group are imported and processed\ to create a simple BED track to annotate the genomic regions associated\ with individual patients.\
\ \ \\ For more information on DECIPHER, please contact\ \ contact@deciphergenomics.\ org\
\ \\ The DECIPHER data access and documentation can be found at\ DECIPHER Downloads.\
\ \\ Firth HV, Richards SM, Bevan AP, Clayton S, Corpas M, Rajan D, Van Vooren S, Moreau Y, Pettett RM,\ Carter NP.\ \ DECIPHER: Database of Chromosomal Imbalance and Phenotype in Humans Using Ensembl Resources.\ Am J Hum Genet. 2009 Apr;84(4):524-33.\ PMID: 19344873; PMC: PMC2667985\
\ phenDis 1 bigDataUrl /gbdb/hg38/decipher/decipherCnv.bb\ filter.size 0\ filterByRange.size on\ filterLimits.size 2:170487333\ filterValues.pathogenicity Benign,Likely Benign,Likely Pathogenic,Pathogenic,Uncertain,Unknown\ filterValues.variant_class Amplification,Copy-Number Gain,Deletion,Duplication,Duplication/Trip\ group phenDis\ html decipherContainer\ itemRgb on\ longLabel DECIPHER CNVs\ mergeSpannedItems on\ mouseOver Position: $chrom:${chromStart}-${chromEnd}\ These tracks contain information relevant to the regulation of transcription from the\ ENCODE Project.\ \
\ These tracks complement each other and together can shed much light on regulatory DNA. The histone\ marks are informative at a high level, but they have a resolution of just ~200 bases and do not\ provide much in the way of functional detail. The DNase hypersensitivity assay is higher in\ resolution at the DNA level and can be done on a large number of cell types since it's just \ a single assay. At the functional level, DNase hypersensitivity suggests that a \ region is very likely to be regulatory in nature, but provides little information beyond that.\ The transcription factor ChIP assay has a high resolution at the DNA level and, due to the very\ specific nature of the transcription factors, is often informative with respect to functional\ detail. However, since each transcription factor must be assayed separately, the information is\ only available for a limited number of transcription factors on a limited number of cell lines. \ Though each assay has its strengths and weaknesses, the fact that all of these assays are \ relatively independent of each other gives increased confidence when multiple tracks are \ suggesting a regulatory function for a region.\
\ \\ For additional information, please click on the hyperlinks for the individual tracks above.\ Also note that additional histone marks and transcription information is available in other\ ENCODE tracks. This integrative supertrack just shows a selection of the most informative data of\ most general interest.\
\ \\ By default, the transcription and histone mark displays use a transparent overlay method of \ displaying data from a number of cell lines in a single track. Each of the cell lines in this track\ is associated with a particular color, and these colors are relatively light and saturated so\ as to work best with the transparent overlay. The color of the transcription and histone mark tracks\ match their versions from their lifted source on the hg19 assembly.
\\ The DNase tracks, which were not lifted from hg19, are colored differently \ to reflect similarity of cell types. There are three DNase tracks starting with a transparent\ overlay DNase Signal Track to allow viewing signals from all 95 cell types in one track.\ The individual signals and the same coloring scheme can also be found in the DNase HS Track\ where processed peaks and hotspots are also called out as gray boxes with the darkness of\ each box reflecting the underlying signal value. Lastly, in the DNase Clusters track all observed\ hypersensitive regions in the different cell lines at the same location were clustered into a single box\ where a number to the left of the box indicates how many cell types showed a hypersensitivity \ region and the darkness of the grey box is proportional to the maximum value seen from one of\ the underlying cell lines. Clicking on these item takes you to a details page where\ additional information displays, such as the list of cell types that combined to form\ the cluster in the DNase Clusters track.\
\ \\ The raw data for ENCODE 3 Regulation tracks can be accessed from \ \ Table Browser or combined with other data-sets through \ Data Integrator. For automated analysis and downloads, the track data files can be downloaded \ from our downloads server or queried\ using the JSON API or the \ Public SQL Individual regions or the whole genome \ annotation can be accessed as text using our utility bigBedToBed. Instructions for downloading \ the utility can be found \ here. That \ utility can also be used to obtain features within a given range, e.g. \ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/wgEncodeRegDnase/wgEncodeRegDnaseUwA549Hotspot.broadPeak.bb -chrom=chr21 -start=0 -end=100000000 stdout
\\ For sorting transcription factor binding sites by cell type, we recommend you use the following\ download \ file for hg38.\
\ \ \\ Specific labs and contributors for these datasets are listed in the Credits section \ of the individual tracks in this super-track. The integrative view presented here was developed by Jim Kent at UCSC.
\ \Users may freely download, analyze and publish results based on any ENCODE data without \ restrictions.\ Researchers using unpublished ENCODE data are encouraged to contact the data producers to discuss possible coordinated publications; however, this is optional.
\ Users of ENCODE datasets are requested to cite the ENCODE Consortium and ENCODE\ production laboratory(s) that generated the datasets used, as described in\ Citing ENCODE.\ regulation 1 canPack On\ group regulation\ longLabel Integrated Regulation from ENCODE\ priority 1\ shortLabel ENCODE Regulation\ superTrack on show\ track wgEncodeReg\ cCREregistry ENCODE4 cCREs bigBed 9 + 5 ENCODE4 Registry of candidate Cis-Regulatory Elements (cCREs) 4 1 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$\ This track displays the ENCODE Registry of candidate cis-Regulatory Elements (cCREs) \ in the human genome from ENCODE 4. A total of 2,348,854 elements identified and classified by the \ ENCODE Data Analysis Center according to biochemical signatures. Most cCREs are anchored on \ DNase hypersensitive sites further annotated with histone modifications (H3K4me3 and H3K27ac) \ or CTCF binding measured by ChIP-seq experiments. In this latest version of the Registry (V4), \ the representative DNase hypersensitive sites (rDHSs) were supplemented \ with 86,748 representative transcription factor ChIP-seq peaks (TF \ rPeaks)—peaks that represent binding sites for at least five TFs. The Registry of cCREs is \ one of the core components of the integrative level of the ENCODE Encyclopedia of DNA Elements.
\ \Additional exploration of the cCREs and underlying raw ENCODE signal data can be done with the\ Core Collection track. The data is also available on the SCREEN (Search Candidate cis-Regulatory \ Elements) web tool, designed specifically for the Registry, accessible by item mouseovers and linkouts from the \ track details page.
\ \\ Each cCRE is displayed as a colored box by type, which reflects its putative functional assignment \ based on biochemical signatures and genomic context:
\\

\ Mousing over the data will display the accession ID, the assigned cCRE class type, and the Max-Z scores\ for the various underlying biosignals (DNase, H3K4me3, H3K27ac, CTCF). A track filter is also available\ to selectively show items based on their cCRE class type.
\ \\ Candidate cis-regulatory elements (cCREs) were first anchored on nucleosome-sized DNase \ hypersensitive sites (rDHSs) identified from DNase-seq data. These rDHSs were then annotated \ using ChIP-seq data for histone modifications—H3K4me3 and H3K27ac, marking promoters and \ enhancers, respectively—and CTCF, marking insulators. To supplement rDHS-anchored cCRE \ definitions, transcription factor ChIP-seq peaks were incorporated, enabling identification \ of cCREs even in regions of low chromatin accessibility. Although not used for anchoring, \ ATAC-seq data were used to assess chromatin accessibility in biosamples lacking DNase-seq.
\ \\ Classification of cCRE's was performed based on the following criteria:
\\ The ENCODE accession numbers of the constituent datasets at the ENCODE Portal are available from the cCRE details page.
\\
The data in this track can be interactively explored with the Table Browser or the Data Integrator. The data can be accessed from \
scripts through our a
\
For automated download and analysis, this annotation is stored in a bigBed file \
that can be downloaded from our download server. \
The file for this track is called cCREregistry.bb. Individual regions or the whole genome \
annotation can be obtained using our tool bigBedToBed which can be compiled from the source \
code or downloaded as a precompiled binary for your system. Instructions for downloading \
source code and binaries can be found here. \
The tool can also be used to obtain only features within a given range, e.g.
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/ccre/cCREregistry.bb -chrom=chr21 -start=0 -end=100000000 stdout
\ Data were generated by the ENCODE Consortium. The data were further processed for \ visualization through a collaborative effort between the Weng lab and the Moore lab at UMass Chan Medical \ School (funded by NIH grant HG012343). Integration and visualization were developed \ by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School, who were \ part of the ENCODE Data Analysis Center. We thank the ENCODE production labs \ for generating the data.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J, Kawli T,\ Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N, Fu Y et\ al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/encodeCcreRegistry.bb\ dataVersion ENCODE Registry version 4, 2024. (ENCODE4 data includes ENCODE2, ENCODE3, and the Roadmap Epigenomics Project)\ filterType.cCRE_class multipleListOr\ filterValues.cCRE_class CA|Chromatin accessibility (CA),CA-CTCF|Chromatin accessibility + CTCF (CA-CTCF),CA-H3K4me3|Chromatin accessibility + H3K4me3 (CA-H3K4me3),CA-TF|Chromatin accessibility + transcription factor (CA-TF),Distal enhancer|Distal enhancer,Proximal enhance|Proximal enhance,Promoter|Promoter,TF|Transcription factor (TF)\ longLabel ENCODE4 Registry of candidate Cis-Regulatory Elements (cCREs)\ mouseOver ID: ${name}\ These tracks represent the experimentally validated promoters generated by \ the Eukaryotic Promoter Database.\
\ \\ Each item in the track is a representation of the promoter sequence identified by EPD. The\ "thin" part of the element represents the 49 bp upstream of the annotated transcription\ start site (TSS) whereas the "thick" part represents the TSS plus 10 bp downstream. The\ relative position of the thick and thin parts define the orientation of the promoter.
\\ Note that the EPD team has created a public track hub containing\ promoter and supporting annotations for human, mouse, and other vertebrate and model organism\ genomes.
\ \\ Briefly, gene transcript coordinates were obtained from multiple sources (HGNC, GENCODE, Ensembl,\ RefSeq) and validated using data from CAGE and RAMPAGE experimental studies obtained from FANTOM 5,\ UCSC, and ENCODE. Peak calling, clustering and filtering based on relative expression were applied\ to identify the most expressed promoters and those present in the largest number of samples.
\\ For the methodology and principles used by EPD to predict TSSs, refer to Dreos et al.\ (2013) in the References section below. A more detailed description of how this data was\ generated can be found at the following links:\ \
\ Data was generated by the EPD team at the \ Swiss Institute of Bioinformatics. \ For inquiries, contact the EPD team using this on-line form \ or email \ \ philipp.\ bucher@epfl.\ ch\ \ .\
\ \\ Dreos R, Ambrosini G, Perier RC, Bucher P.\ \ EPD and EPDnew, high-quality promoter resources in the\ next-generation sequencing era. Nucleic Acids\ Res. 2013 Jan 1;41(D1):D157-64. PMID: 23193273.\
\ \ expression 1 bigDataUrl /gbdb/hg38/bbi/epdNewHuman006.hg38.bb\ color 50,50,200\ dataVersion EPDNew Human Version 006 (May 2018)\ longLabel Promoters from EPDnew human version 006\ parent epdNew on\ priority 1\ shortLabel EPDnew v6\ track epdNewPromoter\ url https://epd.epfl.ch/cgi-bin/get_doc?db=hgEpdNew&format=genome&entry=$$\ fixSeqLiftOverPsl Fix Patches psl Reference Assembly Fix Patch Sequence Alignments 3 1 231 203 21 243 229 138 0 0 0\ This track shows alignments of fix patch sequences to\ main chromosome sequences in the reference genome assembly.\ When errors are corrected in the reference genome assembly, the\ Genome Reference Consortium\ (GRC) adds fix patch sequences containing the corrected regions.\ This strikes a balance between providing the most complete and correct genome\ sequence, while maintaining stable chromosome coordinates for the original assembly\ sequences.\
\\ Fix patches are often associated with incident reports displayed in the GRC Incidents\ track.\
\ \\ This track follows the display conventions for\ \ PSL alignment tracks.\ Mismatching bases are highlighted in red.\ Several types of alignment gap may also be colored;\ for more information, see\ \ Alignment Insertion/Deletion Display Options.\
\ \\ The alignments were provided by NCBI as GFF files and translated into the PSL\ representation for browser display by UCSC.\
\ map 1 baseColorDefault diffBases\ baseColorUseSequence db\ color 231,203,21\ darkerLabels on\ group map\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Reference Assembly Fix Patch Sequence Alignments\ parent patchesPsl\ pennantIcon p14 black https://genome-blog.gi.ucsc.edu/blog/patches/ "Includes annotations on GRCh38.p14 patch sequences"\ priority 1\ shortLabel Fix Patches\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ track fixSeqLiftOverPsl\ type psl\ visibility pack\ knownGene GENCODE V49 bigGenePred knownGenePep knownGeneMrna GENCODE V49 3 1 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 49, September 2025) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The v49 release was derived from the GTF file that contains annotations only on the main\ chromosomes. Statistics for this build and information on how they were generated can be found on\ the GENCODE site.
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is mostly based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \ \\ Within a gene using the pack display mode, transcripts below a specified rank will be\ condensed into a view similar to squish mode. The transcript ranking approach is\ preliminary and will change in future releases. The transcripts rankings are defined by the\ following criteria for protein-coding and non-coding genes:
\ Protein_coding genes\\
The GENCODE v49 track was built from the GENCODE downloads file \
gencode.v49.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources\
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney. This version of the track was\ generated by Jonathan Casper.
\ \\ Mudge JM, Carbonell-Sala S, Diekhans M, Martinez JG, Hunt T, Jungreis I, Loveland JE, Arnan C,\ Barnes I, Bennett R et al.\ \ GENCODE 2025: reference gene annotation for human and mouse.\ Nucleic Acids Res. 2025 Jan 6;53(D1):D966-D975.\ PMID: 39565199; PMC: PMC11701607\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV49.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ downloadUrl.1 "GFF Format" https://hgdownload.soe.ucsc.edu/goldenPath/hg38/bigZips/genes/hg38.knownGene.gtf.gz\ group genes\ hgsid on\ html knownGeneV49\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V49\ maxItems 50000\ priority 1\ searchIndex name\ shortLabel GENCODE V49\ squishyPackField rank\ squishyPackLabel Number of transcripts shown at full height (ranked by GENCODE transcript ranking)\ squishyPackPoint 1\ table knownGene\ track knownGene\ type bigGenePred knownGenePep knownGeneMrna\ visibility pack\ pliByGene Gene LoF bigBed 12 + gnomAD Predicted Loss of Function Constraint Metrics By Gene (LOEUF and pLI) v2.1.1 3 1 0 0 0 127 127 127 0 0 0 https://gnomad.broadinstitute.org/gene/$$?dataset=gnomad_r2_1 varRep 1 bigDataUrl /gbdb/hg38/gnomAD/pLI/pliByGene.bb\ defaultLabelFields geneName\ filter._pli 0:1\ filterByRange._pli on\ filterLabel._pli Show only items between this pLI range\ itemRgb on\ labelFields name,geneName\ longLabel gnomAD Predicted Loss of Function Constraint Metrics By Gene (LOEUF and pLI) v2.1.1\ mouseOver LOEUF: $_loeuf\ GnomAD 4 used the whole-genome data from gnomAD 3 and added more exomes.\ The current v4.1 release includes a fix for the allele number\ issue.\ The v4.1 track shows variants from 807,162 individuals, including 730,947\ exomes and 76,215 genomes. This includes the 76,156 genomes from the gnomAD v3.1.2 release as well\ as new exome data from 416,555 UK Biobank individuals. For more detailed information on gnomAD\ v4.1, see the related blog post.\
\ \\ Following the conventions on the gnomAD browser, items are shaded according to their Annotation\ type:\
| pLoF | |
| Missense | |
| Synonymous | |
| Other |
\ Mouse hover on an item will display the following details about each variant:
\\ Clicking on an item will display additional details on the variant, including a population frequency\ table showing allele count in each sub-population.\
\ \\ To maintain consistency with the gnomAD website, variants are by default labeled according\ to their chromosomal start position followed by the reference and alternate alleles,\ for example "chr1-1234-T-CAG". dbSNP rsID's are also available as an additional\ label, if the variant is present in dbSnp.\
\ \\ Three filters are available for this track:\
\\ The gnomAD v4.1 data is unfiltered.
\ \\ For the full steps used to create the gnomAD tracks at UCSC, please see the\ hg38 gnomad makedoc.\
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below).
\ \\ The underlying bigBed only contains enough information necessary to use the track in the browser.\ The extra data like VEP annotations and CADD scores are available in the\ same directory\ as the bigBed but in the files details.tab.gz and details.tab.gz.gzi. The\ details.tab.gz contains the gzip compressed extra data in JSON format, and the .gzi file is\ available to speed searching of this data. Each variant has an associated md5sum in the name field\ of the bigBed which can be used along with the _dataOffset and _dataLen fields to get the\ associated external data. For example:
\ \\
# find an item of interest, the last two fields are _dataOffset and _dataLen:\
bigBedToBed genomes.bb stdout | head -4 | tail -1\
chr1 12416 12417 854246d79dc5d02dcdbd5f5438542b6e [..omitted..] 67293 902\
\
# use _dataOffset and _dataLen (add one to _dataLen for the newline character):\
bgzip -b 67293 -s 903 gnomad.v4.1.genomes.details.tab.gz\
854246d79dc5d02dcdbd5f5438542b6e {"DDX11L1": {"cons": ["non_coding_transcript_variant"...\
\
\
\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\ varRep 1 compositeTrack on\ configureByPopup off\ dataVersion Release v4.1 (April 19, 2024)\ html gnomadV4.1\ longLabel Genome Aggregation Database (gnomAD) Genome and Exome Variants v4.1\ maxItems 50000\ maxWindowCoverage 200000\ parent gnomadVariants\ priority 1\ shortLabel gnomAD v4.1\ track gnomadVariantsV4.1\ type bigBed 9 +\ visibility squish\ gnomadGenomesVariantsV4_1 gnomAD v4.1 Genomes bigBed 9 + Genome Aggregation Database (gnomAD) Genome Variants v4.1 4 1 0 0 0 127 127 127 0 0 0 https://gnomad.broadinstitute.org/variant/$s-$<_startPos>-$-$\ GnomAD 4 used the whole-genome data from gnomAD 3 and added more exomes.\ The current v4.1 release includes a fix for the allele number\ issue.\ The v4.1 track shows variants from 807,162 individuals, including 730,947\ exomes and 76,215 genomes. This includes the 76,156 genomes from the gnomAD v3.1.2 release as well\ as new exome data from 416,555 UK Biobank individuals. For more detailed information on gnomAD\ v4.1, see the related blog post.\
\ \\ Following the conventions on the gnomAD browser, items are shaded according to their Annotation\ type:\
| pLoF | |
| Missense | |
| Synonymous | |
| Other |
\ Mouse hover on an item will display the following details about each variant:
\\ Clicking on an item will display additional details on the variant, including a population frequency\ table showing allele count in each sub-population.\
\ \\ To maintain consistency with the gnomAD website, variants are by default labeled according\ to their chromosomal start position followed by the reference and alternate alleles,\ for example "chr1-1234-T-CAG". dbSNP rsID's are also available as an additional\ label, if the variant is present in dbSnp.\
\ \\ Three filters are available for this track:\
\\ The gnomAD v4.1 data is unfiltered.
\ \\ For the full steps used to create the gnomAD tracks at UCSC, please see the\ hg38 gnomad makedoc.\
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below).
\ \\ The underlying bigBed only contains enough information necessary to use the track in the browser.\ The extra data like VEP annotations and CADD scores are available in the\ same directory\ as the bigBed but in the files details.tab.gz and details.tab.gz.gzi. The\ details.tab.gz contains the gzip compressed extra data in JSON format, and the .gzi file is\ available to speed searching of this data. Each variant has an associated md5sum in the name field\ of the bigBed which can be used along with the _dataOffset and _dataLen fields to get the\ associated external data. For example:
\ \\
# find an item of interest, the last two fields are _dataOffset and _dataLen:\
bigBedToBed genomes.bb stdout | head -4 | tail -1\
chr1 12416 12417 854246d79dc5d02dcdbd5f5438542b6e [..omitted..] 67293 902\
\
# use _dataOffset and _dataLen (add one to _dataLen for the newline character):\
bgzip -b 67293 -s 903 gnomad.v4.1.genomes.details.tab.gz\
854246d79dc5d02dcdbd5f5438542b6e {"DDX11L1": {"cons": ["non_coding_transcript_variant"...\
\
\
\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/v4.1/genomes/genomes.bb\ dataVersion Release v4.1 (April 19, 2024)\ defaultLabelFields _displayName\ detailsDynamicTable _jsonVep|Variant Effect Predictor,_jsonPopTable|Population Frequencies,_jsonHapTable|Haplotype Frequencies\ detailsTabUrls _dataOffset=/gbdb/hg38/gnomAD/v4.1/genomes/gnomad.v4.1.genomes.details.tab.gz\ filter.AF 0.0\ filterLabel.AF Minor Allele Frequency Filter\ filterType.FILTER multipleListAnd\ filterType.variation_type multipleListOr\ filterValues.FILTER PASS,InbreedingCoeff,RF,AC0,AS_VQSR,indel_stack (chrM only),npg (chrM only)\ filterValues.annot pLoF,missense,synonymous,other\ filterValues.variation_type 3_prime_UTR_variant,5_prime_UTR_variant,NMD_transcript_variant,coding_sequence_variant,frameshift_variant,incomplete_terminal_codon_variant,inframe_deletion,inframe_insertion,intron_variant,mature_miRNA_variant,missense_variant,non_coding_transcript_exon_variant,non_coding_transcript_variant,protein_altering_variant,splice_acceptor_variant,splice_donor_variant,splice_region_variant,start_lost,start_retained_variant,stop_gained,stop_lost,stop_retained_variant,synonymous_variant,transcript_ablation\ filterValuesDefault.FILTER PASS\ filterValuesDefault.annot pLoF,missense,synonymous\ html gnomadV4.1\ itemRgb on\ labelFields rsId,_displayName\ longLabel Genome Aggregation Database (gnomAD) Genome Variants v4.1\ mouseOver Position: $chrom:${chromStart}-${chromEnd} ($ref/$alt)\ The\ \ NIH Genotype-Tissue Expression (GTEx) project\ was created to establish a sample and data resource for studies on the relationship between \ genetic variation and gene expression in multiple human tissues. \ This track shows median gene expression levels in 52 tissues and 2 cell lines, \ based on RNA-seq data from the GTEx final data release (V8, August 2019).\ This release is based on data from 17,382 tissue samples obtained from 948 adult \ post-mortem individuals.
\ \\
In Full and Pack display modes, expression for each gene is represented by a colored bargraph,\
where the height of each bar represents the median expression level across all samples for a \
tissue, and the bar color indicates the tissue.\
Tissue colors were assigned to conform to the GTEx Consortium publication conventions.\

\
The bargraph display has the same width and tissue order for all genes.\
Mouse hover over a bar will show the tissue and median expression level.\
The Squish display mode draws a rectangle for each gene, colored to indicate the tissue\
with highest expression level if it contributes more than 10% to the overall expression\
(and colored black if no tissue predominates).\
In Dense mode, the darkness of the grayscale rectangle displayed for the gene reflects the total\
median expression level across all tissues.
\ The GTEx transcript model used to quantify expression level is displayed below the graph,\ colored to indicate the transcript class \ (coding, \ noncoding, \ pseudogene, \ problem), \ following GENCODE conventions.\
\\ Click-through on a graph displays a boxplot of expression level quartiles with outliers, \ per tissue, along with a link to the corresponding gene page on the GTEx Portal.
\ The track configuration page provides controls to limit the genes and tissues displayed,\ and to select raw or log transformed expression level display.\ \\ RNA-seq was performed by the GTEx Laboratory, Data Analysis and Coordinating Center \ (LDACC) at the Broad Institute.\ The Illumina TruSeq protocol was used to create an unstranded polyA+ library sequenced\ on the Illumina HiSeq 2000 and HiSeq 2500 platforms to produce 76-bp paired end reads with a coverage\ goal of 50M (median achieved was ~82M total reads).\
\ Sequence reads were aligned to the hg38/GRCh38 human genome using STAR v2.5.3a\ assisted by the GENCODE 26 transcriptome definition. \ The alignment pipeline is available\ here.\ \\ Gene annotations were produced using a custom isoform collapsing procedure that excluded\ retained intron and read through transcripts, merged overlapping exon intervals and then excluded\ exon intervals overlapping between genes.\ Gene expression levels in TPM were called via the RNA-SeQC tool (v1.1.9), after filtering for \ unique mapping, proper pairing, and exon overlap.\ For further method details, see the \ \ GTEx Portal Documentation page.
\\ UCSC obtained the gene-level expression files, gene annotations and sample metadata from the \ GTEx Portal Download page.\ Median expression level in TPM was computed per gene/per tissue.
\ \\ The scientific goal of the GTEx project required that the donors and their biospecimen \ present with no evidence of disease. \ The tissue types collected were chosen based on their clinical significance, logistical \ feasibility and their relevance to the scientific goal of the project and the \ research community. \ Summary plots of GTEx sample characteristics are available at the \ \ GTEx Portal Tissue Summary page.
\ \ \\ The raw data for the GTEx Gene expression track can be accessed interactively through the \ \ Table Browser or Data Integrator. Metadata can be \ found in the connected tables below.\
\
For automated analysis and downloads, the track data files can be downloaded from \
our downloads server\
or the JSON API.\
Individual regions or the whole genome annotation can be accessed as text using our utility\
bigBedToBed. Instructions for downloading the utility can be found \
here. \
That utility can also be used to obtain features within a given range, e.g. \
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gtex/gtexGeneV8.bb -chrom=chr21\
-start=0 -end=100000000 stdout
\ Data can also be obtained directly from GTEx at the following link:\ \ https://gtexportal.org/home/datasets
\ \\ Statistical analysis and data interpretation was performed by The GTEx Consortium Analysis \ Working Group. \ Data was provided by the GTEx LDACC at The Broad Institute of MIT and Harvard.
\ \\ GTEx Consortium. \ \ The GTEx Consortium atlas of genetic regulatory effects across human tissues.\ Science. 2020 Sep 11;369(6509):1318-1330.\ PMID: 32913098; \ PMC: PMC7737656
\ \ \\ GTEx Consortium.\ \ The Genotype-Tissue Expression (GTEx) project.\ Nat Genet. 2013 Jun;45(6):580-5.\ PMID: 23715323; \ PMC: PMC4010069
\ \\ Carithers LJ, Ardlie K, Barcus M, Branton PA, Britton A, Buia SA, Compton CC, DeLuca DS, \ Peter-Demchok J, Gelfand ET et al.\ \ A Novel Approach to High-Quality Postmortem Tissue Procurement: The GTEx Project.\ Biopreserv Biobank. 2015 Oct;13(5):311-9.\ PMID: 26484571; \ PMC: PMC4675181
\ \ Melé M, Ferreira PG, Reverter F, DeLuca DS, Monlong J, Sammeth M, Young TR, Goldmann JM,\ Pervouchine DD, Sullivan TJ et al.\ \ Human genomics. The human transcriptome across tissues and individuals.\ Science. 2015 May 8;348(6235):660-5.\ PMID: 25954002; PMC: PMC4547472\ \\ DeLuca DS, Levin JZ, Sivachenko A, Fennell T, Nazaire MD, Williams C, Reich M, Winckler W, Getz G.\ \ RNA-SeQC: RNA-seq metrics for quality control and process optimization.\ Bioinformatics. 2012 Jun 1;28(11):1530-2.\ PMID: 22539670; PMC: PMC3356847
\ \ expression 1 group expression\ longLabel Gene Expression in 54 tissues from GTEx RNA-seq of 17382 samples, 948 donors (V8, Aug 2019)\ maxItems 200\ priority 1\ shortLabel GTEx Gene V8\ spectrum on\ track gtexGeneV8\ type bed 6 +\ visibility pack\ h1hescInsitu H1-hESC In situ hic In situ Hi-C Chromatin Structure on H1-hESC 0 1 0 0 0 127 127 127 0 0 0 regulation 1 bigDataUrl /gbdb/hg38/bbi/hic/4DNFIQYQWPF5.hic\ longLabel In situ Hi-C Chromatin Structure on H1-hESC\ parent hicAndMicroC off\ shortLabel H1-hESC In situ\ track h1hescInsitu\ type hic\ haqers HAQERS bigBed 4 + HAQERS: 1580 Human Ancestor Quickly Evolved Regions 0 1 0 0 0 127 127 127 0 0 0 compGeno 1 bigDataUrl /gbdb/hg38/unusualcons/haqers.bb\ longLabel HAQERS: 1580 Human Ancestor Quickly Evolved Regions\ parent unusualcons on\ shortLabel HAQERS\ track haqers\ type bigBed 4 +\ chainHprcGCA_018466845v1 HG02257.mat chain GCA_018466845.1 HG02257.mat HG02257.pri.mat.f1_v2 (May 2021 GCA_018466845.1_HG02257.pri.mat.f1_v2) HPRC project computed Chained Alignments 3 1 0 0 0 255 255 0 1 0 0 hprc 1 longLabel HG02257.mat HG02257.pri.mat.f1_v2 (May 2021 GCA_018466845.1_HG02257.pri.mat.f1_v2) HPRC project computed Chained Alignments\ otherDb GCA_018466845.1\ parent hprcChainNetViewchain off\ priority 18\ shortLabel HG02257.mat\ subGroups view=chain sample=s018 population=afr subpop=acb hap=mat\ track chainHprcGCA_018466845v1\ type chain GCA_018466845.1\ humanMethylationAtlasSummary Human Methylation Atlas Summary bigBed 4 . Human Methylation Atlas summary regions and enhancers 3 1 0 0 0 127 127 127 0 0 0\ The Human Methylation Atlas tracks display genome-wide DNA methylation profiles from \ deep whole-genome bisulfite sequencing (WGBS) of 39 primary human cell types \ sorted from 205 healthy tissue samples. This comprehensive resource enables fragment-level \ analysis across thousands of unique markers, providing a detailed reference for \ cell-type-specific methylation patterns.\
\ \ Human Methylation Atlas Summary consists of the following subtracks:\\ Unsupervised clustering of these methylomes recapitulates key elements of tissue ontogeny and\ developmental lineage relationships.\
\ \\ Tracks are colored by tissue/cell type category as follows:\
\ \| Color | Cell Type(s) |
|---|---|
| Neurons | |
| Oligodendrocytes | |
| Thyroid Epithelium | |
| Prostate Epithelium | |
| Bladder Epithelium | |
| Heart Cardiomyocytes | |
| Smooth Muscle | |
| Heart Fibroblasts | |
| Skeletal Muscle | |
| Erythrocyte Progenitors | |
| Blood Granulocytes | |
| Blood Monocytes/Macrophages | |
| Blood T Cells | |
| Blood B Cells | |
| Blood NK Cells | |
| Pancreas Beta Cells | |
| Pancreas Alpha Cells | |
| Pancreas Delta Cells | |
| Pancreas Duct Cells | |
| Pancreas Acinar Cells | |
| Colon Epithelium | |
| Colon Fibroblasts | |
| Small Intestine Epithelium | |
| Gastric Epithelium | |
| Gallbladder | |
| Liver Hepatocytes | |
| Lung Bronchus Epithelium | |
| Lung Alveolar Epithelium | |
| Kidney Epithelium | |
| Endothelial | |
| Breast Basal Epithelium | |
| Breast Luminal Epithelium | |
| Fallopian Epithelium | |
| Ovary Epithelium | |
| Adipocytes | |
| Epidermal Keratinocytes | |
| Dermal Fibroblasts | |
| Bone Osteoblasts | |
| Head Neck Epithelium |
\ Items in these tracks can be filtered by:\
\\ Primary human cells were isolated from freshly dissociated adult healthy tissues using \ fluorescence-activated cell sorting (FACS), yielding high-purity preparations across major \ cell lineages. A total of 205 samples representing 77 primary cell types were collected from\ 137 consenting donors and merged into 39 cell type groups based on methylation similarity.\ Average sample purity exceeded 90% as determined by flow cytometry, gene expression, and\ DNA methylation analysis. Some cell types showed lower purity, including colon fibroblasts (78%),\ smooth muscle cells (82%), endothelial cells (86%), and adipocytes (87%).\
\ \\ Several cell types are absent from the atlas, typically due to limited availability of primary\ material. These include osteoblasts, cholangiocytes, cells of the adrenal gland, urethral\ epithelium, and haematopoietic stem cells. Subpopulations of interest, such as distinct neuronal or\ lymphocyte subtypes, were also not resolved separately.\
\ \\ Whole-genome bisulfite sequencing was performed using 150 bp paired-end reads at an average \ sequencing depth of 30× (minimum 6.62×). Libraries were prepared using the \ Accel-NGS Methyl-Seq DNA library preparation kit and sequenced on the Illumina NovaSeq 6000 \ platform.\
\ \\ Reads were mapped to the human genome (hg38) using bwa-meth, deduplicated with Sambamba, \ and processed into per-CpG methylation calls. The genome was segmented into 7.1 million \ non-overlapping methylation blocks using a multi-channel dynamic programming algorithm \ that identifies regions of homogeneous methylation across samples.\
\ \\ Cell-type-specific differentially methylated regions were identified using a one-versus-all \ comparison approach. Regions uniquely unmethylated in specific cell types were found to be \ enriched for transcriptional enhancers and tissue-specific transcription factor binding motifs.\
\ \\ Data processing was performed using \ wgbstools, an open-source \ computational suite for DNA methylation sequencing data representation, visualization, \ and analysis.\
\ \\ The raw data for these tracks can be explored interactively using the \ Table Browser or the \ Data Integrator. \ For automated analysis, the data may also be queried from our \ REST API.\
\ \\ The complete dataset, including all WGBS data files and processed methylation calls, \ is available from GEO accession \ GSE186458.\
\ \\ For questions regarding the data, please contact \ Prof. Tommy Kaplan at the Hebrew \ University of Jerusalem.\
\ \\ Data generation and analysis were performed at the Hebrew University of Jerusalem by the \ Dor, Kaplan, and Glaser laboratories and collaborators. Sample collection involved \ collaboration with Hadassah Medical Center, Oregon Health & Science University, \ Karolinska Institute, and University of Alberta.\
\ \\ Loyfer N, Magenheim J, Peretz A, Cann G, Bredno J, Klochendler A, Fox-Fisher I, \ Shabi-Porat S, Hecht M, Pelet T et al.\ \ A DNA methylation atlas of normal human cell types.\ Nature. 2023 Jan;613(7943):355-364.\ PMID: 36599988\
\ \\ Loyfer N, Rosenski J, Kaplan T.\ \ wgbstools: a computational suite for DNA methylation sequencing data analysis.\ Life Sci Alliance. 2026 Apr;9(4):e202503514.\ PMID: 41611450\
\ \ regulation 1 compositeTrack on\ dataVersion Data release version 2\ html methylationAtlas.html\ longLabel Human Methylation Atlas summary regions and enhancers\ parent dnaMethylation\ priority 1\ shortLabel Human Methylation Atlas Summary\ showCfg on\ track humanMethylationAtlasSummary\ type bigBed 4 .\ visibility pack\ platinumHybrid hybrid vcfTabix Platinum genome hybrid 3 1 0 0 0 127 127 127 0 0 23 chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22,chrX, varRep 1 bigDataUrl /gbdb/hg38/platinumGenomes/hg38.hybrid.vcf.gz\ chromosomes chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22,chrX\ configureByPopup off\ group varRep\ longLabel Platinum genome hybrid\ maxWindowToDraw 200000\ parent platinumGenomes\ shortLabel hybrid\ showHardyWeinberg on\ track platinumHybrid\ type vcfTabix\ vcfDoFilter off\ vcfDoMaf off\ visibility pack\ xGen_Research_Probes_V1 IDT xGen V1 P bigBed IDT - xGen Exome Research Panel V1 Probes 0 1 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/xgen-exome-research-panel-probes-hg38.bb\ color 100,143,255\ longLabel IDT - xGen Exome Research Panel V1 Probes\ parent exomeProbesets off\ shortLabel IDT xGen V1 P\ track xGen_Research_Probes_V1\ type bigBed\ jarvis JARVIS bigWig JARVIS: score to prioritize non-coding regions for disease relevance 1 1 150 130 160 202 192 207 0 0 0\ The "Constraint scores" container track includes several subtracks showing the results of\ constraint prediction algorithms. These try to find regions of negative\ selection, where variations likely have functional impact. The algorithms do\ not use multi-species alignments to derive evolutionary constraint, but use\ primarily human variation, usually from variants collected by gnomAD (see the\ gnomAD V2 or V3 tracks on hg19 and hg38) or TOPMED (contained in our dbSNP\ tracks and available as a filter). One of the subtracks is based on UK Biobank\ variants, which are not available publicly, so we have no track with the raw data.\ The number of human genomes that are used as the input for these scores are\ 76k, 53k and 110k for gnomAD, TOPMED and UK Biobank, respectively.\
\ \Note that another important constraint score, gnomAD\ constraint, is not part of this container track but can be found in the hg38 gnomAD\ track.\
\ \ The algorithms included in this track are:\\ JARVIS scores are shown as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The scores were downloaded and converted to a single bigWig file.\ Move the mouse over the bars to display the exact values. A horizontal line is shown at the 0.733\ value which signifies the 90th percentile.
\ See hg19 makeDoc and\ hg38 makeDoc.\\ Interpretation: The authors offer a suggested guideline of > 0.9998 for identifying\ higher confidence calls and minimizing false positives. In addition to that strict threshold, the \ following two more relaxed cutoffs can be used to explore additional hits. Note that these\ thresholds are offered as guidelines and are not necessarily representative of pathogenicity.
\ \\
| Percentile | JARVIS score threshold |
|---|---|
| 99th | 0.9998 |
| 95th | 0.9826 |
| 90th | 0.7338 |
\ HMC scores are displayed as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The highly-constrained cutoff\ of 0.8 is indicated with a line.
\\ Interpretation: \ A protein residue with HMC score <1 indicates that missense variants affecting\ the homologous residues are significantly under negative selection (P-value <\ 0.05) and likely to be deleterious. A more stringent score threshold of HMC<0.8\ is recommended to prioritize predicted disease-associated variants.\
\ \\ Interpretation: The authors suggest the following guidelines for evaluating\ intolerance. By default, the MetaDome track displays a horizontal line at 0.7 which \ signifies the first intolerant bin. For more information see the MetaDome publication.
\ \\
| Classification | MetaDome Tolerance Score |
|---|---|
| Highly intolerant | ≤ 0.175 |
| Intolerant | ≤ 0.525 |
| Slightly intolerant | ≤ 0.7 |
\ MTR data can be found on two tracks, MTR All data and MTR Scores. In the\ MTR Scores track the data has been converted into 4 separate signal tracks\ representing each base pair mutation, with the lowest possible score shown when\ multiple transcripts overlap at a position. Overlaps can happen since this score\ is derived from transcripts and multiple transcripts can overlap. \ A horizontal line is drawn on the 0.8 score line\ to roughly represent the 25th percentile, meaning the items below may be of particular\ interest. It is recommended that the data be explored using\ this version of the track, as it condenses the information substantially while\ retaining the magnitude of the data.
\ \Any specific point mutations of interest can then be researched in the \ MTR All data track. This track contains all of the information from\ \ MTRV2 including more than 3 possible scores per base when transcripts overlap.\ A mouse-over on this track shows the ref and alt allele, as well as the MTR score\ and the MTR score percentile. Filters are available for MTR score, False Discovery Rate\ (FDR), MTR percentile, and variant consequence. By default, only items in the bottom\ 25 percentile are shown. Items in the track are colored according\ to their MTR percentile:
\\ Interpretation: Regions with low MTR scores were seen to be enriched with\ pathogenic variants. For example, ClinVar pathogenic variants were seen to\ have an average score of 0.77 whereas ClinVar benign variants had an average score\ of 0.92. Further validation using the FATHMM cancer-associated training dataset saw\ that scores less than 0.5 contained 8.6% of the pathogenic variants while only containing\ 0.9% of neutral variants. In summary, lower scores are more likely to represent\ pathogenic variants whereas higher scores could be pathogenic, but have a higher chance\ to be a false positive. For more information see the MTR-Viewer publication.
\ \\ Scores were downloaded and converted to a single bigWig file. See the\ hg19 makeDoc and the\ hg38 makeDoc for more info.\
\ \\ Scores were downloaded and converted to .bedGraph files with a custom Python \ script. The bedGraph files were then converted to bigWig files, as documented in our \ makeDoc hg19 build log.
\ \\
The authors provided a bed file containing codon coordinates along with the scores. \
This file was parsed with a python script to create the two tracks. For the first track\
the scores were aggregated for each coordinate, then the lowest score chosen for any\
overlaps and the result written out to bedGraph format. The file was then converted\
to bigWig with the bedGraphToBigWig utility. For the second track the file\
was reorganized into a bed 4+3 and conveted to bigBed with the bedToBigBed\
utility.
\ See the hg19 makeDoc for details including the build script.
\\ The raw MetaDome data can also be accessed via their Zenodo handle.
\ \\ V2\ file was downloaded and columns were reshuffled as well as itemRgb added for the\ MTR All data track. For the MTR Scores track the file was parsed with a python\ script to pull out the highest possible MTR score for each of the 3 possible mutations\ at each base pair and 4 tracks built out of these values representing each mutation.
\\ See the hg19 makeDoc entry on MTR for more info.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/hmc/hmc.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \ \\ Thanks to Jean-Madeleine Desainteagathe (APHP Paris, France) for suggesting the JARVIS, MTR, HMC tracks. Thanks to Xialei Zhang for providing the HMC data file and to Dimitrios Vitsios and Slave Petrovski for helping clean up the hg38 JARVIS files for providing guidance on interpretation. Additional\ thanks to Laurens van de Wiel for providing the MetaDome data as well as guidance on the track development and interpretation. \
\ \ \\ Vitsios D, Dhindsa RS, Middleton L, Gussow AB, Petrovski S.\ \ Prioritizing non-coding regions based on human genomic constraint and sequence context with deep\ learning.\ Nat Commun. 2021 Mar 8;12(1):1504.\ PMID: 33686085; PMC: PMC7940646\
\ \\ Xiaolei Zhang, Pantazis I. Theotokis, Nicholas Li, the SHaRe Investigators, Caroline F. Wright, Kaitlin E. Samocha, Nicola Whiffin, James S. Ware\ \ Genetic constraint at single amino acid resolution improves missense variant prioritisation and gene discovery.\ Medrxiv 2022.02.16.22271023\
\ \\ Wiel L, Baakman C, Gilissen D, Veltman JA, Vriend G, Gilissen C.\ \ MetaDome: Pathogenicity analysis of genetic variants through aggregation of homologous human protein\ domains.\ Hum Mutat. 2019 Aug;40(8):1030-1038.\ PMID: 31116477; PMC: PMC6772141\
\ \\ Silk M, Petrovski S, Ascher DB.\ \ MTR-Viewer: identifying regions within genes under purifying selection.\ Nucleic Acids Res. 2019 Jul 2;47(W1):W121-W126.\ PMID: 31170280; PMC: PMC6602522\
\ \\ Halldorsson BV, Eggertsson HP, Moore KHS, Hauswedell H, Eiriksson O, Ulfarsson MO, Palsson G,\ Hardarson MT, Oddsson A, Jensson BO et al.\ \ The sequences of 150,119 genomes in the UK Biobank.\ Nature. 2022 Jul;607(7920):732-740.\ PMID: 35859178; PMC: PMC9329122\
\ \ \\ Huang YF, Gulko B, Siepel A.\ \ Fast, scalable prediction of deleterious noncoding variants from functional and population genomic\ data.\ Nat Genet. 2017 Apr;49(4):618-624.\ PMID: 28288115; PMC: PMC5395419\
\ \ phenDis 0 bigDataUrl /gbdb/hg38/jarvis/jarvis.bw\ color 150,130,160\ group phenDis\ html constraintSuper\ longLabel JARVIS: score to prioritize non-coding regions for disease relevance\ maxHeightPixels 8:40:128\ maxWindowToDraw 10000000\ mouseOverFunction noAverage\ parent constraintSuper\ priority 1\ shortLabel JARVIS\ track jarvis\ type bigWig\ viewLimits 0.0:1.0\ visibility dense\ yLineMark 0.73\ yLineOnOff on\ jaspar2026 JASPAR 2026 TFBS bigBed 6 + JASPAR CORE 2026 - Predicted Transcription Factor Binding Sites 3 1 0 0 0 127 127 127 1 0 0 http://jaspar.genereg.net/search?q=$$&collection=all&tax_group=all&tax_id=all&type=all&class=all&family=all&version=all regulation 1 bigDataUrl /gbdb/hg38/jaspar/JASPAR2026.bb\ filter.score 400\ filterByRange.score 0:1000\ filterValues.TFName Ahr::Arnt,Alx1,ALX3,Alx4,Ar,ARGFX,Arid3a,Arid3b,Arid5a,Arnt,ARNT2,ARNT::HIF1A,Arntl,Arx,ASCL1,Ascl2,Atf1,ATF2,Atf3,ATF3,ATF4,ATF6,ATF7,Atoh1,ATOH7,BACH1,Bach1::Mafk,BACH2,Banp,BARHL1,BARHL2,BARX1,BARX2,BATF,BATF3,BATF::JUN,BCL11A,Bcl11B,BCL6,BCL6B,Bhlha15,BHLHA15,BHLHE22,BHLHE23,BHLHE40,BHLHE41,BNC2,BSX,CASZ1,CDX1,CDX2,CDX4,CEBPA,CEBPB,CEBPD,CEBPE,CEBPG,CGGBP1,CLOCK,CREB1,CREB3,CREB3L1,Creb3l2,CREB3L3,CREB3L4,Creb5,CREM,Crx,CTCF,CTCFL,CUX1,CUX2,DBP,Ddit3::Cebpa,DLX1,Dlx2,Dlx3,Dlx4,Dlx5,DLX6,Dmbx1,Dmrt1,DMRT3,DMRTA1,DMRTA2,Dmrtb1,DMRTC2,DMTF1,DNTTIP1,DPRX,DRGX,Dux,DUX4,DUXA,Duxbl1,E2F1,E2F2,E2F3,E2F4,E2F6,E2F7,E2F8,EBF1,Ebf2,EBF3,Ebf4,EGR1,EGR2,EGR3,EGR4,EHF,ELF1,ELF2,ELF3,ELF4,Elf5,ELK1,ELK1::HOXA1,ELK1::HOXB13,ELK1::SREBF2,ELK3,ELK4,EMX1,EMX2,EN1,EN2,EOMES,EPAS1,ERF,ERF::FIGLA,ERF::FOXI1,ERF::FOXO1,ERF::HOXB13,ERF::NHLH1,ERF::SREBF2,Erg,ESR1,ESR2,ESRRA,ESRRB,Esrrg,ESX1,ETS1,ETS2,ETV1,ETV2,ETV2::DRGX,ETV2::FIGLA,ETV2::FOXI1,ETV2::HOXB13,ETV3,ETV4,ETV5,ETV5::DRGX,ETV5::FIGLA,ETV5::FOXI1,ETV5::FOXO1,ETV5::HOXA2,ETV6,ETV7,EVX1,EVX2,EWSR1-FLI1,FAM200B,FERD3L,FEV,FEZF2,FIGLA,FIZ1,FLI1,FLI1::DRGX,FLI1::FOXI1,FLYWCH1,FOS,Fosb,FOSB::JUN,FOSB::JUNB,FOS::JUN,FOS::JUNB,FOS::JUND,FOSL1,FOSL1::JUN,FOSL1::JUNB,FOSL1::JUND,FOSL2,FOSL2::JUN,FOSL2::JUNB,FOSL2::JUND,FOXA1,FOXA2,FOXA3,FOXB1,FOXC1,FOXC2,FOXD1,FOXD2,FOXD3,FOXE1,Foxf1,FOXF2,FOXG1,FOXH1,FOXI1,Foxj2,FOXJ2::ELF1,Foxj3,FOXK1,FOXK2,FOXL1,Foxl2,FOXM1,Foxn1,FOXN3,Foxo1,FOXO1::ELF1,FOXO1::ELK1,FOXO1::ELK3,FOXO1::FLI1,Foxo3,FOXO4,FOXO6,FOXP1,FOXP2,FOXP3,FOXP4,Foxq1,FOXS1,GABPA,GATA1,GATA1::TAL1,GATA2,Gata3,GATA4,GATA5,GATA6,GBX1,GBX2,GCM1,GCM2,GFI1,Gfi1B,Gli1,Gli2,GLI3,GLIS1,GLIS2,GLIS3,Gmeb1,GMEB2,GRHL1,GRHL2,GSC,GSC2,GSC::TBX21,GSX1,GSX2,Hand1,Hand1::Tcf3,HAND2,HELT,HES1,HES2,HES5,HES6,HES7,HESX1,HEY1,HEY2,Hic1,HIC2,HIF1A,HIF3A,HINFP,HLF,HMBOX1,Hmga1,Hmx1,Hmx2,Hmx3,Hnf1A,HNF1A,HNF1B,HNF4A,HNF4G,HOMEZ,HOXA1,HOXA10,Hoxa11,Hoxa13,HOXA2,HOXA3,HOXA4,HOXA5,HOXA6,HOXA7,HOXA9,HOXB1,HOXB13,HOXB2,HOXB2::ELK1,HOXB3,HOXB4,HOXB5,HOXB6,HOXB7,HOXB8,HOXB9,HOXC10,HOXC11,HOXC12,HOXC13,HOXC4,HOXC8,HOXC9,HOXD1,HOXD10,HOXD11,HOXD12,HOXD12::ELK1,Hoxd13,HOXD3,HOXD4,HOXD8,HOXD9,HSF1,HSF2,HSF4,HSF5,IKZF1,IKZF2,Ikzf3,INSM1,Irf1,IRF2,IRF3,IRF4,IRF5,IRF6,IRF7,IRF8,IRF9,IRX1,IRX2,IRX5,Isl1,ISL2,ISX,JDP2,Jun,JUN,JUNB,JUND,JUN::JUNB,KLF1,KLF10,KLF11,KLF12,KLF13,KLF14,KLF15,KLF16,KLF17,KLF2,KLF3,KLF4,KLF5,KLF6,KLF7,KLF8,KLF9,LBX1,LBX2,Lef1,LEUTX,Lhx1,LHX2,Lhx3,Lhx4,LHX5,LHX6,Lhx8,LHX9,LIN54,LMX1A,LMX1B,MAF,MAFA,Mafb,MAFF,Mafg,MAFG::NRF1,MAFK,MAF::NFE2,MAX,MAX::MYC,MAZ,Mecom,MEF2A,MEF2B,MEF2C,MEF2D,MEIS1,MEIS2,MEIS3,MEOX1,MEOX2,MGA,MGA::EVX1,MITF,mix-a,MIXL1,MKX,MLX,Mlxip,MLXIPL,MNT,MNX1,MSANTD1,MSANTD3,MSANTD4,MSC,Msgn1,MSX1,MSX2,Msx3,MTF1,MXI1,MYB,MYBL1,MYBL2,MYC,MYCN,MYF5,MYF6,MYOD1,MYOG,MYPOP,MYT1,MYT1L,MZF1,NACC2,Nanog,NEUROD1,Neurod2,NEUROG1,NEUROG2,Nfat5,Nfatc1,Nfatc2,NFATC3,NFATC4,NFE2,Nfe2l2,NFIA,NFIB,NFIC,NFIC::TLX1,NFIL3,NFIX,NFKB1,NFKB2,NFYA,NFYB,NFYC,NHLH1,NHLH2,Nkx2-1,NKX2-2,NKX2-3,NKX2-4,NKX2-5,NKX2-8,Nkx3-1,Nkx3-2,NKX6-1,NKX6-2,NKX6-3,Nobox,NOTO,Npas2,NPAS3,Npas4,NR1D1,NR1D2,Nr1H2,NR1H2::RXRA,Nr1h3,Nr1h3::Rxra,Nr1H4,NR1H4::RXRA,NR1I2,NR1I3,NR2C1,NR2C2,Nr2e1,Nr2e3,NR2F1,NR2F2,Nr2f6,Nr2F6,NR2F6,NR3C1,NR3C2,NR4A1,NR4A2,NR4A2::RXRA,Nr5a1,Nr5a2,NR6A1,Nrf1,NRL,OLIG1,Olig2,OLIG2,OLIG3,ONECUT1,ONECUT2,ONECUT3,OSR1,OSR2,OTX1,OTX2,OVOL1,OVOL2,PATZ1,PAX1,PAX2,PAX3,PAX4,PAX5,PAX6,Pax7,PAX8,PAX9,PBX1,PBX2,PBX3,PDX1,Pgr,PGR,PHOX2A,PHOX2B,PITX1,PITX2,PITX3,PKNOX1,PKNOX2,PLAG1,Plagl1,PLAGL2,POGK,POU1F1,POU2F1,POU2F1::SOX2,POU2F2,POU2F3,POU3F1,POU3F2,POU3F3,POU3F4,POU4F1,POU4F2,POU4F3,POU5F1,POU5F1B,Pou5f1::Sox2,POU6F1,POU6F2,Ppara,PPARA::RXRA,PPARD,PPARG,Pparg::Rxra,PRDM1,PRDM13,Prdm14,Prdm15,Prdm4,Prdm5,PRDM9,PROP1,PROX1,PRRX1,PRRX2,Ptf1A,RARA,RARA::RXRA,RARA::RXRG,Rarb,RARB,Rarg,RARG,RAX,RAX2,RBPJ,REL,RELA,RELB,REST,RFX1,RFX2,RFX3,RFX4,RFX5,Rfx6,RFX7,Rhox11,RHOXF1,RORA,RORB,RORC,RREB1,Runx1,RUNX2,RUNX3,Rxra,RXRA::VDR,RXRB,RXRG,SALL3,SATB1,SCAND3,SCRT1,SCRT2,SHOX,Shox2,SIX1,SIX2,Six3,Six4,SLC2A4RG,SMAD2,SMAD3,Smad4,SMAD5,SNAI1,SNAI2,SNAI3,SOHLH2,Sox1,SOX10,Sox11,SOX12,SOX13,SOX14,SOX15,Sox17,SOX18,SOX2,SOX21,Sox3,SOX30,SOX4,Sox5,Sox6,Sox7,SOX8,SOX9,SP1,SP140L,SP2,SP3,SP4,SP5,SP8,SP9,SPDEF,Spi1,SPIB,SPIC,Spz1,SREBF1,SREBF2,SRF,SRY,STAT1,STAT1::STAT2,Stat2,STAT3,Stat4,Stat5a,Stat5a::Stat5b,Stat5b,Stat6,TAL1::TCF3,TBP,TBR1,TBX1,TBX15,TBX18,TBX19,TBX2,TBX20,TBX21,TBX3,TBX4,TBX5,Tbx6,TBXT,Tcf12,TCF12,Tcf21,TCF21,TCF3,TCF4,TCF7,TCF7L1,TCF7L2,TCFL5,TEAD1,TEAD2,TEAD3,TEAD4,TEF,TFAP2A,TFAP2B,TFAP2C,TFAP2E,TFAP4,TFAP4::ETV1,TFAP4::FLI1,TFCP2,Tfcp2l1,TFDP1,TFE3,TFEB,TFEC,TGIF1,TGIF2,TGIF2LX,TGIF2LY,THAP1,Thap11,THRA,THRB,TIGD3,TIGD4,TIGD5,TIGD7,TLX2,TLX3,TP53,TP63,TP73,TPRX1,TRPS1,TWIST1,Twist2,UNCX,USF1,USF2,USF3,VAX1,VAX2,Vdr,VENTX,VEZF1,VSX1,VSX2,Wt1,XBP1,Yy1,YY2,ZBED1,ZBED2,ZBED4,ZBED5,ZBTB11,ZBTB12,ZBTB14,ZBTB17,ZBTB18,Zbtb2,ZBTB21,ZBTB24,ZBTB26,ZBTB32,ZBTB33,ZBTB40,ZBTB41,ZBTB47,ZBTB48,ZBTB5,ZBTB6,ZBTB7A,ZBTB7B,ZBTB7C,ZBTB8A,ZBTB8B,ZEB1,ZEB2,ZFP14,ZFP28,ZFP3,Zfp335,ZFP42,ZFP57,Zfp809,Zfp961,ZFTA,Zfx,ZGLP1,ZIC1,Zic1::Zic2,Zic2,Zic3,ZIC4,ZIC5,ZIM3,ZKSCAN1,ZKSCAN3,ZKSCAN4,ZKSCAN5,ZNF121,ZNF131,ZNF134,ZNF135,ZNF136,ZNF140,ZNF142,ZNF143,ZNF146,ZNF148,ZNF157,ZNF16,ZNF175,ZNF18,ZNF184,ZNF189,ZNF20,ZNF211,ZNF213,ZNF214,ZNF215,ZNF226,ZNF234,ZNF24,ZNF250,ZNF251,ZNF257,ZNF260,ZNF263,ZNF274,ZNF275,ZNF281,ZNF282,ZNF286B,ZNF317,ZNF320,ZNF322,ZNF324,ZNF331,ZNF335,ZNF341,ZNF343,ZNF347,ZNF35,ZNF354A,ZNF354C,ZNF362,ZNF367,ZNF382,ZNF384,ZNF395,ZNF407,ZNF410,ZNF416,ZNF417,ZNF418,ZNF423,ZNF43,ZNF436,ZNF449,ZNF454,ZNF460,ZNF470,ZNF471,ZNF490,ZNF493,ZNF497,ZNF500,ZNF510,ZNF518B,ZNF524,ZNF528,ZNF530,ZNF536,ZNF547,ZNF549,ZNF551,ZNF558,ZNF568,ZNF57,ZNF574,ZNF582,ZNF596,ZNF606,ZNF610,ZNF623,ZNF648,ZNF652,ZNF66,ZNF667,ZNF669,ZNF672,ZNF675,ZNF676,ZNF677,ZNF678,ZNF680,ZNF682,ZNF683,ZNF684,ZNF689,ZNF692,ZNF696,ZNF699,ZNF70,ZNF700,ZNF701,ZNF707,ZNF708,ZNF721,ZNF724,ZNF726,ZNF728,ZNF732,ZNF740,ZNF746,ZNF75A,ZNF75D,ZNF76,ZNF766,ZNF768,ZNF770,ZNF772,ZNF773,ZNF775,ZNF780B,ZNF784,ZNF8,ZNF800,ZNF814,ZNF816,ZNF827,ZNF831,ZNF836,ZNF841,ZNF85,ZNF850,ZNF853,ZNF865,ZNF878,ZNF93,ZSCAN16,ZSCAN2,ZSCAN21,ZSCAN25,ZSCAN29,ZSCAN31,ZSCAN4,ZXDA,ZXDB,ZXDC\ labelFields TFName\ longLabel JASPAR CORE 2026 - Predicted Transcription Factor Binding Sites\ maxItems 100000\ motifPwmTable hgFixed.jasparCore2026\ parent jaspar on\ priority 1\ shortLabel JASPAR 2026 TFBS\ showCfg on\ track jaspar2026\ type bigBed 6 +\ visibility pack\ wgEncodeRegDnaseUwK562Peak K562 Pk narrowPeak K562 lymphoblast chronic myeloid leukemia cell line DNaseI Peaks from ENCODE 1 1 255 85 85 255 170 170 1 0 0 regulation 1 color 255,85,85\ longLabel K562 lymphoblast chronic myeloid leukemia cell line DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak on\ shortLabel K562 Pk\ subGroups view=a_Peaks cellType=K562 treatment=n_a tissue=bone_marrow cancer=cancer\ track wgEncodeRegDnaseUwK562Peak\ wgEncodeRegDnaseUwK562Wig K562 Sg bigWig 0 38914.2 K562 lymphoblast chronic myeloid leukemia cell line DNaseI Signal from ENCODE 0 1 255 85 85 255 170 170 0 0 0 regulation 1 color 255,85,85\ longLabel K562 lymphoblast chronic myeloid leukemia cell line DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig on\ priority 1\ shortLabel K562 Sg\ subGroups cellType=K562 treatment=n_a tissue=bone_marrow cancer=cancer\ table wgEncodeRegDnaseUwK562Signal\ track wgEncodeRegDnaseUwK562Wig\ type bigWig 0 38914.2\ lovdShort LOVD Variants < 50 bp + ins bigBed 4 + LOVD: Leiden Open Variation Database, short < 50 bp variants and insertions of any length 0 1 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/lovd/lovd.hg38.short.bb\ group phenDis\ longLabel LOVD: Leiden Open Variation Database, short < 50 bp variants and insertions of any length\ noScoreFilter on\ parent lovdComp\ shortLabel LOVD Variants < 50 bp + ins\ track lovdShort\ urls id="https://varcache.lovd.nl/redirect/$$"\ visibility hide\ mavedb_align_aa MaveDB AA Align bigPsl Reference-Aligned AA Sequences from MaveDB Experiments 3 1 0 0 0 127 127 127 0 0 0 expression 1 bigDataUrl /gbdb/hg38/maveDB/mavedb_aa.bb\ longLabel Reference-Aligned AA Sequences from MaveDB Experiments\ parent mavedb_align_composite\ shortLabel MaveDB AA Align\ track mavedb_align_aa\ mavedb_align_composite MaveDB Alignments bigPsl MaveDB Experiment Sequence Alignments 3 1 0 0 0 127 127 127 0 0 0\
\
\ The DNA subtrack is also configured to highlight base differences from the reference genome. Due to the\ alignment method, this highlighting is currently unavailable for the peptide alignments.\
\
\ Two DNA sequences (for 00000002-a-2 and 00000053-a-2) weren't sufficiently identical for this process to\ find a good alignment; in those cases, the sequences were instead aligned using BLAT's translated alignment flags.\
\ Peptide sequences went solely through the GENCODE-pslMap path.\
\
\
\ Rubin AF, Stone J, Bianchi AH, Capodanno BJ, Da EY, Dias M, Esposito D, Frazer J, Fu Y, Grindstaff\ SB et al.\ \ MaveDB 2024: a curated community database with over seven million variant effects from multiplexed\ functional assays.\ Genome Biol. 2025 Jan 21;26(1):13.\ PMID: 39838450; PMC: PMC11753097\
\ expression 1 compositeTrack on\ hideEmptySubtracks on\ html mavedb_align\ indelDoubleInsert on\ indelPolyA on\ indelQueryInsert on\ longLabel MaveDB Experiment Sequence Alignments\ parent mavedb\ priority 1\ shortLabel MaveDB Alignments\ track mavedb_align_composite\ type bigPsl\ visibility pack\ MaxCounts_Fwd Max counts of CAGE reads (fwd) bigWig Max counts of CAGE reads forward 2 1 255 0 0 255 127 127 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/fantom5/ctssMaxCounts.fwd.bw\ color 255,0,0\ dataVersion FANTOM5 reprocessed7\ longLabel Max counts of CAGE reads forward\ parent Max_counts_multiwig\ shortLabel Max counts of CAGE reads (fwd)\ subGroups category=max strand=forward\ track MaxCounts_Fwd\ type bigWig\ gnomADPextmean_proportion Mean Proportion bigWig 0 1 gnomAD pext Mean Proportion 2 1 66 139 202 160 197 228 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/exp_prop_mean.bw\ color 66,139,202\ longLabel gnomAD pext Mean Proportion\ parent gnomadPext on\ priority 1\ shortLabel Mean Proportion\ track gnomADPextmean_proportion\ visibility full\ mexbb Mexico Biobank, 6k Array vcfTabix Phased Variants: Mexico Biobank 6k Array 3 1 0 0 0 127 127 127 0 0 0\ This tracks contains variants of individual genotypes, usually phased, from the projects\ Human Diversity Genome Project, Simons Genome Diversity Project, gnomad's HGDP+1000 Genomes callset,\ and the Mexico Biobank.\ The original release of 1000 Genomes has its own, separate track.\ Projects where the released variants are not phased can be found in the container track "SNV Frequencies".\
\ \\ Available on hg19 and hg38:
\\ Available only on hg38:
\\ Full haplotype display:\ In "pack" mode, this track sorts the haplotypes. This can be\ useful for determining the similarity between the samples and inferring\ inheritance at a particular locus.\ Each sample's phased and/or homozygous genotypes are split into haplotypes,\ clustered by similarity around a central variant (in pink), and sorted for\ display by their position in the clustering tree. Click a variant to center on it.\ The tree (as space allows) is drawn in the label area next to the track image.\ Leaf clusters, in which all haplotypes are identical (at least for the variants\ used in clustering), are colored purple. \
\\ For a full description of how the display works, please see our \ Haplotype Display help page.\ \
\ MXB: Allele frequencies by geographical state and ancestry are available via\ the MexVar platform.\ Raw genotype data are available under controlled access at the\ EGA (Study: EGAS00001005797; Dataset: EGAD00010002361). For the VCFs, email\ andres.moreno@cinvestav.mx.\
\ \\ SGDP: The version used was\ https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/,\ merged with bcftools and lifted to hg38 with CrossMap. \
\ \\ MXB: We thank the Center for Research and Advanced Studies (Cinvestav) of Mexico for\ generating and providing the frequency data, the National Institute of Medical\ Sciences and Nutrition (INCMNSZ) for DNA extraction, and the Ministry of Health\ together with the National Institute of Public Health (INSP) for the design and\ implementation of the National Health Survey 2000 (ENSA 2000). We also thank\ the ENSA-Genomics Consortium for their contributions to sample collection and\ data processing that made possible the construction of the MXB genomic\ resource.\
\\ SGDP: This project was funded by the Simons Foundation. Thanks to David Reich and Swapan \ Mallick for help with importing the data.\
\ \\ Barberena-Jonas C, Medina-Muñoz SG, Cedillo-Castelán V, Sepúlveda-Morales T,\ Gonzaga-Jáuregui C, ENSA Genomics Consortium, García-García L, Ioannidis AG,\ Moreno-Estrada A.\ \ Clinical genetic variation across Hispanic populations in the Mexican Biobank.\ Nat Med. 2026 Jan 21;.\ DOI: 10.1038/s41591-025-04100-z; PMID: 41566040\
\ \\ Sohail M, Moreno-Estrada A.\ \ The Mexican Biobank Project promotes genetic discovery, inclusive science and local capacity\ building.\ Dis Model Mech. 2024 Jan 1;17(1).\ PMID: 38299665; PMC: PMC10855211\
\ \\ Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Corés CD, Barberena-Jonas C, Medina-Muñoz SG,\ Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L et al.\ \ Mexican Biobank advances population and medical genomics of diverse ancestries.\ Nature. 2023 Oct;622(7984):775-783.\ PMID: 37821706; PMC: PMC10600006\
\ \\ Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J\ et al.\ \ Insights into human genetic variation and population history from 929 diverse genomes.\ Science. 2020 Mar 20;367(6484).\ PMID: 32193295; PMC: PMC7115999\
\ \\ Koenig Z, Yohannes MT, Nkambule LL, Zhao X, Goodrich JK, Kim HA, Wilson MW, Tiao G, Hao SP, Sahakian\ N et al.\ \ A harmonized public resource of deeply sequenced diverse human genomes.\ Genome Res. 2024 Jun 25;34(5):796-809.\ PMID: 38749656; PMC: PMC11216312\
\ \\ Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S,\ Tandon A et al.\ \ The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.\ Nature. 2016 Oct 13;538(7624):201-206.\ PMID: 27654912; PMC: PMC5161557\
\ \ varRep 1 bigDataUrl /gbdb/hg38/phasedVars/mexbb/MXBv2.vcf.gz\ dataVersion Nov 2025 (hg38 lift)\ hapClusterEnabled true\ html phasedVars.html\ longLabel Phased Variants: Mexico Biobank 6k Array\ parent phasedVars on\ priority 1\ shortLabel Mexico Biobank, 6k Array\ tableBrowser off\ track mexbb\ type vcfTabix\ visibility pack\ mitoMapVars MITOMAP Variants bigBed 9 + 11 MITOMAP Control and Coding Variants 0 1 0 0 0 127 127 127 0 0 2 chrM,chrMT, https://www.mitomap.org/foswiki/bin/view/MITOMAP/$<_varType> phenDis 1 bigDataUrl /gbdb/hg38/bbi/mitoMapVars.bb\ exonNumbers off\ group phenDis\ longLabel MITOMAP Control and Coding Variants\ mouseOverField _mouseOver\ parent mitoMap on\ priority 1\ shortLabel MITOMAP Variants\ track mitoMapVars\ type bigBed 9 + 11\ url https://www.mitomap.org/foswiki/bin/view/MITOMAP/$<_varType>\ urlLabel MITOMAP link\ mprabase MPRA Base bigBed 9 + 11 MPRAs: MPRA Base Enhancer Elements 3 1 0 0 0 127 127 127 0 0 0\ Massively Parallel Reporter Assays (MPRAs) and related methods such as STARR-seq\ enable quantitative testing of thousands of candidate regulatory DNA sequences in\ parallel by linking each sequence to a reporter gene and measuring transcriptional\ output using sequencing.\
\ \\ The MPRA Base track shows 40,938 experimentally tested cis-regulatory elements\ curated from the MPRA Base\ database\ (Zhao et al., 2023),\ drawn from MPRA, STARR-seq, and related reporter assay experiments.\ The database integrates data from multiple studies, assay platforms (lentiMPRA,\ plasmidMPRA, STARR-seq, CRE-seq, and others), and cell types while preserving\ experiment-level resolution. Only elements derived from genomic fragments that can\ be mapped to the reference genome are included; synthetic or designed oligonucleotide\ libraries without genomic coordinates are excluded.\
\\ The track is a curated union of study-specific libraries rather than a uniform\ genome-wide enhancer catalog: each contributing study targeted a distinct set of\ candidate regions, including HepG2 liver-enhancer panels, melanoma GWAS variants,\ human/mouse pluripotent TSSs, and ASD-associated promoter variants. Each item\ represents one experimental measurement, not a full enhancer; longer regulatory\ elements may be represented by multiple adjacent tiles. Item width corresponds\ to the assayed DNA fragment for tile-based studies (most items, 144–200 bp;\ some Klein et al., 2020 elements 354–678 bp) but collapses to a\ single base for variant-centered studies that mark the SNP location rather than\ the surrounding tested window (Choi et al., 2020).\
\\ Note on cell lines: The cell line shown for each element is the reporter\ cell line in which the genomic fragment was assayed. Most rows test human DNA in\ human cells; the exception is Mattioli et al., 2020, where mESC rows assay the\ mouse orthologous sequence in mouse cells, with hg38 coordinates derived from the\ human ortholog by liftOver.\
\\ The biological context of each cell line is summarized below:\
\| Cell line | Biological context |
|---|---|
| HepG2 | Hepatocellular carcinoma; liver enhancer studies |
| HUES64 | Human embryonic stem cells; pluripotent |
| mESC | Mouse embryonic stem cells; pluripotent |
| NPC | H1-derived neural progenitor cells; developing brain |
| HEK293FT | Embryonic kidney; high-transfection-efficiency reference |
| UACC903 | Melanoma cell line |
\ Each item represents a genomic fragment tested within a specific experiment, defined\ as a unique combination of cell line, assay type, and publication (PMID). The same\ genomic region may appear multiple times if tested in different experiments.\
\ \\ Items are colored by percentile rank of the mean raw activity score within each experiment:\
\\ The mouse-over shows the cell line, assay type, raw activity score, percentile rank,\ and citation for each element.\
\ \\ The details page additionally shows the variant allele type for each\ row (reference or alternate for a row that is part of a\ variant comparison, NA for a standard enhancer element that is not a\ variant test) and the tested oligo sequence — the exact DNA\ fragment assayed in the MPRA experiment.\
\ \\ For most studies in this track, the raw score is the log2 ratio of reporter\ RNA to input DNA from the source experiment. A score of 0 means the fragment produced\ RNA in proportion to the input plasmid copies (no measurable activity above baseline),\ positive scores indicate the fragment drove the reporter above baseline (enhancer-like\ activity in the assay), and negative scores indicate sub-baseline output (treated as\ inactive, not as validated transcriptional repression). Linear fold change relative to\ baseline is approximately 2raw_score — for example, a raw score of 0.18\ corresponds to roughly 1.13× baseline output, 1.0 to 2×, and 2.0 to 4×.\
\\ Two studies use a different scale: Mattioli et al., 2020 and Koesterich\ et al., 2023 report the MPRAnalyze induced-transcription rate\ (α), which is a positive-only quantity not directly convertible to a fold\ change. As noted in the Methods section, scoring methodology and the threshold\ used to call an element "active" differ between studies, so the percentile rank\ reflects within-experiment ranking only and does not by itself indicate the\ absolute strength of an element.\
\ \\ Within each experiment, replicate measurements for the same genomic fragment were\ aggregated by computing the mean raw activity score, yielding 40,938 unique\ experiment-level genomic elements.\
\ \\ Elements are ranked by mean raw activity score independently within each experiment,\ and a percentile rank (0–100) is computed per experiment to avoid cross-study\ distortions caused by differing assay dynamic ranges.\
\ \\ Scoring methodology and the threshold used to call an element "active"\ differ between studies, so percentile-rank comparisons across experiments are\ approximate. Lower scores indicate that the fragment did not measurably activate\ transcription in the assay, rather than that it actively represses transcription.\ For any element of interest, users should consult the source publication for the\ original significance and effect-size calls.\
\ \\ Original genomic coordinates from the source studies (mostly hg19, with some\ mm9 and mm10) were lifted to hg38 by the MPRA Base pipeline using the UCSC\ liftOver tool.\
\ \\ The following table lists the experiments represented in this track.\
\ \| PMID | \Author | \Year | \Lab | \Cell type | \Assay | \Elements | \
|---|---|---|---|---|---|---|
| 27831498 | Inoue et al. | 2017 | Shendure Lab | HepG2 | lentiMPRA | 2,241 |
| 30045748 | Klein et al. | 2018 | Shendure Lab | HepG2 | STARR-seq | 6,728 |
| 32483191 | Choi et al. | 2020 | Brown Lab | HEK293FT | lentiMPRA | 840 |
| 32483191 | Choi et al. | 2020 | Brown Lab | UACC903 | lentiMPRA | 840 |
| 32819422 | Mattioli et al. | 2020 | Mele Lab | HUES64 | plasmidMPRA | 6,954 |
| 32819422 | Mattioli et al. | 2020 | Mele Lab | mESC | plasmidMPRA | 6,954 |
| 33046894 | Klein et al. | 2020 | Shendure Lab | HepG2 | lentiMPRA | 8,116 |
| 33046894 | Klein et al. | 2020 | Shendure Lab | HepG2 | plasmidMPRA | 2,228 |
| 33046894 | Klein et al. | 2020 | Shendure Lab | HepG2 | STARR-seq | 2,230 |
| 36834916 | Koesterich et al. | 2023 | Kreimer Lab | NPC | lentiMPRA | 3,807 |
\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator\ and exported from there to spreadsheet or tab-sep tables.\ From scripts, the data can be accessed through our\ API, track=mprabase.\
\\ For automated download and analysis, the genome annotation is stored in a bigBed\ file that can be downloaded from\ our download server.\ The file for this track is called mprabase.bb. Individual\ regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a\ precompiled binary for your system. Instructions for downloading source code and\ binaries can be found\ here.\ The tool can also be used to obtain features within a given range, e.g.\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/mpra/mprabase/mprabase.bb -chrom=chr21 -start=0 -end=100000000 stdout\
\\ The original data can be downloaded from the\ MPRA Base web application.\
\ \\ Thanks to Varda Singhal, Jianyu Zhao, and the\ Ahituv Lab\ at the University of California San Francisco for creating and curating MPRA Base and for creating this track.\
\ \\ Choi J, Zhang T, Vu A, Ablain J, Makowski MM, Colli LM, Xu M, Hennessey RC, Yin J, Rothschild H\ et al.\ \ Massively parallel reporter assays of melanoma risk variants identify MX2 as a gene promoting\ melanoma.\ Nat Commun. 2020 Jun 1;11(1):2718.\ PMID: 32483191; PMC: PMC7264232\
\ \\ Inoue F, Kircher M, Martin B, Cooper GM, Witten DM, McManus MT, Ahituv N, Shendure J.\ \ A systematic comparison reveals substantial differences in chromosomal versus episomal encoding of\ enhancer activity.\ Genome Res. 2017 Jan;27(1):38-52.\ PMID: 27831498; PMC: PMC5204343\
\ \\ Klein JC, Keith A, Agarwal V, Durham T, Shendure J.\ \ Functional characterization of enhancer evolution in the primate lineage.\ Genome Biol. 2018 Jul 25;19(1):99.\ PMID: 30045748; PMC: PMC6060477\
\ \\ Klein JC, Agarwal V, Inoue F, Keith A, Martin B, Kircher M, Ahituv N, Shendure J.\ \ A systematic evaluation of the design and context dependencies of massively parallel reporter\ assays.\ Nat Methods. 2020 Nov;17(11):1083-1091.\ PMID: 33046894; PMC: PMC7727316\
\ \\ Koesterich J, An JY, Inoue F, Sohota A, Ahituv N, Sanders SJ, Kreimer A.\ \ Characterization of De Novo Promoter Variants in Autism Spectrum Disorder with Massively Parallel\ Reporter Assays.\ Int J Mol Sci. 2023 Feb 9;24(4).\ PMID: 36834916; PMC: PMC9959321\
\ \\ Mattioli K, Oliveros W, Gerhardinger C, Andergassen D, Maass PG, Rinn JL, Melé M.\ \ Cis and trans effects differentially contribute to the evolution of promoters and enhancers.\ Genome Biol. 2020 Aug 20;21(1):210.\ PMID: 32819422; PMC: PMC7439725\
\ \\ Zhao J, Baltoumas FA, Konnaris MA, Mouratidis I, Liu Z, Sims J, Agarwal V, Pavlopoulos GA,\ Georgakopoulos-Soares I, Ahituv N.\ \ MPRAbase: A Massively Parallel Reporter Assay Database.\ bioRxiv. 2023 Nov 22;.\ PMID: 38045264; PMC: PMC10690217\
\ \ regulation 1 bigDataUrl /gbdb/hg38/mpra/mprabase/mprabase.bb\ dataVersion MPRA Base 2026-05-27 refresh\ defaultLabelFields name\ filter.percentile_rank 0:100\ filterByRange.percentile_rank on\ filterLabel.percentile_rank Filter by activity percentile rank (within experiment)\ filterLimits.percentile_rank 0:100\ filterValues.assay lentiMPRA (LM),plasmidMPRA (PM),STARR-seq (ST)\ filterValues.cell_line HepG2,HUES64,mESC,NPC,HEK293FT,UACC903\ filterValues.variant_type alternate,NA,reference\ itemRgb on\ labelFields name,variant_type,cell_line,assay,author_lab\ longLabel MPRAs: MPRA Base Enhancer Elements\ mouseOver Element: $name\ This track shows multiple alignments of 90 human genomes generated by the Minigraph-Cactus\ pangenome pipeline, which creates pangenomes directly from whole-genome alignments. This method\ builds graphs containing all forms of genetic variation while allowing use of current mapping and\ genotyping tools.\
\ \\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ size of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ Note that excluding species from the pairwise display does not alter the\ the conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. The following\ conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment. The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\ \\ The MAF was obtained from the HPRC v1.0 minigraph-cactus HAL file (renamed\ to replace all "." characters in sample names with "#" using\ halRenameGenomes) using cactus v2.6.4 as follows.\
\ cactus-hal2maf ./js ./hprc-v1.0-mc-grch38.h\ al hprc-v1.0-mc-grch38.maf.gz --noAncestors --refGenome GRCh38\ --filterGapCausingDupes --chunkSize 100000 --batchCores 96 --batchCount 1\ 0 --noAncestors --batchParallelTaf 32 --batchSystem slurm --logFile\ hprc-v1.0-mc-grch38.maf.gz.log\ \ zcat hprc-v1.0-mc-grch38.maf.gz | mafDuplicateFilter -m - -k | bgzip >\ hprc-v1.0-mc-grch38-single-copy.maf.gz\ \ \
\ Thank you to Glenn Hickey for providing the HAL file from the HPRC project.\
\ \\ Liao WW, Asri M, Ebler J, Doerr D, Haukness M, Hickey G, Lu S, Lucas JK, Monlong J, Abel HJ et\ al.\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ DOI: 10.1038/s41586-023-05896-x; PMID: 37165242; PMC: PMC10172123\
\ \\ Hickey G, Monlong J, Ebler J, Novak AM, Eizenga JM, Gao Y, Human Pangenome Reference Consortium,\ Marschall T, Li H, Paten B.\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nat Biotechnol. 2023 May 10;.\ DOI: 10.1038/s41587-023-01793-w; PMID: 37165083; PMC: PMC10638906\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q, Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ DOI: 10.1038/s41586-020-2871-y; PMID: 33177663; PMC: PMC7673649\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ DOI: 10.1101/gr.123356.111;\ PMID: 21665927; PMC: PMC3166836\
\ hprc 1 altColor 0,90,10\ color 0, 10, 100\ irows on\ itemFirstCharCase noChange\ longLabel Multiple Alignment on 90 human genome assemblies\ mafDot on\ noInherit on\ parent consHprc90wayViewalign\ sGroup_Afr_Carib_Barabdos GCA_018466835.1 GCA_018466845.1 GCA_018466855.1 GCA_018466985.1 GCA_018467005.1 GCA_018467015.1 GCA_018467155.1 GCA_018467165.1 GCA_018505825.1 GCA_018505855.1 GCA_018505865.1 GCA_018506125.1 GCA_018852585.1 GCA_018852595.1\ sGroup_African_SW_USA GCA_018504625.1 GCA_018504635.1\ sGroup_Columbia_Medellin GCA_018469405.1 GCA_018469665.1 GCA_018469675.1 GCA_018469685.1 GCA_018469695.1 GCA_018469705.1 GCA_018469865.1 GCA_018469965.1\ sGroup_Esan_Nigeria GCA_018469415.1 GCA_018469425.1\ sGroup_Gambian GCA_018469875.1 GCA_018469925.1 GCA_018469935.1 GCA_018469945.1 GCA_018469955.1 GCA_018470425.1 GCA_018470435.1 GCA_018470445.1 GCA_018470455.1 GCA_018470465.1 GCA_018473295.1 GCA_018473315.1 GCA_018503575.1 GCA_018503585.1 GCA_018504065.1 GCA_018504075.1\ sGroup_HAPMAP GCA_018504655.1 GCA_018504665.1\ sGroup_Han_SoChina GCA_018471515.1 GCA_018472565.1 GCA_018472575.1 GCA_018472585.1 GCA_018472595.1 GCA_018472605.1\ sGroup_Mende_Sierra_Leone GCA_018472825.1 GCA_018472835.1 GCA_018472855.1 GCA_018473305.1 GCA_018503245.1 GCA_018503525.1 GCA_018506155.1 GCA_018506165.1\ sGroup_Peru_Lima GCA_018471525.1 GCA_018471535.1 GCA_018471545.1 GCA_018471555.1 GCA_018472695.1 GCA_018472705.1 GCA_018472845.1 GCA_018472865.1\ sGroup_Puerto_Rico GCA_018471065.1 GCA_018471075.1 GCA_018471085.1 GCA_018471095.1 GCA_018471105.1 GCA_018471345.1 GCA_018472685.1 GCA_018472715.1 GCA_018472725.1 GCA_018472765.1 GCA_018504045.1 GCA_018504365.1 GCA_018504375.1 GCA_018504645.1 GCA_018506955.1 GCA_018506975.1\ sGroup_Punjabo_Pakis GCA_018505835.1 GCA_018505845.1\ sGroup_T2T hs1\ sGroup_Vietnam_Kinh GCA_018504055.1 GCA_018504085.1\ sGroup_Yoruba_Nigeria GCA_018503255.1 GCA_018503285.1\ shortLabel Multiple Alignment\ speciesCodonDefault hg38\ speciesGroups T2T HAPMAP Yoruba_Nigeria Esan_Nigeria Gambian Mende_Sierra_Leone Afr_Carib_Barabdos African_SW_USA Puerto_Rico Peru_Lima Columbia_Medellin Han_SoChina Vietnam_Kinh Punjabo_Pakis\ speciesLabels GCA_018466835.1="HG02257.pat" GCA_018466845.1="HG02257.mat" GCA_018466855.1="HG02559.pat" GCA_018466985.1="HG02559.mat" GCA_018467005.1="HG02486.pat" GCA_018467015.1="HG02486.mat" GCA_018467155.1="HG01891.mat" GCA_018467165.1="HG01891.pat" GCA_018469405.1="HG01258.mat" GCA_018469415.1="HG03516.pat" GCA_018469425.1="HG03516.mat" GCA_018469665.1="HG01123.mat" GCA_018469675.1="HG01258.pat" GCA_018469685.1="HG01361.mat" GCA_018469695.1="HG01123.pat" GCA_018469705.1="HG01361.pat" GCA_018469865.1="HG01358.mat" GCA_018469875.1="HG02622.mat" GCA_018469925.1="HG02622.pat" GCA_018469935.1="HG02717.mat" GCA_018469945.1="HG02630.pat" GCA_018469955.1="HG02630.mat" GCA_018469965.1="HG01358.pat" GCA_018470425.1="HG02717.pat" GCA_018470435.1="HG02572.pat" GCA_018470445.1="HG02572.mat" GCA_018470455.1="HG02886.mat" GCA_018470465.1="HG02886.pat" GCA_018471065.1="HG01175.pat" GCA_018471075.1="HG01106.pat" GCA_018471085.1="HG01175.mat" GCA_018471095.1="HG00741.mat" GCA_018471105.1="HG00741.pat" GCA_018471345.1="HG01106.mat" GCA_018471515.1="HG00438.mat" GCA_018471525.1="HG02148.pat" GCA_018471535.1="HG02148.mat" GCA_018471545.1="HG01952.mat" GCA_018471555.1="HG01952.pat" GCA_018472565.1="HG00673.mat" GCA_018472575.1="HG00621.pat" GCA_018472585.1="HG00673.pat" GCA_018472595.1="HG00438.pat" GCA_018472605.1="HG00621.mat" GCA_018472685.1="HG01071.mat" GCA_018472695.1="HG01928.mat" GCA_018472705.1="HG01928.pat" GCA_018472715.1="HG00735.pat" GCA_018472725.1="HG01071.pat" GCA_018472765.1="HG00735.mat" GCA_018472825.1="HG03579.mat" GCA_018472835.1="HG03579.pat" GCA_018472845.1="HG01978.pat" GCA_018472855.1="HG03453.mat" GCA_018472865.1="HG01978.mat" GCA_018473295.1="HG03540.mat" GCA_018473305.1="HG03453.pat" GCA_018473315.1="HG03540.pat" GCA_018503245.1="HG03486.pat" GCA_018503255.1="NA18906.mat" GCA_018503285.1="NA18906.pat" GCA_018503525.1="HG03486.mat" GCA_018503575.1="HG02818.pat" GCA_018503585.1="HG02818.mat" GCA_018504045.1="HG01243.pat" GCA_018504055.1="HG02080.pat" GCA_018504065.1="HG02723.mat" GCA_018504075.1="HG02723.pat" GCA_018504085.1="HG02080.mat" GCA_018504365.1="HG01109.mat" GCA_018504375.1="HG01243.mat" GCA_018504625.1="NA20129.pat" GCA_018504635.1="NA20129.mat" GCA_018504645.1="HG01109.pat" GCA_018504655.1="NA21309.mat" GCA_018504665.1="NA21309.pat" GCA_018505825.1="HG02109.mat" GCA_018505835.1="HG03492.pat" GCA_018505845.1="HG03492.mat" GCA_018505855.1="HG02055.pat" GCA_018505865.1="HG02109.pat" GCA_018506125.1="HG02055.mat" GCA_018506155.1="HG03098.pat" GCA_018506165.1="HG03098.mat" GCA_018506955.1="HG00733.pat" GCA_018506975.1="HG00733.mat" GCA_018852585.1="HG02145.mat" GCA_018852595.1="HG02145.pat" hs1="T2T-CHM13v2.0"\ subGroups view=align\ summary hprc90waySummary\ track hprc90way\ treeImage phylo/hprc_90way.png\ type wigMaf 0.0 1.0\ viewUi on\ cons100wayViewalign Multiz Alignments bed 4 UCSC 100 Vertebrates - 100 vertebrate genomes aligned with MultiZ by the UCSC Browser Group 3 1 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel UCSC 100 Vertebrates - 100 vertebrate genomes aligned with MultiZ by the UCSC Browser Group\ parent cons100way\ shortLabel Multiz Alignments\ track cons100wayViewalign\ view align\ viewUi on\ visibility pack\ caddA Mutation: A bigWig CADD 1.6 Score: Mutation is A 1 1 100 130 160 177 192 207 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/cadd/a.bw\ longLabel CADD 1.6 Score: Mutation is A\ maxHeightPixels 128:20:8\ parent cadd on\ shortLabel Mutation: A\ track caddA\ type bigWig\ viewLimits 10:50\ viewLimitsMax 0:100\ visibility dense\ promoterAiA Mutation: A bigWig PromoterAI: Mutation is A 1 1 200 0 0 0 0 200 0 0 0 phenDis 0 altColor 0,0,200\ alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/_promoterAi/a.bw\ color 200,0,0\ longLabel PromoterAI: Mutation is A\ maxHeightPixels 128:40:8\ maxWindowToDraw 10000000\ maxWindowToQuery 500000\ mouseOverFunction noAverage\ parent promoterAi on\ shortLabel Mutation: A\ track promoterAiA\ type bigWig\ viewLimits -1:1\ viewLimitsMax -1:1\ visibility dense\ cadd1_7_A Mutation: A bigWig CADD 1.7 Score: Mutation is A 1 1 100 130 160 177 192 207 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/cadd1.7/a.bw\ longLabel CADD 1.7 Score: Mutation is A\ maxHeightPixels 128:20:8\ parent cadd1_7 on\ setColorWith /gbdb/hg38/cadd1.7/a.color.bb\ shortLabel Mutation: A\ track cadd1_7_A\ type bigWig\ viewLimits 10:50\ viewLimitsMax 0:100\ visibility dense\ revelA Mutation: A bigWig REVEL: Mutation is A 1 1 150 80 200 202 167 227 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/revel/a.bw\ longLabel REVEL: Mutation is A\ maxHeightPixels 128:20:8\ maxWindowToDraw 10000000\ maxWindowToQuery 500000\ mouseOverFunction noAverage\ parent revel on\ setColorWith /gbdb/hg38/revel/a.color.bb\ shortLabel Mutation: A\ track revelA\ type bigWig\ viewLimits 0:1.0\ viewLimitsMax 0:1.0\ visibility dense\ alphaMissense_A Mutation: A bigWig AlphaMissense Score: Mutation is A 1 1 100 130 160 177 192 207 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/alphaMissense/a.bw\ longLabel AlphaMissense Score: Mutation is A\ maxHeightPixels 128:20:8\ parent alphaMissense on\ setColorWith /gbdb/hg38/alphaMissense/a.color.bb\ shortLabel Mutation: A\ track alphaMissense_A\ type bigWig\ viewLimits 0:1\ visibility dense\ mutScoreA Mutation: A bigWig MutScore: Mutation is A 2 1 50 80 200 152 167 227 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/mutscore/mutscoreA.bw\ longLabel MutScore: Mutation is A\ maxHeightPixels 128:20:8\ maxWindowToDraw 10000000\ maxWindowToQuery 500000\ mouseOverFunction noAverage\ parent mutScore on\ shortLabel Mutation: A\ track mutScoreA\ type bigWig\ viewLimits 0:1.0\ viewLimitsMax 0:1.0\ visibility full\ neuronMerged Neurons Merged bigWig Methylation Atlas: Neurons Merged Samples 2 1 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/neuronMerged.bw\ color 138,43,226\ longLabel Methylation Atlas: Neurons Merged Samples\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals on\ priority 1\ shortLabel Neurons Merged\ subGroups cellType=Neuron dataType=Merged\ track neuronMerged\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ omimContainer OMIM Online Mendelian Inheritance in Man 0 1 0 0 0 127 127 127 0 0 0\ OMIM is a compendium of human genes and genetic phenotypes. The full-text, \ referenced overviews in OMIM contain information on all known Mendelian \ disorders and over 12,000 genes. OMIM is authored and edited at the McKusick-Nathans \ Institute of Genetic Medicine, Johns Hopkins University School of Medicine, under \ the direction of Dr. Ada Hamosh. This database was initiated in the early 1960s \ by Dr. Victor A. McKusick as a catalog of Mendelian traits and disorders, \ entitled Mendelian Inheritance in Man (MIM).
\ \\
The OMIM data are separated into three separate tracks:
\
\
OMIM Alellic Variant Phenotypes (OMIM Alleles) - Variants in the OMIM \
database that have associated dbSNP identifiers.
\
\
OMIM Gene Phenotypes (OMIM Genes) - The genomic positions of gene \
entries in the OMIM database. The coloring indicates the associated OMIM phenotype map key.
\
\
OMIM Cytogenetic Loci Phenotypes: Gene Unknown (OMIM Cyto Loci) - Regions \
known to be associated with a phenotype, but for which no specific gene is known \
to be causative. This track also includes known multi-gene syndromes.
\
\
Clicking into the individual tracks provides additional information including display conventions.\
NOTE:
\
OMIM is intended for use primarily by physicians and other\
professionals concerned with genetic disorders, by genetics researchers, and\
by advanced students in science and medicine. While the OMIM database is\
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions. Further, please be\
sure to click through to omim.org for the very latest, as they are continually \
updating data.
NOTE ABOUT DOWNLOADS:
\
OMIM is the property \
of Johns Hopkins University and is not available for download or mirroring \
by any third party without their permission. Please see \
OMIM\
for downloads.
OMIM is a compendium of human genes and genetic phenotypes. The full-text,\ referenced overviews in OMIM contain information on all known Mendelian\ disorders and over 12,000 genes. OMIM is authored and edited at the\ McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University\ School of Medicine, under the direction of Dr. Ada Hamosh. This database\ was initiated in the early 1960s by Dr. Victor A. McKusick as a catalog\ of Mendelian traits and disorders, entitled Mendelian Inheritance\ in Man (MIM).\
\ \\ The OMIM data are separated into three separate tracks:\
\ \OMIM Alellic Variant Phenotypes (OMIM Alleles)\
Variants in the OMIM database that have associated \
dbSNP identifiers.\
\
OMIM Gene Phenotypes (OMIM Genes)\
The genomic positions of gene entries in the OMIM \
database. The coloring indicates the associated OMIM phenotype map key.\
OMIM Cytogenetic Loci Phenotypes - Gene Unknown (OMIM Cyto Loci)\
Regions known to be associated with a phenotype, \
but for which no specific gene is known to be causative. This track \
also includes known multi-gene syndromes.\
\ This track shows the allelic variants in the Online Mendelian Inheritance in Man\ (OMIM) database that have associated\ dbSNP identifiers.\
\ \Genomic positions of OMIM allelic variants are marked by solid blocks, which appear\ as tick marks when zoomed out. \
The details page for each variant displays the allelic variant description, the amino\ acid replacement, and the associated\ dbSNP and/or\ ClinVar identifiers with links to the\ variant's details at those resources.\
\The descriptions of OMIM entries are shown on the main browser display when Full display\ mode is chosen. In Pack mode, the descriptions are shown when mousing over each entry.\
\ \\ This track was constructed as follows: \
\ Because OMIM has only allowed Data queries within individual chromosomes, no download files are\ available from the Genome Browser. Full genome datasets can be downloaded directly from the\ OMIM Downloads page.\ All genome-wide downloads are freely available from OMIM after registration.
\\ If you need the OMIM data in exactly the format of the UCSC Genome Browser,\ for example if you are running a UCSC Genome Browser local installation (a partial "mirror"),\ please create a user account on omim.org and contact OMIM via\ https://omim.org/contact. Send them your OMIM\ account name and request access to the UCSC Genome Browser 'entitlement'. They will\ then grant you access to a MySQL/MariaDB data dump that contains all UCSC\ Genome Browser OMIM tables.
\\ UCSC offers queries within chromosomes from\ Table Browser that include a variety\ of filtering options and cross-referencing other datasets using our\ Data Integrator tool.\ UCSC also has an API\ that can be used to retrieve data in JSON format from a particular chromosome range.
\\ Please refer to our searchable\ mailing list archives\ for more questions and example queries, or our\ Data Access FAQ\ for more information.
\ \\ Thanks to OMIM and NCBI for the use of their data. This track was constructed by Fan Hsu,\ Robert Kuhn, and Brooke Rhead of the UCSC Genome Bioinformatics Group.
\ \\ Amberger J, Bocchini CA, Scott AF, Hamosh A.\ McKusick's Online Mendelian Inheritance in Man (OMIM).\ Nucleic Acids Res. 2009 Jan;37(Database issue):D793-6.\ PMID: 18842627; PMC: PMC2686440\
\ \\ Hamosh A, Scott AF, Amberger JS, Bocchini CA, McKusick VA.\ \ Online Mendelian Inheritance in Man (OMIM), a knowledgebase of human genes and genetic\ disorders.\ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D514-7.\ PMID: 15608251; PMC: PMC539987\
\ phenDis 1 color 0, 80, 0\ hgsid on\ longLabel OMIM Allelic Variant Phenotypes\ noGenomeReason Distribution restrictions by OMIM. See the track documentation for details. You can download the complete OMIM dataset for free from omim.org\ parent omimContainer\ priority 1\ shortLabel OMIM Alleles\ tableBrowser noGenome omimAv omimAvRepl\ track omimAvSnp\ type bed 4\ url http://www.omim.org/entry/\ visibility dense\ panelAppGenes PanelApp GE Genes bigBed 9 + Genomics England PanelApp Genes 3 1 0 0 0 127 127 127 0 0 0 https://panelapp.genomicsengland.co.uk/panels/$\ This container track helps call out sections of the genome that often cause problems or\ confusion when working with the genome. The hg19 genome has a track with the same name, but with\ more subtracks, as the GeT-RM and Genome-in-a-Bottle artifact variants do not exist \ for hg38.\ \
\ The Problematic Regions track contains the following subtracks:\
\ The Highly Reproducible Regions track highlights regions and variants\ from eight samples that can be used to assess variant detection pipelines. The\ "Highly Reproducible Regions" subtrack comprises the intersection of the reproducible\ regions across all eight samples, while the "Variants" subtracks contain the reproducible\ variants from each assayed sample. Both tracks contain data from the following samples:\
\The Genome in a Bottle (GIAB) Problematic Regions tracks provide stratifications of the\ genome to evaluate variant calls in complex regions. It is designed for use with Global Alliance\ for Genomic Health (GA4GH) benchmarking tools like\ hap.py\ and includes regions with low complexity, segmental duplications, functional regions,\ and difficult-to-sequence areas. Developed in collaboration with GA4GH, the\ Genome in a Bottle (GIAB) consortium, and the\ Telomere-to-Telomere Consortium (T2T), the dataset aims to standardize the\ analysis of genetic variation by offering pre-defined BED files for stratifying true and false\ positives in genomic studies, facilitating accurate assessments in complex areas of the genome.
\ \\ The creation of the GIAB Problematic Regions tracks involves using a pipeline and configuration to\ generate stratification BED files that categorize genomic regions based on specific challenges,\ such as low complexity or difficult mapping, to facilitate accurate benchmarking of variant calls.\ For more information on the pipeline and configuration used, please visit the following webpage:\ \ https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/genome-stratifications/v3.5/README.md.\ If you have questions or comments, please write to Justin Zook (jzook@nist.gov).
\ \\ The Panmask Easy 151b Regions subtrack contains a set of sample-agnostic easy regions where\ short-read variant calling reaches high accuracy. Easy regions are derived for variant filtration\ agnostic to individual samples. They are genomic intervals where general variant callers achieve\ high accuracy without sophisticated filtering.
\\ A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that\ could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples\ were inspected to exclude regions with excessively high or low coverage or those enriched with\ low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure\ are referred to as pm151:lenient, where "pm" stands for panmask. In addition, low\ complexity regions identified by SDUST were removed.
\The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and\ other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38,\ 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter\ variant calls for clinical or research human samples. Like the HighRepro track in this container\ (see above), it shows regions that are easy to sequence, not those that are problematic. The data\ was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.
\ \\ Each track contains a set of regions of varying length with no special configuration options. \ The UCSC Unusual Regions track has a mouse-over description, all other tracks have at most\ a name field, which can be shown in pack mode. The tracks are usually kept in dense mode.\
\ \\ The Hide empty subtracks control hides subtracks with no data in the browser window.\ Changing the browser window by zooming or scrolling may result in the display of a different\ selection of tracks.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator.\ \
\
For automated download and analysis, the genome annotation is stored in bigBed files that\
can be downloaded from\
our download server.\
Individual\
regions or the whole genome annotation can be obtained using our tool bigBedToBed\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g. \
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/problematic/comments.bb -chrom=chr21 -start=0 -end=100000000 stdout
\
\ Files were downloaded from the respective databases and converted to bigBed format.\ The procedure is documented in our\ hg38 makeDoc file.\
\ \\ Thanks to Anna Benet-Pagès, Max Haeussler, Angie Hinrichs, Daniel Schmelter, and Jairo\ Navarro at the UCSC Genome Browser for planning, building, and testing these tracks. The\ underlying data comes from the\ ENCODE Blacklist and some parts were copied manually from the HGNC and NCBI\ RefSeq tracks.\
\ \\ Amemiya HM, Kundaje A, Boyle AP.\ \ The ENCODE Blacklist: Identification of Problematic Regions of the Genome.\ Sci Rep. 2019 Jun 27;9(1):9354.\ PMID: 31249361; PMC: PMC6597582\
\ \\ Dwarshuis N, Kalra D, McDaniel J, Sanio P, Alvarez Jerez P, Jadhav B, Huang WE, Mondal R, Busby B,\ Olson ND et al.\ \ The GIAB genomic stratifications resource for human reference genomes.\ Nat Commun. 2024 Oct 19;15(1):9029.\ PMID: 39424793; PMC: PMC11489684\
\ \\ Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, Gonzalez-Porta M, Eberle MA,\ Tezak Z, Lababidi S et al.\ \ Best practices for benchmarking germline small-variant calls in human genomes.\ Nat Biotechnol. 2019 May;37(5):555-560.\ PMID: 30858580; PMC: PMC6699627\
\ \\ Li H.\ \ Finding easy regions for short-read variant calling from pangenome data.\ ArXiv. 2025 Aug 8;.\ PMID: 40799803; PMC: PMC12340882\
\ \\ Pan B, Ren L, Onuchic V, Guan M, Kusko R, Bruinsma S, Trigg L, Scherer A, Ning B, Zhang C et\ al.\ \ Assessing reproducibility of inherited variants detected with short-read whole genome\ sequencing.\ Genome Biol. 2022 Jan 3;23(1):2.\ PMID: 34980216; PMC: PMC8722114\
\ map 1 compositeTrack on\ hideEmptySubtracks on\ longLabel Problematic/special genomic regions for sequencing or very variable regions\ parent problematicSuper\ priority 1\ shortLabel Problematic Regions\ track problematic\ type bigBed 3 +\ visibility pack\ yale_parents Pseudogene Parents bigBed 9 + Yale Pseudogene Parents 3 1 0 0 0 127 127 127 0 0 0\ These tracks contain pseudogene predictions and their parents as identified by PseudoPipe.\ PseudoPipe is a homology-based\ computational pipeline that can search a mammalian genome and identify pseudogene sequences\ comprehensively and consistently.\
\\ Pseudogenes are genomic sequences that bear similarity to specific protein-coding genes, but are\ unable to produce functional proteins due to the existence of frameshifts, premature stop codons, or\ other deleterious mutations. They arise from gene duplication or retrotransposition events and are\ important resources in understanding the evolutionary history of genes and genomes.
\ \This composite track consists of two subtracks: the Pseudogenes track and the Pseudogene\ Parents track.
\\ The Pseudogene Parents track displays parent genes and pseudogenes\ labeled with their HUGO\ IDs, which were derived from Ensembl gene IDs provided by the Gerstein lab after dataset creation. It includes indicators for pseudogenes. \ These indicators do not show pseudogene locations directly but instead indicate how many pseudogenes\ are associated with each gene and link to their genomic regions in the Pseudogenes track.
\\ The Pseudogenes track shows pseudogenes labeled with their parent HUGO ID and colored\ according to pseudogene type. The authors assigned PGOHUMG IDs to genes and PGOHUMT IDs to\ transcripts. Note: Not all PseudoPipe IDs could be mapped back to their original Ensembl\ IDs. In these cases, the gene ID is listed as NA.
\ \ Pseudogene types:\Each parent gene is shown with associated pseudogenes represented as grey blocks. These blocks\ do not reflect actual pseudogene locations but rather indicate the count of pseudogenes linked to\ the gene.\
\\ If a parent gene has four grey blocks beneath it, this indicates the presence of four pseudogenes\ elsewhere in the genome. Hovering over an item displays the gene type, ID (Ensembl transcript ID\ or PseudoPipe transcript ID), and the genome position of the gene or pseudogene, with a link to\ that genomic region.\
\ \Pseudogenes are colored by type.
\\ Hovering over a pseudogene item shows the pseudogene type, parent HUGO gene symbol, and the Ensembl\ parent transcript ID, which links to the genome position of the parent gene.
\ \\ The PseudoPipe pipeline identifies pseudogenes through a series of steps. It first uses BLAST to\ rapidly cross-reference potential parent proteins against the intergenic regions of the genome. The\ resulting raw hits are then processed by removing redundancies, clustering neighboring sequences,\ and aligning each cluster with a unique parent gene. Finally, pseudogenes are classified based on a\ combination of criteria, including homology, intron-exon structure, and the presence of stop codons\ or frameshifts. This method is designed to detect pseudogenes that are unable to be translated into\ proteins.
\\ These tracks were generated using a Bash script that processes a GTF file with pseudogene\ annotations by removing duplicates, correcting overlapping exons, and converting the data to BED\ format with pseudoPipeToBed.py. This script extracts gene and transcript IDs, merges overlapping\ exons, assigns colors based on pseudogene type, and outputs a BED file with gene and parent\ annotations. PseudoPipeParents.py then links pseudogenes to their functional genes by determining\ parent gene coordinates, updating pseudogene entries with interactive browser links and generating a\ parent BED file. The final data are formatted into pseudoPipePgenes.bb and pseudoPipeParents.bb BigBed\ files. The detailed documentation (makeDoc) and \ Python scripts are available in our GitHub repository.\
\ \The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ The data may also be explored interactively using our\ REST API.
\For automated download and analysis, the genome annotation is stored at UCSC in bigBed files\ that can be downloaded from the\ download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed which can be compiled from the source code or downloaded as a precompiled\ binary for your system.
\\ Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/hg38/pseudogenes/pseudoPipePgenes.bb -chrom=chr21 -start=0 -end=10000000 stdout\ \ \Thanks to the Gerstein lab at Yale University for making this data available, and to Cristina\ Sisu for providing data in GTF format with parent annotations.
\ \\ Zhang Z, Carriero N, Zheng D, Karro J, Harrison PM, Gerstein M.\ \ PseudoPipe: an automated pseudogene identification pipeline.\ Bioinformatics. 2006 Jun 15;22(12):1437-9.\ PMID: 16574694\
\ genes 1 bigDataUrl /gbdb/hg38/pseudogenes/pseudoPipeParents.bb\ html pseudogenes.html\ itemRgb on\ labelFields hugo\ labelSeparator " "\ longLabel Yale Pseudogene Parents\ mouseOver Gene type: ${geneType}\ The recombination rate track represents calculated rates of recombination based\ on the genetic maps from deCODE (Halldorsson et al., 2019) and 1000 Genomes\ (2013 Phase 3 release, lifted from hg19). The deCODE map is more recent, has a higher \ resolution and was natively created on hg38 and therefore recommended. \ For the Recomb. deCODE average track, the recombination rates for chrX represent the female rate.\
\ \This track also includes a subtrack with all the\ individual deCODE recombination events and another subtrack with several thousand\ de-novo mutations found in the deCODE sequencing data. These two tracks are hidden by\ default and have to be switched on explicitly on the configuration page.\
\ \\ This is a super track that contains different subtracks, three with the deCODE\ recombination rates (paternal, maternal and average) and one with the 1000\ Genomes recombination rate (average). These tracks are in \ signal graph\ (wiggle) format. By default, to show most recombination hotspots, their maximum\ value is set to 100 cM, even though many regions have values higher than 100.\ The maximum value can be changed on the configuration pages of the tracks.\
\ \\ There are two more tracks that show additional details provided by deCODE: one\ subtrack with the raw data of all cross-overs tagged with their proband ID and\ another one with around 8000 human de-novo mutation variants that are linked to\ cross-over changes.\
\ \\ The deCODE genetic map was created at \ deCODE Genetics. It is based \ on microarrays assaying 626,828 SNP markers that allowed to identify 1,476,140 crossovers in\ 56,321 paternal meioses and 3,055,395 crossovers in 70,086 maternal meioses.\ In total, the data is based on 4,531,535 crossovers in 126,427 meioses. By\ using WGS data with 9,305,070 SNPs, the boundaries for 761,981 crossovers were\ refined: 247,942 crossovers in 9423 paternal meioses and 514,039 crossovers in\ 11,750 maternal meioses. The average resolution of the genetic map is 682 base\ pairs (bp): 655 and 708 bp for the paternal and maternal maps, respectively.\
\ \The 1000 Genomes genetic map is based on the IMPUTE genetic map based on 1000 Genomes Phase 3, on hg19 coordinates. It\ was converted to hg38 by Po-Ru Loh at the Broad Institute. After a run of \ liftOver, he post-processed the data to deal with situations in which\ consecutive map locations became much closer/farther after lifting. The\ heuristic used is sufficient for statistical phasing but may not be optimal for\ other analyses. For this reason, and because of its higher resolution, the DeCODE\ map is therefore recommended for hg38.\
\ \As with all other tracks, the data conversion commands and pointers to the\ original data files are documented in the \ makeDoc file of this track.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr17 -start=45941345 -end=45942345 http://hgdownload.soe.ucsc.edu/gbdb/hg38/recombRate/recombAvg.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ This track was produced at UCSC using data that are freely available for\ the deCODE\ and 1000 Genomes genetic maps. Thanks to Po-Ru Loh at the\ Broad Institute for providing the code to lift the hg19 1000 Genomes map data to hg38.\
\ \\ 1000 Genomes Project Consortium., Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA,\ Hurles ME, McVean GA.\ \ A map of human genome variation from population-scale sequencing.\ Nature. 2010 Oct 28;467(7319):1061-73.\ PMID: 20981092; PMC: PMC3042601\
\ \\ Halldorsson BV, Palsson G, Stefansson OA, Jonsson H, Hardarson MT, Eggertsson HP, Gunnarsson B,\ Oddsson A, Halldorsson GH, Zink F et al.\ \ Characterizing mutagenic effects of recombination through a sequence-level genetic map.\ Science. 2019 Jan 25;363(6425).\ PMID: 30679340\
\ map 0 bigDataUrl /gbdb/hg38/recombRate/recombAvg.bw\ html recombRate2.html\ longLabel Recombination rate: deCODE Genetics, average from paternal and maternal (mat for chrX)\ maxHeightPixels 128:60:8\ parent recombRate2\ priority 1\ shortLabel Recomb. deCODE Avg\ track recombAvg\ type bigWig\ viewLimits 0.0:100\ viewLimitsMax 0:150000\ visibility full\ ncbiRefSeq RefSeq All genePred NCBI RefSeq genes, curated and predicted (NM_*, XM_*, NR_*, XR_*, NP_*, YP_*) 1 1 12 12 120 133 133 187 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 12,12,120\ idXref ncbiRefSeqLink mrnaAcc name\ longLabel NCBI RefSeq genes, curated and predicted (NM_*, XM_*, NR_*, XR_*, NP_*, YP_*)\ parent refSeqComposite off\ priority 1\ shortLabel RefSeq All\ track ncbiRefSeq\ ReMapDensity ReMap density bigWig ReMap density 0 1 0 0 0 127 127 127 0 0 0\ This track represents the ReMap Atlas of regulatory regions, which consists of a\ large-scale integrative analysis of all Public ChIP-seq data for transcriptional\ regulators from GEO, ArrayExpress, and ENCODE. \
\ \\ Below is a schematic diagram of the types of regulatory regions: \
\
\
\ This 4th release of ReMap (2022) presents the analysis of a total of 8,103 \ quality controlled ChIP-seq (n=7,895) and ChIP-exo (n=208) data sets from public\ sources (GEO, ArrayExpress, ENCODE). The ChIP-seq/exo data sets have been mapped\ to the GRCh38/hg38 human assembly. The data set is defined as a ChIP-seq \ experiment in a given series (e.g. GSE46237), for a given TF (e.g. NR2C2), in a\ particular biological condition (i.e. cell line, tissue type, disease state, or\ experimental conditions; e.g. HELA). Data sets were labeled by concatenating\ these three pieces of information, such as GSE46237.NR2C2.HELA. \ \
\Those merged analyses cover a total of 1,211 DNA-binding proteins\ (transcriptional regulators) such as a variety of transcription factors (TFs),\ transcription co-activators (TCFs), and chromatin-remodeling factors (CRFs) for\ 182 million peaks. \
\ \
\
\
\
Public ChIP-seq data sets were extracted from Gene Expression Omnibus (GEO) and\
ArrayExpress (AE) databases. For GEO, the query\
\
'('chip seq' OR 'chipseq' OR\
'chip sequencing') AND 'Genome binding/occupancy profiling by high throughput\
sequencing' AND 'homo sapiens'[organism] AND NOT 'ENCODE'[project]'\
\
was used to return a list of all potential data sets to analyze, which were then manually \
assessed for further analyses. Data sets involving polymerases (i.e. Pol2 and\
Pol3), and some mutated or fused TFs (e.g. KAP1 N/C terminal mutation, GSE27929)\
were excluded.\
\ Available ENCODE ChIP-seq data sets for transcriptional regulators from the\ ENCODE portal were processed with the\ standardized ReMap pipeline. The list of ENCODE data was retrieved as FASTQ files from the\ ENCODE portal\ using the following filters:\
\ Both Public and ENCODE data were processed similarly. Bowtie 2 (PMC3322381) (version 2.2.9) with options -end-to-end -sensitive was used to align all\ reads on the genome. Biological and technical\ replicates for each unique combination of GSE/TF/Cell type or Biological condition\ were used for peak calling. TFBS were identified using MACS2 peak-calling tool\ (PMC3120977) (version 2.1.1.2) in order to follow ENCODE ChIP-seq guidelines,\ with stringent thresholds (MACS2 default thresholds, p-value: 1e-5). An input data\ set was used when available.\
\ \ \\ To assess the quality of public data sets, a score was computed based on the\ cross-correlation and the FRiP (fraction of reads in peaks) metrics developed by\ the ENCODE Consortium (https://genome.ucsc.edu/ENCODE/qualityMetrics.html). Two\ thresholds were defined for each of the two cross-correlation ratios (NSC,\ normalized strand coefficient: 1.05 and 1.10; RSC, relative strand coefficient:\ 0.8 and 1.0). Detailed descriptions of the ENCODE quality coefficients can be\ found at https://genome.ucsc.edu/ENCODE/qualityMetrics.html. The\ phantompeak tools suite was used\ (https://code.google.com/p/phantompeakqualtools/) to compute\ RSC and NSC.\
\\ Please refer to the ReMap 2022, 2020, and 2018 publications for more details\ (citation below).\
\ \ \ \\ ReMap Atlas of regulatory regions data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ Individual BED files for specific TFs, cells/biotypes, or data sets can be\ found and downloaded on the ReMap website.\
\ \\ Chèneby J, Gheorghe M, Artufel M, Mathelier A, Ballester B.\ \ ReMap 2018: an updated atlas of regulatory regions from an integrative analysis of DNA-binding ChIP-\ seq experiments.\ Nucleic Acids Res. 2018 Jan 4;46(D1):D267-D275.\ PMID: 29126285; PMC: PMC5753247\
\\ Chèneby J, Ménétrier Z, Mestdagh M, Rosnet T, Douida A, Rhalloussi W, Bergon A, Lopez\ F, Ballester B.\ \ ReMap 2020: a database of regulatory regions from an integrative analysis of Human and Arabidopsis\ DNA-binding sequencing experiments.\ Nucleic Acids Res. 2020 Jan 8;48(D1):D180-D188.\ PMID: 31665499; PMC: PMC7145625\
\\ Griffon A, Barbier Q, Dalino J, van Helden J, Spicuglia S, Ballester B.\ \ Integrative analysis of public ChIP-seq experiments reveals a complex multi-cell regulatory\ landscape.\ Nucleic Acids Res. 2015 Feb 27;43(4):e27.\ PMID: 25477382; PMC: PMC4344487\
\\ Hammal F, de Langen P, Bergon A, Lopez F, Ballester B.\ \ ReMap 2022: a database of Human, Mouse, Drosophila and Arabidopsis regulatory regions from an\ integrative analysis of DNA-binding sequencing experiments.\ Nucleic Acids Res. 2022 Jan 7;50(D1):D316-D325.\ PMID: 34751401; PMC: PMC8728178\
\ \ regulation 0 autoScale on\ bigDataUrl /gbdb/hg38/reMap/reMapDensity2022.bw\ html ../reMap\ longLabel ReMap density\ parent ReMap on\ priority 1\ shortLabel ReMap density\ track ReMapDensity\ type bigWig\ visibility hide\ rmsk RepeatMasker rmsk Repeating Elements by RepeatMasker 1 1 0 0 0 127 127 127 1 0 0\ This track was created by using Arian Smit's\ RepeatMasker\ program, which screens DNA sequences\ for interspersed repeats and low complexity DNA sequences. The program\ outputs a detailed annotation of the repeats that are present in the\ query sequence (represented by this track), as well as a modified version\ of the query sequence in which all the annotated repeats have been masked\ (generally available on the\ Downloads page). RepeatMasker uses the\ Repbase Update library of repeats from the\ Genetic \ Information Research Institute (GIRI).\ Repbase Update is described in Jurka (2000) in the References section below.
\ \This track and the masking information in our \ hg38 genome download FASTA files was created in 2010 with the original RepBase library from 2010-03-02 and RepeatMasker 3.0.1.\ Since April 2019, RepBase is under a commercial license, we cannot distribute\ it or update the track using the RepBase library without a license. Therefore, and for\ compatibility with past results, given how central the masking is for many other\ annotations, we decided to not update the repeatmasking of hg38. However, you can show the\ small differences between the RepeatMasker 3/RepBase from 2010 and RepeatMasker 4/DFAM\ from 2020 using the track "RepeatMasker Viz" in the same track group. It\ contains two subtracks, one with the old and one with the new data. Also, these\ tracks have many more visualization options than the original RepeatMasker\ track.\
\ \However, the last track update time of this track at UCSC is not 2010, because we had to add\ repeatmasking annotations to the rarely used _alt and _fix "patch" sequences of\ the hg38 genome. The repeatmasking annotations of the main chromosomes were unaffected\ and have not changed since 2010.\ For more information on genome patches, see our blog post.\
\ \\ In full display mode, this track displays up to ten different classes of repeats:\
\ The level of color shading in the graphical display reflects the amount of\ base mismatch, base deletion, and base insertion associated with a repeat\ element. The higher the combined number of these, the lighter the shading.\
\ \\ A "?" at the end of the "Family" or "Class" (for example, DNA?) signifies that\ the curator was unsure of the classification. At some point in the future,\ either the "?" will be removed or the classification will be changed.
\ \\ Data are generated using the RepeatMasker -s flag. Additional flags\ may be used for certain organisms. Repeats are soft-masked. Alignments may\ extend through repeats, but are not permitted to initiate in them.\ See the FAQ for more information.\
\ \\ Thanks to Arian Smit, Robert Hubley and GIRI for providing the tools and\ repeat libraries used to generate this track.\
\ \\ Smit AFA, Hubley R, Green P. RepeatMasker Open-3.0.\ \ https://www.repeatmasker.org/. 1996-2010.\
\ \\ Repbase Update is described in:\
\ \\ Jurka J.\ \ Repbase Update: a database and an electronic journal of repetitive elements.\ Trends Genet. 2000 Sep;16(9):418-420.\ PMID: 10973072\
\ \\ For a discussion of repeats in mammalian genomes, see:\
\ \\ Smit AF.\ \ Interspersed repeats and other mementos of transposable elements in mammalian genomes.\ Curr Opin Genet Dev. 1999 Dec;9(6):657-63.\ PMID: 10607616\
\ \\ Smit AF.\ \ The origin of interspersed repeats in the human genome.\ Curr Opin Genet Dev. 1996 Dec;6(6):743-8.\ PMID: 8994846\
\ rep 0 canPack off\ group rep\ html rmsk\ longLabel Repeating Elements by RepeatMasker\ maxWindowToDraw 10000000\ priority 1\ shortLabel RepeatMasker\ spectrum on\ track rmsk\ type rmsk\ visibility dense\ gnomad31XPercentage Sample % > 1X bigWig gnomAD Percentage of Genome Samples with at least 1X Coverage v3.0.1 2 1 255 0 0 255 127 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_1.bw\ color 255,0,0\ longLabel gnomAD Percentage of Genome Samples with at least 1X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 1\ shortLabel Sample % > 1X\ track gnomad31XPercentage\ viewLimits 0:1\ gnomad4Exome1XPercentage Sample % > 1X bigWig gnomAD Percentage of Exome Samples with at least 1X Coverage v4.0 2 1 255 0 0 255 127 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_1.bw\ color 255,0,0\ longLabel gnomAD Percentage of Exome Samples with at least 1X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 1\ shortLabel Sample % > 1X\ track gnomad4Exome1XPercentage\ viewLimits 0:1\ miRnaAtlasSample1BarChart Sample 1 bigBarChart miRNA Tissue Atlas microRna Expression 2 1 0 0 0 127 127 127 0 0 0\ The Human miRNA Tissue Atlas is a\ catalog of tissue-specific microRNA (miRNA) expression across 62 tissues. This track contains\ quantile normalized miRNA expression data sampled from two individuals and mapped to\ miRBase v21 coordinates. The track contains two subtracks, one\ for each individual sampled.
\ \\ The Tissue Specificity Index (TSI) is analogous to the "tau" value for mRNA expression,\ and is calculated as described in the\ \ associated publication. Values closer to 0 indicate miRNAs expressed in many or all tissues,\ while values closer to 1 indicate miRNAs expressed only in a specific tissue or tissues. To\ browse miRNAs by TSI value, please see the\ miRNA Tissue Atlas.
\ \\ This track is formatted as a barChart track,\ similar to the GTEx or the\ TCGA Cancer Expression tracks, where the\ heights of each bar indicate the expression value for the miRNA in a specific tissue. The tissues\ sampled are described in the table below:\
\| Bar Color | Sample 1 | Sample 2 |
| Adipocyte | Adipocyte | |
| Artery | Artery | |
| Colon | Colon | |
| Dura mater | Dura mater | |
| Kidney | Kidney | |
| Liver | Liver | |
| Lung | Lung | |
| Muscle | Muscle | |
| Myocardium | Myocardium | |
| Skin | Skin | |
| Spleen | Spleen | |
| Stomach | Stomach | |
| Testis | Testis | |
| Thyroid | Thyroid | |
| Small intestine | ||
| Bone | ||
| Gallbladder | ||
| Fascia | ||
| Bladder | ||
| Epididymis | ||
| Tunica albuginea | ||
| Nervus intercostalis | ||
| Arachnoid mater | ||
| Brain | ||
| Small intestine duodenum | ||
| Small intestine jejunum | ||
| Pancreas | ||
| Kidney glandula suprarenalis | ||
| Kidney cortex renalis | ||
| Esophagus | ||
| Prostate | ||
| Bone marrow | ||
| Vein | ||
| Lymph node | ||
| Nerve not specified | ||
| Pleura | ||
| Pituitary gland | ||
| Spinal cord | ||
| Thalamus | ||
| Brain white matter | ||
| Nucleus caudatus | ||
| Kidney medulla renalis | ||
| Brain gray_matter | ||
| Cerebral cortex temporal | ||
| Cerebral cortex frontal | ||
| Cerebral cortex occipital | ||
| Cerebellum |
\ The 14 shared tissues sampled across both individuals are presented in the same order for easier comparison.\
\ \\ The underlying expression matrix and TSI values can be obtained from the\ miRNA tissue atlas website, in the\ data_matrix_quantile.txt and tsi_quantile.csv files.\
\ \\ Ludwig N, Leidinger P, Becker K, Backes C, Fehlmann T, Pallasch C, Rheinheimer S, Meder B,\ Stähler C, Meese E et al.\ \ Distribution of miRNA expression across human tissues.\ Nucleic Acids Res. 2016 May 5;44(8):3865-77.\ PMID: 26921406; PMC: PMC4856985\
\ expression 1 barChartBars adipocyte artery colon dura_mater kidney liver lung muscle myocardium skin spleen stomach testis thyroid small_intestine bone gallbladder fascia bladder epididymis tunica_albuginea nerve_nervus_intercostalis arachnoid_mater brain\ barChartColors #F7A028 #F73528 #DEBE98 #86BF80 #CDB79E #CDB79E #9ACD32 #7A67AE #9745AC #1E90FF \\#CDB79E #FFD39B #A6A6A6 #008B45 #CDB79E #BD34D7 #CDA7FE #4C7CD7 #CBD79E #A6F6A1 \\#A6CEA4 #FFD700 #86BF10 #EEEE00\ barChartLabel Tissue\ barChartMatrixUrl /gbdb/hgFixed/human/expMatrix/miRnaAtlasSample1Matrix.txt\ barChartSampleUrl /gbdb/hgFixed/human/expMatrix/miRnaAtlasSample1.txt\ barChartUnit Quantile_Normalized_Expression\ bigDataUrl /gbdb/hg38/bbi/miRnaAtlasSample1.bb\ configurable on\ group expression\ html miRnaAtlas\ longLabel miRNA Tissue Atlas microRna Expression\ maxLimit 52000\ parent miRnaAtlasSample1\ searchIndex name\ shortLabel Sample 1\ subGroups view=a_A\ track miRnaAtlasSample1BarChart\ url2 http://www.mirbase.org/cgi-bin/query.pl?terms=$$\ url2Label miRBase v21 Precursor Accession:\ visibility full\ covidHgiGwasR4PvalA2 Severe COVID vars bigLolly 9 + Severe respiratory COVID risk variants from the COVID-19 HGI GWAS Analysis A2 (4336 cases, 12 studies, Rel 4: Oct 2020) 0 1 0 0 0 127 127 127 0 0 22 chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22, phenDis 1 bigDataUrl /gbdb/hg38/covidHgiGwas/covidHgiGwasR4.A2.hg38.bb\ longLabel Severe respiratory COVID risk variants from the COVID-19 HGI GWAS Analysis A2 (4336 cases, 12 studies, Rel 4: Oct 2020)\ parent covidHgiGwasR4Pval on\ priority 1\ shortLabel Severe COVID vars\ track covidHgiGwasR4PvalA2\ snpediaAll SNPedia all bigBed 9 + SNPedia all SNPs (including empty pages) 0 1 50 0 100 152 127 177 0 0 0 https://www.snpedia.com/index.php/$$ phenDis 1 bigDataUrl /gbdb/hg38/bbi/snpediaAll.bb\ color 50,0,100\ exonNumbers off\ itemRgb on\ longLabel SNPedia all SNPs (including empty pages)\ mouseOverField note\ parent snpedia\ searchIndex name\ shortLabel SNPedia all\ track snpediaAll\ type bigBed 9 +\ url https://www.snpedia.com/index.php/$$\ urlLabel Link to SNPedia page:\ spliceAiAccPlus SpliceAI Acceptor Plus bigWig 0 1 SpliceAI Splice Acceptor Sites, Plus Strand 2 1 0 0 0 127 127 127 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/bbi/spliceAi/wildtype/spliceAiAcceptorPlus.bw\ longLabel SpliceAI Splice Acceptor Sites, Plus Strand\ parent spliceAIWt on\ priority 1\ shortLabel SpliceAI Acceptor Plus\ track spliceAiAccPlus\ type bigWig 0 1\ spliceAIsnvs SpliceAI SNVs bigBed 9 + SpliceAI SNVs (unmasked) 1 1 0 0 0 127 127 127 0 0 0\ SpliceAI is an open-source deep\ learning algorithm that predicts splicing probability for nucleotides and \ as a result can score DNA variants for splicing impact.\ Such variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a \ lookup tool \ is provided by the Broad institute.\
\ \The spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added and the new sequence is\ scored again. The "wildtype" container track shows the scores for the genome\ sequence itself and the "variants" container track shows the impact of all\ possible variants close to known splice sites. The "wildtype" subtracks are\ useful when looking at new transcript models, to evaluate how likely exon\ boundaries are. The "variants" subtracks are used to evaluate the impact of\ variants onto splicing, typically in medical diagnostics.\
\ \\ SpliceAI only annotates variants close to splice sites of genes defined by the \ Gencode gene annotation track. Additionally, SpliceAI does not annotate variants if they are\ close to chromosome ends (5kb on either side), deletions of length greater than\ twice the input parameter -D, or inconsistent with the reference fasta file.\
\ \\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \\ Variants are colored according to Walker et al. 2023 splicing imact:\
\\ The scores range from 0 to 1 and can be interpreted as the \ probability of the variant being splice-altering. In the paper, a detailed characterization is \ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\
The data were downloaded from Illumina. \
The spliceAI scores are represented in the VCF INFO field as \
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain \
\ Since most of the values are 0 or almost 0, we selected only those variants \ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \ \\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are \ made available by Illumina only for academic or not-for-profit research only. \ By accessing the SpliceAI data, you acknowledge and agree that you may only \ use this data for your own personal academic or not-for-profit research only, \ and not for any other purposes. You may not use this data for any for-profit, \ clinical, or other commercial purpose without obtaining a commercial license \ from Illumina, Inc.\
\ \\ Thanks to Illumina for making the data available. Thanks to Michael Hiller, Francois Lecoquierre and\ Jean-Madeleine de Sainte Agathe for making available and suggesting the SpliceAI wildtype tracks.\
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/spliceAIsnvs.bb\ filter.AIscore 0.02\ filterLabel.spliceType Splice type\ filterLimits.AIscore 0.02:1\ filterValues.spliceType donor_gain|Donor gain,donor_loss|Donor loss,acceptor_gain|Acceptor gain,acceptor_loss|Acceptor loss\ html spliceAI\ itemRgb on\ longLabel SpliceAI SNVs (unmasked)\ mouseOver Change: $name\ SpliceAI is an open-source deep\ learning algorithm that predicts splicing probability for nucleotides and \ as a result can score DNA variants for splicing impact.\ Such variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a \ lookup tool \ is provided by the Broad institute.\
\ \The spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added and the new sequence is\ scored again. The "wildtype" container track shows the scores for the genome\ sequence itself and the "variants" container track shows the impact of all\ possible variants close to known splice sites. The "wildtype" subtracks are\ useful when looking at new transcript models, to evaluate how likely exon\ boundaries are. The "variants" subtracks are used to evaluate the impact of\ variants onto splicing, typically in medical diagnostics.\
\ \\ SpliceAI only annotates variants close to splice sites of genes defined by the \ Gencode gene annotation track. Additionally, SpliceAI does not annotate variants if they are\ close to chromosome ends (5kb on either side), deletions of length greater than\ twice the input parameter -D, or inconsistent with the reference fasta file.\
\ \\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \\ Variants are colored according to Walker et al. 2023 splicing imact:\
\\ The scores range from 0 to 1 and can be interpreted as the \ probability of the variant being splice-altering. In the paper, a detailed characterization is \ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\
The data were downloaded from Illumina. \
The spliceAI scores are represented in the VCF INFO field as \
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain \
\ Since most of the values are 0 or almost 0, we selected only those variants \ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \ \\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are \ made available by Illumina only for academic or not-for-profit research only. \ By accessing the SpliceAI data, you acknowledge and agree that you may only \ use this data for your own personal academic or not-for-profit research only, \ and not for any other purposes. You may not use this data for any for-profit, \ clinical, or other commercial purpose without obtaining a commercial license \ from Illumina, Inc.\
\ \\ Thanks to Illumina for making the data available. Thanks to Michael Hiller, Francois Lecoquierre and\ Jean-Madeleine de Sainte Agathe for making available and suggesting the SpliceAI wildtype tracks.\
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ phenDis 1 compositeTrack on\ dataVersion Illumina SpliceAI Score v1.3 (August 22, 2024)\ group phenDis\ longLabel SpliceAI: Splice Variant Prediction Score\ parent spliceImpactSuper on\ priority 1\ shortLabel SpliceAI Variants\ tableBrowser off\ track spliceAI\ type bigBed 9 +\ visibility dense\ unipAliSwissprot SwissProt Aln. bigPsl UCSC alignment of SwissProt proteins to genome (dark blue: main isoform, light blue: alternative isoforms) 3 1 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorTickColor contrastingColor\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/uniprot/unipAliSwissprot.bb\ indelDoubleInsert on\ indelQueryInsert on\ itemRgb on\ labelFields name,acc,uniprotName,geneName,hgncSym,refSeq,refSeqProt,ensProt\ longLabel UCSC alignment of SwissProt proteins to genome (dark blue: main isoform, light blue: alternative isoforms)\ mouseOver UniProt record accession: $accThis track displays regulatory regions in the human genome identified using ENCODE \ data, specifically spanning ENCODE phases 2 through 4. It highlights genomic \ regions bound by DNA-associated proteins involved in transcriptional regulation, \ such as RNA polymerase, transcription factors (TFs), and chromatin remodeling \ proteins. Sequence-specific TFs bind directly to short DNA motifs via their \ DNA-binding domains, while other DNA-associated proteins interact with DNA \ indirectly through protein-protein interactions with sequence-specific TFs. Chromatin \ immunoprecipitation followed by sequencing (ChIP-seq) is a high-throughput method \ for mapping genome-wide protein-DNA interactions. Regions of high ChIP signal, \ commonly referred to as ChIP-seq peaks, indicate protein binding sites. For each DNA\ -associated protein, all ENCODE ChIP-seq peaks across biosamples were integrated to generate \ a set of representative peaks (rPeaks). This track displays these rPeaks alongside \ detected DNA motif sites.
\ \Each rPeak is represented as a gray box, with the shade of gray corresponding \ to the maximum ChIP-seq signal observed across contributing biosamples. The HGNC \ gene name of the associated protein is displayed to the left of the box. If the \ rPeak overlaps a cognate TF motif site in the collection built previously (PMID: \ 37104580 DOI: 10.1126/science.abn7930), \ the motif site is highlighted in green.
\ \Clicking on an rPeak provides detailed information about the biosamples where the \ rPeak was detected, including the count of biosamples with contributing ChIP-seq peaks \ and the total number of biosamples assayed for the protein. Links to relevant ENCODE \ ChIP-seq experiments and overlapping ENCODE candidate cis-regulatory elements (cCREs) \ are also provided.
\ \By default, rPeaks for all 912 DNA-associated proteins with ENCODE ChIP-seq data \ are displayed. Users can customize the display by selecting specific DNA-associated \ proteins in the track settings.
\ \2,509 ENCODE ChIP-seq experiments were integrated from 912 DNA-associated \ proteins across 1,152 unique biosamples to produce representative peaks (rPeaks) \ for each protein. The processing steps were as follows:
\ \The raw data for the ENCODE TF rPeak track will soon be available.
\ \\ The raw data can be explored interactively with the Table Browser,\ for download, intersection or correlations with other tracks. To join this track with others\ based on the chromosome positions, use the Data Integrator.\ \
\ Regarding access to this data track in the Genome Browser, for automated download \ and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ The file for this track is called TFrPeakClusters.bb. Individual\ regions or the whole genome annotation can be obtained using our tool bigBedToBed\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool\ can also be used to obtain only features within a given range, e.g.\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/ENCODE4/TFrPeakClusters.bb -chrom=chr21 -start=0 -end=100000000 stdout
\ \ \For automated access, this track like all others, is also available via our\ API. However, for bulk processing in\ pipelines, downloading the data and/or using bigBed files as described above is\ usually faster.
\ \This track was made possible thanks to the efforts of the ENCODE Consortium, \ ENCODE ChIP-seq production laboratories, and the ENCODE Data Coordination Center \ for generating and processing the ChIP-seq datasets. The ENCODE accession numbers \ for the constituent datasets are accessible from the peak details page. Special thanks \ to Drs. Mingshi Gao, Greg Andrews, Jill Moore, and Zhiping Weng at UMass Chan Medical \ School, who were members of the ENCODE Data Analysis Center, for developing this track, \ including providing the rPeak and motif datasets and associated metadata and building the \ track. We also extend our gratitude to Max Haeussler and Jonathan Casper from the UCSC \ Genome Browser Project Team for their assistance in developing this track. For updates \ on the track, please contact the Weng lab.
\ regulation 1 bigDataUrl /gbdb/hg38/bbi/ENCODE4/TFrPeakClusters.bb\ decorator.default.bigDataUrl /gbdb/hg38/bbi/ENCODE4/TFrPeakClustersDecorator.bb\ detailsDynamicTable json_table|experiments support this rPeak\ filterType.factor multipleListOr\ filterValues.factor ADNP,AFF1,AFF4,AGO1,AGO2,AHDC1,AHR,AKAP8,AKNA,ARHGAP35,ARID1B,ARID2,ARID3A,ARID4A,ARID4B,ARID5B,ARNTL,ARNT,AR,ASH1L,ASH2L,ATF1,ATF2,ATF3,ATF4,ATF5,ATF6,ATF7,ATOH8,BACH1,BATF,BCL11A,BCL11B,BCL3,BCL6B,BCL6,BCLAF1,BCOR,BDP1,BHLHE40,BMI1,BNC2,BRCA1,BRCA2,BRD4,BRD9,C11orf30,CAMTA2,CBFA2T2,CBFA2T3,CBFB,CBX1,CBX2,CBX3,CBX5,CBX8,CC2D1A,CCAR2,CCNT2,CDC5L,CEBPA,CEBPB,CEBPD,CEBPG,CEBPZ,CERS6,CGGBP1,CHAMP1,CHD1,CHD2,CHD4,CHD7,CLOCK,CREB1,CREB3L1,CREB3,CREB5,CREM,CSDC2,CSDE1,CSRNP3,CTBP1,CTBP2,CTCFL,CTCF,CUX1,DACH1,DBP,DDIT3,DDX20,DEAF1,DEK,DIDO1,DLX6,DMAP1,DMTF1,DNMT1,DNMT3B,DPF2,DR1,DRAP1,DZIP1,E2F1,E2F2,E2F3,E2F4,E2F5,E2F6,E2F7,E2F8,E4F1,EBF1,EEA1,EED,EGR1,EGR2,EHF,EHMT2,ELF1,ELF2,ELF3,ELF4,ELK1,ELK4,EMX1,EP300,EP400,ERF,ERG,ESR1,ESRRA,ESRRG,ETS1,ETV1,ETV4,ETV5,ETV6,EZH2phosphoT487,EZH2,FEZF1,FIP1L1,FOSB,FOSL1,FOSL2,FOS,FOXA1,FOXA2,FOXA3,FOXC1,FOXF2,FOXJ2,FOXJ3,FOXK1,FOXK2,FOXM1,FOXO1,FOXO4,FOXP1,FOXP2,FOXP4,FOXS1,FUBP1,FUBP3,FUS,GABPA,GABPB1,GATA1,GATA2,GATA3,GATA4,GATAD1,GATAD2A,GATAD2B,GFI1B,GFI1,GLI2,GLI4,GLIS1,GLIS2,GLIS3,GLYR1,GMEB1,GMEB2,GPN1,GTF2A2,GTF2B,GTF2E2,GTF2F1,GTF2I,GTF3A,GTF3C2,GZF1,HBP1,HCFC1,HDAC1,HDAC2,HDAC3,HDAC6,HDAC8,HDGF,HES1,HES2,HES4,HEYL,HHEX,HIC1,HIC2,HINFP,HIVEP1,HLF,HLTF,HMBOX1,HMG20A,HMG20B,HMGA2,HMGN3,HMGXB3,HMGXB4,HNF1A,HNF1B,HNF4A,HNF4G,HNRNPH1,HNRNPK,HNRNPLL,HNRNPL,HNRNPUL1,HOMEZ,HOXA3,HOXA5,HOXA7,HOXB13,HOXB5,HOXD1,HSF1,HSF2,HSF4,ID3,IKZF1,IKZF2,IKZF3,IKZF4,IKZF5,ILF3,INSM2,IRF1,IRF2,IRF3,IRF4,IRF5,IRF9,ISL1,ISL2,ISX,JRK,JUNB,JUND,JUN,KAT2B,KAT7,KAT8,KDM1A,KDM2A,KDM2B,KDM3A,KDM4A,KDM4B,KDM5A,KDM5B,KDM6A,KHSRP,KIAA2018,KLF10,KLF11,KLF12,KLF13,KLF16,KLF17,KLF1,KLF4,KLF5,KLF6,KLF7,KLF8,KLF9,KMT2A,KMT2B,L3MBTL2,LARP7,LBX2,LCORL,LCOR,LEF1,LIN54,MAF1,MAFF,MAFG,MAFK,MAX,MAZ,MBD1,MBD2,MBD4,MED13,MED1,MEF2A,MEF2B,MEF2C,MEF2D,MEIS1,MEIS2,MGA,MIER1,MIER2,MIER3,MITF,MIXL1,MLLT1,MLXIP,MLX,MNT,MNX1,MSX2,MTA1,MTA2,MTA3,MTF1,MTF2,MXD1,MXD3,MXD4,MXI1,MYBL2,MYB,MYC,MYNN,MYRF,MZF1,NAIF1,NANOG,NBN,NCOA1,NCOA2,NCOA3,NCOA6,NCOR1,NEUROD1,NFAT5,NFATC1,NFATC3,NFATC4,NFE2L1,NFE2L2,NFE2,NFIA,NFIB,NFIC,NFIL3,NFIX,NFKB2,NFKBIZ,NFRKB,NFXL1,NFYA,NFYB,NFYC,NKRF,NKX3-1,NONO,NR0B2,NR1H2,NR2C1,NR2C2,NR2E3,NR2F1,NR2F2,NR2F6,NR3C1,NR4A1,NR5A1,NR5A2,NRF1,NRL,NUFIP1,ONECUT1,ONECUT2,OSR2,OTX2,OVOL1,OVOL3,PAF1,PATZ1,PAWR,PAX5,PAX8,PAXIP1,PBX1,PBX2,PBX3,PCBP1,PCBP2,PHB2,PHB,PHF20,PHF21A,PHF5A,PHF8,PITX1,PKNOX1,PLSCR1,PML,POGZ,POLR2AphosphoS2,POLR2AphosphoS5,POLR2A,POLR2B,POLR2G,POLR2H,POLR3A,POU2F2,POU5F1,POU6F1,PPARG,PRDM10,PRDM15,PRDM1,PRDM4,PRDM6,PRPF4,PRRX2,PTBP1,PTRF,PTTG1,PYGO2,RAD21,RAD51,RARA,RARB,RARG,RB1,RBAK,RBBP5,RBFOX2,RBM14,RBM22,RBM25,RBM39,RBPJ,RCOR1,RCOR2,RELA,RELB,REPIN1,RERE,REST,RFX1,RFX3,RFX5,RFXANK,RFXAP,RLF,RNF219,RNF2,RORA,RREB1,RUNX1,RUNX3,RXRA,RXRB,SAFB2,SAFB,SAP130,SAP30,SATB2,SCRT1,SCRT2,SETDB1,SFPQ,SHOX2,SIN3A,SIN3B,SIRT6,SIX1,SIX4,SIX5,SKIL,SKI,SMAD1,SMAD3,SMAD4,SMAD5,SMAD9,SMARCA4,SMARCA5,SMARCB1,SMARCC1,SMARCC2,SMARCE1,SMC3,SNAI1,SNAI2,SNAPC4,SNIP1,SOX13,SOX18,SOX5,SOX6,SP110,SP140L,SP1,SP2,SP3,SP4,SP5,SP7,SPDEF,SPEN,SPI1,SREBF1,SREBF2,SRF,SRSF1,SRSF3,SRSF4,SSRP1,STAG1,STAT1,STAT3,STAT5A,STAT5B,STAT6,SUPT5H,SUZ12,TAF15,TAF1,TAF7,TAF9B,TAL1,TARDBP,TBL1XR1,TBPL1,TBP,TBX18,TBX21,TBX2,TBX3,TCF12,TCF15,TCF3,TCF4,TCF7L2,TCF7,TEAD1,TEAD2,TEAD3,TEAD4,TEF,TFAP4,TFCP2L1,TFCP2,TFDP1,TFDP2,TFE3,TGIF2,THAP11,THAP12,THAP1,THAP7,THAP8,THAP9,THRAP3,THRA,THRB,TIGD3,TIGD6,TOE1,TOX2,TOX,TP53,TP63,TRAFD1,TRIM22,TRIM24,TRIM25,TRIM28,TSC22D2,TSHZ1,TSHZ2,U2AF1,U2AF2,UBTF,USF1,USF2,VEZF1,WIZ,WRNIP1,WT1,XBP1,XRCC5,YBX1,YEATS2,YEATS4,YY1,YY2,ZBED1,ZBED4,ZBED5,ZBTB10,ZBTB11,ZBTB12,ZBTB14,ZBTB17,ZBTB1,ZBTB20,ZBTB21,ZBTB24,ZBTB25,ZBTB26,ZBTB2,ZBTB33,ZBTB37,ZBTB38,ZBTB39,ZBTB3,ZBTB40,ZBTB42,ZBTB43,ZBTB44,ZBTB46,ZBTB48,ZBTB49,ZBTB4,ZBTB5,ZBTB6,ZBTB7A,ZBTB7B,ZBTB8A,ZBTB9,ZC3H10,ZC3H11A,ZC3H4,ZC3H8,ZCCHC11,ZEB1,ZEB2,ZFHX2,ZFP14,ZFP1,ZFP36L2,ZFP36,ZFP37,ZFP3,ZFP41,ZFP64,ZFP69B,ZFP82,ZFP91,ZFX,ZFY,ZGPAT,ZHX1,ZHX2,ZHX3,ZIC2,ZIK1,ZKSCAN1,ZKSCAN5,ZKSCAN8,ZMAT3,ZMIZ1,ZMYM2,ZMYM3,ZNF101,ZNF10,ZNF121,ZNF124,ZNF12,ZNF133,ZNF134,ZNF138,ZNF140,ZNF142,ZNF143,ZNF146,ZNF148,ZNF157,ZNF165,ZNF16,ZNF175,ZNF17,ZNF180,ZNF184,ZNF189,ZNF18,ZNF197,ZNF202,ZNF205,ZNF207,ZNF20,ZNF215,ZNF217,ZNF219,ZNF221,ZNF223,ZNF224,ZNF225,ZNF230,ZNF232,ZNF234,ZNF239,ZNF24,ZNF251,ZNF256,ZNF257,ZNF25,ZNF263,ZNF264,ZNF266,ZNF26,ZNF274,ZNF276,ZNF280A,ZNF280B,ZNF280D,ZNF281,ZNF282,ZNF296,ZNF2,ZNF30,ZNF311,ZNF316,ZNF317,ZNF318,ZNF319,ZNF324,ZNF329,ZNF331,ZNF333,ZNF335,ZNF337,ZNF33A,ZNF33B,ZNF341,ZNF343,ZNF34,ZNF350,ZNF354B,ZNF354C,ZNF362,ZNF366,ZNF367,ZNF383,ZNF384,ZNF391,ZNF394,ZNF395,ZNF397,ZNF398,ZNF3,ZNF407,ZNF414,ZNF416,ZNF41,ZNF423,ZNF426,ZNF430,ZNF431,ZNF433,ZNF441,ZNF444,ZNF445,ZNF446,ZNF449,ZNF44,ZNF451,ZNF460,ZNF462,ZNF483,ZNF484,ZNF485,ZNF488,ZNF48,ZNF490,ZNF501,ZNF503,ZNF507,ZNF510,ZNF511,ZNF512B,ZNF512,ZNF513,ZNF518A,ZNF526,ZNF529,ZNF530,ZNF532,ZNF543,ZNF547,ZNF548,ZNF549,ZNF550,ZNF552,ZNF555,ZNF556,ZNF557,ZNF558,ZNF561,ZNF569,ZNF570,ZNF572,ZNF574,ZNF576,ZNF579,ZNF580,ZNF583,ZNF584,ZNF585B,ZNF589,ZNF592,ZNF596,ZNF598,ZNF600,ZNF605,ZNF607,ZNF608,ZNF609,ZNF610,ZNF614,ZNF615,ZNF616,ZNF619,ZNF623,ZNF624,ZNF629,ZNF639,ZNF644,ZNF646,ZNF652,ZNF654,ZNF660,ZNF664,ZNF671,ZNF674,ZNF677,ZNF678,ZNF680,ZNF687,ZNF691,ZNF692,ZNF697,ZNF700,ZNF703,ZNF707,ZNF709,ZNF70,ZNF710,ZNF713,ZNF737,ZNF740,ZNF75A,ZNF761,ZNF764,ZNF766,ZNF768,ZNF76,ZNF770,ZNF772,ZNF773,ZNF775,ZNF776,ZNF777,ZNF778,ZNF781,ZNF782,ZNF784,ZNF785,ZNF786,ZNF788,ZNF791,ZNF792,ZNF79,ZNF7,ZNF800,ZNF816,ZNF830,ZNF839,ZNF83,ZNF843,ZNF84,ZNF850,ZNF865,ZNF883,ZNF891,ZNF8,ZSCAN12,ZSCAN16,ZSCAN18,ZSCAN20,ZSCAN21,ZSCAN22,ZSCAN23,ZSCAN29,ZSCAN30,ZSCAN31,ZSCAN32,ZSCAN4,ZSCAN5A,ZSCAN5C,ZSCAN9,ZXDB,ZXDC,ZZZ3\ html TFrPeakClusters.html\ itemRgb on\ labelFields factor\ longLabel Transcription Factor Representative Peak (rPeak) Clusters (912 factors in 1152 biosamples) from ENCODE 4\ mouseOverField factor\ parent wgEncodeReg\ scoreMax 1000\ scoreMin 1\ shortLabel TF rPeak Clusters\ spectrum on\ track TFrPeakClusters\ type bigBed 12 +\ urls exp="https://www.encodeproject.org/experiments/$$/" cCRE="https://screen.wenglab.org/search/?q=$$&assembly=GRCh38" factor="https://www.factorbook.org/tf/human/$$/function"\ visibility hide\ TotalCounts_Fwd Total counts of CAGE reads (fwd) bigWig Total counts of CAGE reads forward 2 1 255 0 0 255 127 127 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/fantom5/ctssTotalCounts.fwd.bw\ color 255,0,0\ dataVersion FANTOM5 reprocessed7\ longLabel Total counts of CAGE reads forward\ parent Total_counts_multiwig\ shortLabel Total counts of CAGE reads (fwd)\ subGroups category=total strand=forward\ track TotalCounts_Fwd\ type bigWig\ cons100way UCSC 100 Vertebrates bed 4 UCSC 100 Vertebrates - 100 vertebrate genomes aligned with MultiZ by the UCSC Browser Group 2 1 0 0 0 127 127 127 0 0 0\ Downloads for data in this track are available:\
\ This track shows multiple alignments of 100 vertebrate\ species and measurements of evolutionary conservation using\ two methods (phastCons and phyloP) from the\ \ PHAST package, for all species.\ The multiple alignments were generated using multiz and\ other tools in the UCSC/Penn State Bioinformatics\ comparative genomics alignment pipeline.\ Conserved elements identified by phastCons are also displayed in\ this track.\ PHAST/Multiz are built from chains ("alignable") and nets ("syntenic"), see the documentation of the Chain/Net tracks for a description of the complete\ alignment process.\
\\ PhastCons is a hidden Markov model-based method that estimates the probability that each\ nucleotide belongs to a conserved element, based on the multiple alignment.\ It considers not just each individual alignment column, but also its\ flanking columns. By contrast, phyloP separately measures conservation at\ individual columns, ignoring the effects of their neighbors. As a\ consequence, the phyloP plots have a less smooth appearance than the\ phastCons plots, with more "texture" at individual sites. The two methods\ have different strengths and weaknesses. PhastCons is sensitive to "runs"\ of conserved sites, and is therefore effective for picking out conserved\ elements. PhyloP, on the other hand, is more appropriate for evaluating\ signatures of selection at particular nucleotides or classes of nucleotides\ (e.g., third codon positions, or first positions of miRNA target sites).\
\\ Another important difference is that phyloP can measure acceleration\ (faster evolution than expected under neutral drift) as well as\ conservation (slower than expected evolution). In the phyloP plots, sites\ predicted to be conserved are assigned positive scores (and shown in blue),\ while sites predicted to be fast-evolving are assigned negative scores (and\ shown in red). The absolute values of the scores represent -log p-values\ under a null hypothesis of neutral evolution. The phastCons scores, by\ contrast, represent probabilities of negative selection and range between 0\ and 1.\
\\ Both phastCons and phyloP treat alignment gaps and unaligned nucleotides as\ missing data, and both were run with the same parameters.\
\\ See also: lastz parameters and other details\ and chain minimum score and gap parameters used in these alignments.\
\ \\ UCSC has repeatmasked and aligned all genome assemblies, and\ provides all the sequences for download. For genome assemblies\ not available in the genome browser, there are alternative assembly hub\ genome browsers. Missing sequence in any assembly\ is highlighted in the track display by regions of yellow when\ zoomed out and by Ns when displayed at base level (see Gap Annotation, below).
\\
\ \\
\ Primate subset \ Organism Species Release date UCSC version Alignment type \ Baboon Papio hamadryas Mar 2012 Baylor Panu_2.0/papAnu2 Reciprocal best net \ Bushbaby Otolemur garnettii Mar 2011 Broad/otoGar3 Syntenic net \ Chimp Pan troglodytes Feb 2011 CSAC 2.1.4/panTro4 Syntenic net \ Crab-eating macaque Macaca fascicularis Jun 2013 Macaca_fascicularis_5.0/macFas5 Syntenic net \ Gibbon Nomascus leucogenys Oct 2012 GGSC Nleu3.0/nomLeu3 Syntenic net \ Gorilla Gorilla gorilla gorilla May 2011 gorGor3.1/gorGor3 Reciprocal best net \ Green monkey Chlorocebus sabaeus Mar 2014 Chlorocebus_sabeus 1.1/chlSab2 Syntenic net \ Human Homo sapiens Dec 2013 GRCh38/hg38 reference species \ Marmoset Callithrix jacchus Mar 2009 WUGSC 3.2/calJac3 Syntenic net \ Orangutan Pongo pygmaeus abelii July 2007 WUGSC 2.0.2/ponAbe2 Reciprocal best net \ Rhesus Macaca mulatta Oct 2010 BGI CR_1.0/rheMac3 Syntenic net \ Squirrel monkey Saimiri boliviensis Oct 2011 Broad/saiBol1 Syntenic net \ Euarchontoglires subset \ Brush-tailed rat Octodon degus Apr 2012 OctDeg1.0/octDeg1 Syntenic net \ Chinchilla Chinchilla lanigera May 2012 ChiLan1.0/chiLan1 Syntenic net \ Chinese hamster Cricetulus griseus Jul 2013 C_griseus_v1.0/criGri1 Syntenic net \ Chinese tree shrew Tupaia chinensis Jan 2013 TupChi_1.0/tupChi1 Syntenic net \ Golden hamster Mesocricetus auratus Mar 2013 MesAur1.0/mesAur1 Syntenic net \ Guinea pig Cavia porcellus Feb 2008 Broad/cavPor3 Syntenic net \ Lesser Egyptian jerboa Jaculus jaculus May 2012 JacJac1.0/jacJac1 Syntenic net \ Mouse Mus musculus Dec 2011 GRCm38/mm10 Syntenic net \ Naked mole-rat Heterocephalus glaber Jan 2012 Broad HetGla_female_1.0/hetGla2 Syntenic net \ Pika Ochotona princeps May 2012 OchPri3.0/ochPri3 Syntenic net \ Prairie vole Microtus ochrogaster Oct 2012 MicOch1.0/micOch1 Syntenic net \ Rabbit Oryctolagus cuniculus Apr 2009 Broad/oryCun2 Syntenic net \ Rat Rattus norvegicus Jul 2014 RGSC 6.0/rn6 Syntenic net \ Squirrel Spermophilus tridecemlineatus Nov 2011 Broad/speTri2 Syntenic net \ Laurasiatheria subset \ Alpaca Vicugna pacos Mar 2013 Vicugna_pacos-2.0.1/vicPac2 Syntenic net \ Bactrian camel Camelus ferus Dec 2011 CB1/camFer1 Syntenic net \ Big brown bat Eptesicus fuscus Jul 2012 EptFus1.0/eptFus1 Syntenic net \ Black flying-fox Pteropus alecto Aug 2012 ASM32557v1/pteAle1 Syntenic net \ Cat Felis catus Nov 2014 ICGSC Felis_catus 8.0/felCat8 Syntenic net \ Cow Bos taurus Jun 2014 Bos_taurus_UMD_3.1.1/bosTau8 Syntenic net \ David's myotis bat Myotis davidii Aug 2012 ASM32734v1/myoDav1 Syntenic net \ Dog Canis lupus familiaris Sep 2011 Broad CanFam3.1/canFam3 Syntenic net \ Dolphin Tursiops truncatus Oct 2011 Baylor Ttru_1.4/turTru2 Reciprocal best net \ Domestic goat Capra hircus May 2012 CHIR_1.0/capHir1 Syntenic net \ Ferret Mustela putorius furo Apr 2011 MusPutFur1.0/musFur1 Syntenic net \ Hedgehog Erinaceus europaeus May 2012 EriEur2.0/eriEur2 Syntenic net \ Horse Equus caballus Sep 2007 Broad/equCab2 Syntenic net \ Killer whale Orcinus orca Jan 2013 Oorc_1.1/orcOrc1 Syntenic net \ Megabat Pteropus vampyrus Jul 2008 Broad/pteVam1 Reciprocal best net \ Little brown bat Myotis lucifugus Jul 2010 Broad Institute Myoluc2.0/myoLuc2 Syntenic net \ Pacific walrus Odobenus rosmarus divergens Jan 2013 Oros_1.0/odoRosDiv1 Syntenic net \ Panda Ailuropoda melanoleuca Dec 2009 BGI-Shenzhen 1.0/ailMel1 Syntenic net \ Pig Sus scrofa Aug 2011 SGSC Sscrofa10.2/susScr3 Syntenic net \ Sheep Ovis aries Aug 2012 ISGC Oar_v3.1/oviAri3 Syntenic net \ Shrew Sorex araneus Aug 2008 Broad/sorAra2 Syntenic net \ Star-nosed mole Condylura cristata Mar 2012 ConCri1.0/conCri1 Syntenic net \ Tibetan antelope Pantholops hodgsonii May 2013 PHO1.0/panHod1 Syntenic net \ Weddell seal Leptonychotes weddellii Mar 2013 LepWed1.0/lepWed1 Reciprocal best net \ White rhinoceros Ceratotherium simum May 2012 CerSimSim1.0/cerSim1 Syntenic net \ Afrotheria subset \ Aardvark Orycteropus afer afer May 2012 OryAfe1.0/oryAfe1 Syntenic net \ Cape elephant shrew Elephantulus edwardii Aug 2012 EleEdw1.0/eleEdw1 Syntenic net \ Cape golden mole Chrysochloris asiatica Aug 2012 ChrAsi1.0/chrAsi1 Syntenic net \ Elephant Loxodonta africana Jul 2009 Broad/loxAfr3 Syntenic net \ Manatee Trichechus manatus latirostris Oct 2011 Broad v1.0/triMan1 Syntenic net \ Tenrec Echinops telfairi Nov 2012 Broad/echTel2 Syntenic net \ Mammal subset \ Armadillo Dasypus novemcinctus Dec 2011 Baylor/dasNov3 Syntenic net \ Opossum Monodelphis domestica Oct 2006 Broad/monDom5 Net \ Platypus Ornithorhynchus anatinus Mar 2007 WUGSC 5.0.1/ornAna1 Reciprocal best net \ Tasmanian devil Sarcophilus harrisii Feb 2011 WTSI Devil_ref v7.0/sarHar1 Net \ Wallaby Macropus eugenii Sep 2009 TWGS Meug_1.1/macEug2 Reciprocal best net \ Aves subset \ Budgerigar Melopsittacus undulatus Sep 2011 WUSTL v6.3/melUnd1 Net \ Chicken Gallus gallus Nov 2011 ICGSC Gallus_gallus-4.0/galGal4 Net \ Collared flycatcher Ficedula albicollis Jun 2013 FicAlb1.5/ficAlb2 Net \ Mallard duck Anas platyrhynchos Apr 2013 BGI_duck_1.0/anaPla1 Net \ Medium ground finch Geospiza fortis Apr 2012 GeoFor_1.0/geoFor1 Net \ Parrot Amazona vittata Jan 2013 AV1/amaVit1 Net \ Peregrine falcon Falco peregrinus Feb 2013 F_peregrinus_v1.0/falPer1 Net \ Rock pigeon Columba livia Feb 2013 Cliv_1.0/colLiv1 Net \ Saker falcon Falco cherrug Feb 2013 F_cherrug_v1.0/falChe1 Net \ Scarlet macaw Ara macao Jun 2013 SMACv1.1/araMac1 Net \ Tibetan ground jay Pseudopodoces humilis Jan 2013 PseHum1.0/pseHum1 Net \ Turkey Meleagris gallopavo Dec 2009 TGC Turkey_2.01/melGal1 Net \ White-throated sparrow Zonotrichia albicollis Apr 2013 ASM38545v1/zonAlb1 Net \ Zebra finch Taeniopygia guttata Feb 2013 WashU taeGut324/taeGut2 Net \ Sarcopterygii subset \ American alligator Alligator mississippiensis Aug 2012 allMis0.2/allMis1 Net \ Chinese softshell turtle Pelodiscus sinensis Oct 2011 PelSin_1.0/pelSin1 Net \ Coelacanth Latimeria chalumnae Aug 2011 Broad/latCha1 Net \ Green seaturtle Chelonia mydas Mar 2013 CheMyd_1.0/cheMyd1 Net \ Lizard Anolis carolinensis May 2010 Broad AnoCar2.0/anoCar2 Net \ Painted turtle Chrysemys picta bellii Mar 2014 v3.0.3/chrPic2 Net \ Spiny softshell turtle Apalone spinifera May 2013 ASM38561v1/apaSpi1 Net \ X. tropicalis Xenopus tropicalis Sep 2012 JGI 7.0/xenTro7 Net \ Fish subset \ Atlantic cod Gadus morhua May 2010 Genofisk GadMor_May2010/gadMor1 Net \ Burton's mouthbreeder Haplochromis burtoni Oct 2011 AstBur1.0/hapBur1 Net \ Fugu Takifugu rubripes Oct 2011 FUGU5/fr3 Net \ Lamprey Petromyzon marinus Sep 2010 WUGSC 7.0/petMar2 Net \ Medaka Oryzias latipes Oct 2005 NIG/UT MEDAKA1/oryLat2 Net \ Mexican tetra (cavefish) Astyanax mexicanus Apr 2013 Astyanax_mexicanus-1.0.2/astMex1 Net \ Nile tilapia Oreochromis niloticus Jan 2011 Broad oreNil1.1/oreNil2 Net \ Princess of Burundi Neolamprologus brichardi May 2011 NeoBri1.0/neoBri1 Net \ Pundamilia nyererei Pundamilia nyererei Oct 2011 PunNye1.0/punNye1 Net \ Southern platyfish Xiphophorus maculatus Jan 2012 Xiphophorus_maculatus-4.4.2/xipMac1 Net \ Spotted gar Lepisosteus oculatus Dec 2011 LepOcu1/lepOcu1 Net \ Stickleback Gasterosteus aculeatus Feb 2006 Broad/gasAcu1 Net \ Tetraodon Tetraodon nigroviridis Mar 2007 Genoscope 8.0/tetNig2 Net \ Yellowbelly pufferfish Takifugu flavidus May 2013 version 1 of Takifugu flavidus genome/takFla1 Net \ Zebra mbuna Maylandia zebra Mar 2012 MetZeb1.1/mayZeb1 Net \ Zebrafish Danio rerio Sep 2014 GRCz10/danRer10 Net
\ Table 1. Genome assemblies included in the 100-way Conservation track.
\
\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ size of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ The names of selected species are colored according to their clade,\ alternating between blue and green.\ Note that excluding species from the pairwise display does not alter the\ conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. The following\ conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment. The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\\ Codon translation is available in base-level display mode if the\ displayed region is identified as a coding segment. To display this annotation,\ select the species for translation from the pull-down menu in the Codon\ Translation configuration section at the top of the page. Then, select one of\ the following modes:\
\ Codon translation uses the following gene tracks as the basis for translation:\ \
\ \\
\ Table 2. Gene tracks used for codon translation.\\ Gene Track Species \ UCSC Genes Human, Mouse \ RefSeq Genes Cow, Frog (X. tropicalis) \ Ensembl Genes v73 Atlantic cod, Bushbaby, Cat, Chicken, Chimp, Coelacanth, Dog, Elephant, Ferret, Fugu, Gorilla, Horse, Lamprey, Little brown bat, Lizard, Mallard duck, Marmoset, Medaka, Megabat, Orangutan, Panda, Pig, Platypus, Rat, Soft-shell Turtle, Southern platyfish, Squirrel, Tasmanian devil, Tetraodon, Zebrafish \ no annotation Aardvark, Alpaca, American alligator, Armadillo, Baboon, Bactrian camel, Big brown bat, Black flying-fox, Brush-tailed rat, Budgerigar, Burton's mouthbreeder, Cape elephant shrew, Cape golden mole, Chinchilla, Chinese hamster, Chinese tree shrew, Collared flycatcher, Crab-eating macaque, David's myotis (bat), Dolphin, Domestic goat, Gibbon, Golden hamster, Green monkey, Green seaturtle, Hedgehog, Killer whale, Lesser Egyptian jerboa, Manatee, Medium ground finch, Mexican tetra (cavefish), Naked mole-rat, Nile tilapia, Pacific walrus, Painted turtle, Parrot, Peregrine falcon, Pika, Prairie vole, Princess of Burundi, Pundamilia nyererei, Rhesus, Rock pigeon, Saker falcon, Scarlet Macaw, Sheep, Shrew, Spiny softshell turtle, Spotted gar, Squirrel monkey, Star-nosed mole, Tawny puffer fish, Tenrec, Tibetan antelope, Tibetan ground jay, Wallaby, Weddell seal, White rhinoceros, White-throated sparrow, Zebra Mbuna, Zebra finch
\ Pairwise alignments with the human genome were generated for\ each species using lastz from repeat-masked genomic sequence.\ Pairwise alignments were then linked into chains using a dynamic programming\ algorithm that finds maximally scoring chains of gapless subsections\ of the alignments organized in a kd-tree.\ The scoring matrix and parameters for pairwise alignment and chaining\ were tuned for each species based on phylogenetic distance from the reference.\ High-scoring chains were then placed along the genome, with\ gaps filled by lower-scoring chains, to produce an alignment net.\ For more information about the chaining and netting process and\ parameters for each species, see the description pages for the Chain and Net\ tracks.
\\ An additional filtering step was introduced in the generation of the 100-way\ conservation track to reduce the number of paralogs and pseudogenes from the\ high-quality assemblies and the suspect alignments from the low-quality\ assemblies:\ the pairwise alignments of high-quality mammalian\ sequences (placental and marsupial) were filtered based on synteny;\ those for 2X mammalian genomes were filtered to retain only\ alignments of best quality in both the target and query ("reciprocal\ best").
\\ The resulting best-in-genome pairwise alignments\ were progressively aligned using multiz/autoMZ,\ following the tree topology diagrammed above, to produce multiple alignments.\ The multiple alignments were post-processed to\ add annotations indicating alignment gaps, genomic breaks,\ and base quality of the component sequences.\ The annotated multiple alignments, in MAF format, are available for\ bulk download.\ An alignment summary table containing an entry for each\ alignment block in each species was generated to improve\ track display performance at large scales.\ Framing tables were constructed to enable\ visualization of codons in the multiple alignment display.
\ \\ Both phastCons and phyloP are phylogenetic methods that rely\ on a tree model containing the tree topology, branch lengths representing\ evolutionary distance at neutrally evolving sites, the background distribution\ of nucleotides, and a substitution rate matrix.\ The\ all-species tree model for this track was\ generated using the phyloFit program from the PHAST package\ (REV model, EM algorithm, medium precision) using multiple alignments of\ 4-fold degenerate sites extracted from the 100-way alignment\ (msa_view). The 4d sites were derived from the RefSeq (Reviewed+Coding) gene\ set, filtered to select single-coverage long transcripts.\
\\ This same tree model was used in the phyloP calculations; however, the\ background frequencies were modified to maintain reversibility.\ The resulting tree model:\ all species.\
\\ The phastCons program computes conservation scores based on a phylo-HMM, a\ type of probabilistic model that describes both the process of DNA\ substitution at each site in a genome and the way this process changes from\ one site to the next (Felsenstein and Churchill 1996, Yang 1995, Siepel and\ Haussler 2005). PhastCons uses a two-state phylo-HMM, with a state for\ conserved regions and a state for non-conserved regions. The value plotted\ at each site is the posterior probability that the corresponding alignment\ column was "generated" by the conserved state of the phylo-HMM. These\ scores reflect the phylogeny (including branch lengths) of the species in\ question, a continuous-time Markov model of the nucleotide substitution\ process, and a tendency for conservation levels to be autocorrelated along\ the genome (i.e., to be similar at adjacent sites). The general reversible\ (REV) substitution model was used. Unlike many conservation-scoring programs,\ phastCons does not rely on a sliding window\ of fixed size; therefore, short highly-conserved regions and long moderately\ conserved regions can both obtain high scores.\ More information about\ phastCons can be found in Siepel et al. 2005.
\\ The phastCons parameters used were: expected-length=45,\ target-coverage=0.3, rho=0.3.
\ \\ The phyloP program supports several different methods for computing\ p-values of conservation or acceleration, for individual nucleotides or\ larger elements (\ http://compgen.cshl.edu/phast/). Here it was used\ to produce separate scores at each base (--wig-scores option), considering\ all branches of the phylogeny rather than a particular subtree or lineage\ (i.e., the --subtree option was not used). The scores were computed by\ performing a likelihood ratio test at each alignment column (--method LRT),\ and scores for both conservation and acceleration were produced (--mode\ CONACC).\
\\ The conserved elements were predicted by running phastCons with the\ --viterbi option. The predicted elements are segments of the alignment\ that are likely to have been "generated" by the conserved state of the\ phylo-HMM. Each element is assigned a log-odds score equal to its log\ probability under the conserved model minus its log probability under the\ non-conserved model. The "score" field associated with this track contains\ transformed log-odds scores, taking values between 0 and 1000. (The scores\ are transformed using a monotonic function of the form a * log(x) + b.) The\ raw log odds scores are retained in the "name" field and can be seen on the\ details page or in the browser when the track's display mode is set to\ "pack" or "full".\
\ \This track was created using the following programs:\
The phylogenetic tree is based on Murphy et al. (2001) and general\ consensus in the vertebrate phylogeny community. Thanks to Giacomo Bernardi for\ help with the fish relationships.\
\ \\ Felsenstein J, Churchill GA.\ A Hidden Markov Model approach to variation among sites in rate of\ evolution. Mol Biol Evol. 1996 Jan;13(1):93-104.\ PMID: 8583911\
\ \\ Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A.\ \ Detection of nonneutral substitution rates on mammalian phylogenies.\ Genome Res. 2010 Jan;20(1):110-21.\ PMID: 19858363; PMC: PMC2798823\
\ \\ Siepel A, Bejerano G, Pedersen JS, Hinrichs AS, Hou M, Rosenbloom K,\ Clawson H, Spieth J, Hillier LW, Richards S, et al.\ Evolutionarily conserved elements in vertebrate, insect, worm,\ and yeast genomes.\ Genome Res. 2005 Aug;15(8):1034-50.\ PMID: 16024819; PMC: PMC1182216\
\ \\ Siepel A, Haussler D.\ Phylogenetic Hidden Markov Models.\ In: Nielsen R, editor. Statistical Methods in Molecular Evolution.\ New York: Springer; 2005. pp. 325-351.\ DOI: 10.1007/0-387-27733-1_12\
\ \\ Yang Z.\ A space-time process model for the evolution of DNA\ sequences.\ Genetics. 1995 Feb;139(2):993-1005.\ PMID: 7713447; PMC: PMC1206396\
\ \\ Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D.\ Evolution's cauldron:\ duplication, deletion, and rearrangement in the mouse and human genomes.\ Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9.\ PMID: 14500911; PMC: PMC208784\
\ \\ Blanchette M, Kent WJ, Riemer C, Elnitski L, Smit AF, Roskin KM,\ Baertsch R, Rosenbloom K, Clawson H, Green ED, et al.\ Aligning multiple genomic sequences with the threaded blockset aligner.\ Genome Res. 2004 Apr;14(4):708-15.\ PMID: 15060014; PMC: PMC383317\
\ \\ Chiaromonte F, Yap VB, Miller W.\ Scoring pairwise genomic sequence alignments.\ Pac Symp Biocomput. 2002:115-26.\ PMID: 11928468\
\ \\ Harris RS.\ Improved pairwise alignment of genomic DNA.\ Ph.D. Thesis. Pennsylvania State University, USA. 2007.\
\ \\ Schwartz S, Kent WJ, Smit A, Zhang Z, Baertsch R, Hardison RC,\ Haussler D, Miller W.\ Human-mouse alignments with BLASTZ.\ Genome Res. 2003 Jan;13(1):103-7.\ PMID: 12529312; PMC: PMC430961\
\ \ \\ Murphy WJ, Eizirik E, O'Brien SJ, Madsen O, Scally M, Douady CJ, Teeling E,\ Ryder OA, Stanhope MJ, de Jong WW, Springer MS.\ Resolution of the early placental mammal radiation using Bayesian phylogenetics.\ Science. 2001 Dec 14;294(5550):2348-51.\ PMID: 11743200\
\ compGeno 1 compositeTrack on\ dragAndDrop subTracks\ group compGeno\ longLabel UCSC 100 Vertebrates - 100 vertebrate genomes aligned with MultiZ by the UCSC Browser Group\ priority 1\ shortLabel UCSC 100 Vertebrates\ subGroup1 view Views align=Multiz_Alignments phyloP=Basewise_Conservation_(phyloP) phastcons=Element_Conservation_(phastCons) elements=Conserved_Elements\ track cons100way\ type bed 4\ visibility full\ liftOverHg19 UCSC liftOver to hg19 chain hg19 UCSC liftOver alignments to hg19 0 1 0 0 0 127 127 127 0 0 0 map 1 chainLinearGap medium\ chainMinScore 3000\ longLabel UCSC liftOver alignments to hg19\ matrix 16 91,-114,-31,-123,-114,100,-125,-31,-31,-125,100,-114,-123,-31,-114,91\ matrixHeader A, C, G, T\ otherDb hg19\ parent liftHg19\ priority 1\ shortLabel UCSC liftOver to hg19\ track liftOverHg19\ type chain hg19\ comments UCSC Unusual Regions bigBed 9 + UCSC unusual regions on assembly structure (manually annotated) 3 1 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/comments.bb\ longLabel UCSC unusual regions on assembly structure (manually annotated)\ mouseOverField note\ noScoreFilter on\ parent problematic\ priority 1\ searchIndex name\ searchTrix /gbdb/hg38/problematic/notes.ix\ shortLabel UCSC Unusual Regions\ track comments\ type bigBed 9 +\ umap24 Umap S24 bigBed 6 Single-read mappability with 24-mers 4 1 80 20 240 167 137 247 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k24.Unique.Mappability.bb\ color 80,20,240\ longLabel Single-read mappability with 24-mers\ parent umapBigBed on\ priority 1\ shortLabel Umap S24\ subGroups view=SR\ track umap24\ wgEncodeReg4Dnase DNase (Layered) bigWig Chromatin accessibility from DNase-seq signal, averaged by organ/tissue 0 1.1 0 0 0 127 127 127 0 0 0\ DNase I hypersensitivity identifies regions of open chromatin, which are often associated with\ regulatory elements such as promoters, enhancers, and insulators. This track displays\ genome-wide DNase I hypersensitivity signal, as determined by ENCODE DNase-seq data across all\ phases of the project. Higher signal intensity indicates greater chromatin accessibility,\ suggesting potential regulatory activity. Regulatory elements -- especially promoters -- tend\ to be strongly DNase-sensitive. The data are processed following the\ ENCODE\ DNase-seq pipeline. Additional chromatin accessibility and transcription factor binding\ datasets are available at the\ ENCODE portal.
\ \\ For each organ, this track provides up to two subtracks averaging DNase signal:
\\ Whether one or two subtracks appear for an organ depends on which kinds of biosamples have\ been assayed:
\| Organ/Tissue | \Tissue and Primary Cell Subtrack | \All Biosamples Subtrack | \
|---|---|---|
| adipose | – | ✓ |
| adrenal gland | – | ✓ |
| blood | ✓ | ✓ |
| blood vessel | – | ✓ |
| bone | ✓ | ✓ |
| bone marrow | ✓ | ✓ |
| brain | ✓ | ✓ |
| breast | ✓ | ✓ |
| connective tissue | ✓ | ✓ |
| embryo | ✓ | ✓ |
| epithelium | ✓ | ✓ |
| esophagus | – | ✓ |
| eye | ✓ | ✓ |
| gallbladder | – | ✓ |
| heart | ✓ | ✓ |
| kidney | ✓ | ✓ |
| large intestine | ✓ | ✓ |
| limb | – | ✓ |
| liver | ✓ | ✓ |
| lung | ✓ | ✓ |
| lymphoid tissue | ✓ | – |
| mouth | ✓ | ✓ |
| muscle | ✓ | ✓ |
| nerve | – | ✓ |
| nose | – | ✓ |
| ovary | – | ✓ |
| pancreas | ✓ | ✓ |
| penis | ✓ | ✓ |
| placenta | ✓ | ✓ |
| prostate | ✓ | ✓ |
| skin | ✓ | ✓ |
| small intestine | – | ✓ |
| spinal cord | – | ✓ |
| spleen | – | ✓ |
| stomach | – | ✓ |
| testis | ✓ | ✓ |
| thymus | – | ✓ |
| thyroid | – | ✓ |
| urinary bladder | – | ✓ |
| uterus | ✓ | ✓ |
| vagina | – | ✓ |
\ By default, this track uses a transparent overlay to visualize data from multiple organs or tissues within\ the same vertical space. For each organ or tissue, signals from all associated experiments were\ averaged to generate the displayed track. Each organ or tissue is assigned a distinct\ color following the\ ENCODE color\ mapping convention,\ selected to be light and saturated to maintain clarity when overlaid. Initially, each layered\ track displays an overlay of five representative organs: blood, brain, kidney, liver, and\ muscle. Clicking on the track opens a details page where you can view and select organs or\ tissues.
\ \ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/regulation/organAve/adiposeDNase.bw stdout
\ Data were generated by the ENCODE Consortium. We thank the production labs for generating the\ data: Drs. Gregory Crawford (Duke) and John Stamatoyannopoulos (UW). The data were further\ processed for visualization through a collaborative effort between the\ Weng lab and the\ Moore lab\ at UMass Chan Medical School (funded by NIH grant HG012343). Integration and visualization\ were developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 aggregate transparentOverlay\ allButtonPair on\ autoScale on\ container multiWig\ dragAndDrop subtracks\ html wgEncodeReg4Dnase.html\ longLabel Chromatin accessibility from DNase-seq signal, averaged by organ/tissue\ maxHeightPixels 100:50:11\ noInherit on\ priority 1.1\ shortLabel DNase (Layered)\ showSubtrackColorOnUi on\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4Dnase\ type bigWig\ viewLimits 0:50\ visibility hide\ lincRNAsAllCellTypeTopView lincRNA RNA-Seq bed 5 + lincRNA RNA-Seq reads expression abundances 1 1.1 0 0 0 127 127 127 0 0 0This track displays the Human Body Map lincRNAs (large intergenic non\ coding RNAs) and TUCPs (transcripts of uncertain coding potential), as well as their\ expression levels across 22 human tissues and cell lines. The Human Body Map catalog was generated\ by integrating previously existing annotation sources with transcripts that were de-novo assembled\ from RNA-Seq data. These transcripts were collected from ~4 billion RNA-Seq reads across 24 tissues \ and cell types.
\ \Expression abundance was estimated by Cufflinks (Trapnell et al., 2010) based on RNA-Seq. \ Expression abundances were estimated on the gene locus level, rather than for each transcript \ separately and are given as raw FPKM. The prefixes tcons_ and tcons_l2_ are used to describe \ lincRNAs and TUCP transcripts, respectively. Specific details about the catalog generation and data \ sets used for this study can be found in Cabili et al (2011). Extended \ characterization of each transcript in the human body map catalog can be found at the Human lincRNA\ Catalog website.
\ \Expression abundance scores range from 0 to 1000, and are displayed from light blue to dark blue\ respectively:
\ \ \01000
\ \The body map RNA-Seq data was kindly provided by the Gene Expression\ Applications research group at Illumina.
\ \\ Cabili MN, Trapnell C, Goff L, Koziol M, Tazon-Vega B, Regev A, Rinn JL.\ \ Integrative annotation of human large intergenic noncoding RNAs reveals global properties and\ specific subclasses.\ Genes Dev. 2011 Sep 15;25(18):1915-27.\ PMID: 21890647; PMC: PMC3185964\
\ \\ Trapnell C, Williams BA, Pertea G, Mortazavi A, Kwan G, van Baren MJ, Salzberg SL, Wold BJ, Pachter\ L.\ \ Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform\ switching during cell differentiation.\ Nat Biotechnol. 2010 May;28(5):511-5.\ PMID: 20436464; PMC: PMC3146043\
\ genes 1 compositeTrack on\ configurable on\ dimensions dimensionY=tissueType\ dragAndDrop subTracks\ html lincRNAs\ longLabel lincRNA RNA-Seq reads expression abundances\ noInherit on\ onlyVisibility dense\ origAssembly hg19\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 1.1\ shortLabel lincRNA RNA-Seq\ sortOrder view=+ tissueType=+\ subGroup1 view Views lincRNAsRefseqExp=RefSeq_Expression_Ratio\ subGroup2 tissueType Tissue_Type adipose=Adipose adrenal=Adrenal brain=Brain brain_r=Brain_R breast=Breast colon=Colon foreskin_r=Foreskin_R heart=Heart hlf_r1=hLF_r1 hlf_r2=hLF_r2 kidney=Kidney liver=Liver lung=Lung lymphnode=LymphNode ovary=Ovary placenta_r=Placenta_R prostate=Prostate skeletalmuscle=SkeletalMuscle testes=Testes testes_r=Testes_R thyroid=Thyroid whitebloodcell=WhiteBloodCell\ superTrack nonCodingRNAs dense\ track lincRNAsAllCellTypeTopView\ type bed 5 +\ lincRNAsAllCellType lincRNAsCellType bed 5 + lincRNA RNA-Seq reads expression abundances 1 1.1 0 60 120 127 157 187 1 0 0 genes 1 color 0, 60, 120\ longLabel lincRNA RNA-Seq reads expression abundances\ origAssembly hg19\ parent lincRNAsAllCellTypeTopView\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel lincRNAsCellType\ track lincRNAsAllCellType\ useScore 1\ view lincRNAsRefseqExp\ visibility dense\ spliceAIindels SpliceAI indels bigBed 9 + SpliceAI Indels (unmasked) 1 1.1 0 0 0 127 127 127 0 0 0\ SpliceAI is an open-source deep\ learning algorithm that predicts splicing probability for nucleotides and \ as a result can score DNA variants for splicing impact.\ Such variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a \ lookup tool \ is provided by the Broad institute.\
\ \The spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added and the new sequence is\ scored again. The "wildtype" container track shows the scores for the genome\ sequence itself and the "variants" container track shows the impact of all\ possible variants close to known splice sites. The "wildtype" subtracks are\ useful when looking at new transcript models, to evaluate how likely exon\ boundaries are. The "variants" subtracks are used to evaluate the impact of\ variants onto splicing, typically in medical diagnostics.\
\ \\ SpliceAI only annotates variants close to splice sites of genes defined by the \ Gencode gene annotation track. Additionally, SpliceAI does not annotate variants if they are\ close to chromosome ends (5kb on either side), deletions of length greater than\ twice the input parameter -D, or inconsistent with the reference fasta file.\
\ \\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \\ Variants are colored according to Walker et al. 2023 splicing imact:\
\\ The scores range from 0 to 1 and can be interpreted as the \ probability of the variant being splice-altering. In the paper, a detailed characterization is \ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\
The data were downloaded from Illumina. \
The spliceAI scores are represented in the VCF INFO field as \
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain \
\ Since most of the values are 0 or almost 0, we selected only those variants \ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \ \\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are \ made available by Illumina only for academic or not-for-profit research only. \ By accessing the SpliceAI data, you acknowledge and agree that you may only \ use this data for your own personal academic or not-for-profit research only, \ and not for any other purposes. You may not use this data for any for-profit, \ clinical, or other commercial purpose without obtaining a commercial license \ from Illumina, Inc.\
\ \\ Thanks to Illumina for making the data available. Thanks to Michael Hiller, Francois Lecoquierre and\ Jean-Madeleine de Sainte Agathe for making available and suggesting the SpliceAI wildtype tracks.\
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/spliceAIindels.bb\ filter.AIscore 0.02\ filterLabel.spliceType Splice type\ filterLimits.AIscore 0.02:1\ filterValues.spliceType donor_gain|Donor gain,donor_loss|Donor loss,acceptor_gain|Acceptor gain,acceptor_loss|Acceptor loss\ html spliceAI\ itemRgb on\ longLabel SpliceAI Indels (unmasked)\ mouseOver Change: $name\ This track shows transcription levels for several cell types as assayed by high-throughput\ sequencing of polyadenylated RNA (RNA-seq).\ Additional views of this dataset and additional documentation on the methods used\ for this track are available at the\ ENCODE Caltech RNA-seq\ page. The data shown here are derived from the Raw Signal view from the paired \ 75-mer 200 bp insert size reads. The two replicates of the signal were pooled and normalized\ so that the total genome-wide signal sums to 10 billion.\
\ \\ By default, this track uses a transparent overlay method of displaying data from a number of cell\ lines in the same vertical space. Each of the cell lines in this track\ is associated with a particular color, and these colors are relatively light and saturated so\ as to work best with the transparent overlay. The color of these tracks\ match their versions from their lifted source on the hg19 assembly. The colors are consistent with the\ other hg19 lifted tracks located in the ENCODE Regulation\ supertrack, with the exception being the DNase tracks, as they were not lifted from hg19 and are\ colored to reflect similarity of cell types.\
\ \\ This track shows data from the\ Wold Lab at Caltech,\ as part of the ENCODE Consortium. \
\ \\ This is release 2 (July 2012) of this track which includes two new subtracks for HeLa-S3 and HepG2.\
\ \\ Primary ENCODE data produced during the 2007-2012 production phase were subject to a restriction\ period. However, the data here are past those restrictions and are freely available.\ The full data release policy for ENCODE is available\ here.\
\ regulation 1 aggregate transparentOverlay\ allButtonPair on\ container multiWig\ dragAndDrop subTracks\ longLabel Transcription Levels Assayed by RNA-seq on 9 Cell Lines from ENCODE\ maxHeightPixels 100:30:11\ noInherit on\ origAssembly hg19\ parent wgEncodeReg\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 1.1\ shortLabel Transcription\ showSubtrackColorOnUi on\ track wgEncodeRegTxn\ transformFunc LOG\ type bigWig 0 65500\ viewLimits 0:8\ visibility hide\ wgEncodeReg4Atac ATAC (Layered) bigWig Chromatin accessibility from ATAC-seq signal, averaged by organ/tissue 0 1.2 0 0 0 127 127 127 0 0 0\ The Assay for Transposase-Accessible Chromatin using sequencing (ATAC-seq) identifies regions\ of open chromatin, which are often associated with active regulatory elements such as promoters\ and enhancers. This track displays genome-wide chromatin accessibility signal, as determined by\ ENCODE ATAC-seq data across all phases of the project. Higher signal intensity indicates\ increased chromatin accessibility, suggesting potential regulatory activity. The data are\ processed following the\ ENCODE\ ATAC-seq pipeline. Additional chromatin accessibility and transcription factor binding\ datasets are available at the\ ENCODE portal.
\ \\ For each organ, this track provides up to two subtracks averaging ATAC signal:
\\ Whether one or two subtracks appear for an organ depends on which kinds of biosamples have\ been assayed:
\| Organ/Tissue | \Tissue and Primary Cell Subtrack | \All Biosamples Subtrack | \
|---|---|---|
| adipose | – | ✓ |
| adrenal gland | – | ✓ |
| blood | ✓ | ✓ |
| blood vessel | – | ✓ |
| bone marrow | – | ✓ |
| brain | ✓ | ✓ |
| breast | ✓ | ✓ |
| esophagus | – | ✓ |
| gallbladder | – | ✓ |
| heart | – | ✓ |
| large intestine | ✓ | ✓ |
| liver | ✓ | ✓ |
| lung | ✓ | ✓ |
| muscle | – | ✓ |
| nerve | – | ✓ |
| ovary | – | ✓ |
| pancreas | ✓ | ✓ |
| prostate | ✓ | ✓ |
| skin | – | ✓ |
| small intestine | – | ✓ |
| spleen | – | ✓ |
| stomach | – | ✓ |
| testis | – | ✓ |
| thyroid | – | ✓ |
| urinary bladder | – | ✓ |
| uterus | – | ✓ |
\ By default, this track uses a transparent overlay to visualize data from multiple organs or tissues within\ the same vertical space. For each organ or tissue, signals from all associated experiments were\ averaged to generate the displayed track. Each organ or tissue is assigned a distinct\ color following the\ ENCODE color\ mapping convention,\ selected to be light and saturated to maintain clarity when overlaid. Initially, each layered\ track displays an overlay of five representative organs: blood, brain, kidney, liver, and\ muscle. Clicking on the track opens a details page where you can view and select organs or\ tissues.
\ \ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/regulation/organAve/adiposeATAC.bw stdout
\ Data were generated by the ENCODE Consortium. We thank the production labs for generating the\ data: Drs. Barbara Wold (Caltech), Michael Snyder, Stephen Montgomery,\ and Will Greenleaf (Stanford), and Yin Shen (UCSF). The data were further processed for\ visualization through a\ collaborative effort between the\ Weng lab and the\ Moore lab\ at UMass Chan Medical School (funded by NIH grant HG012343). Integration and visualization\ were developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 aggregate transparentOverlay\ allButtonPair on\ autoScale on\ container multiWig\ dragAndDrop subtracks\ html wgEncodeReg4Atac.html\ longLabel Chromatin accessibility from ATAC-seq signal, averaged by organ/tissue\ maxHeightPixels 100:50:11\ noInherit on\ priority 1.2\ shortLabel ATAC (Layered)\ showSubtrackColorOnUi on\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4Atac\ type bigWig\ viewLimits 0:50\ visibility hide\ wgEncodeRegMarkH3k4me1 Layered H3K4Me1 bigWig 0 10000 H3K4Me1 Mark (Often Found Near Regulatory Elements) on 7 cell lines from ENCODE 0 1.2 0 0 0 127 127 127 0 0 0\ Chemical modifications (e.g., methylation and acetylation) to the histone proteins\ present in chromatin influence gene expression by changing how\ accessible the chromatin is to transcription. A specific modification of\ a specific histone protein is called a histone mark.\ This track shows the levels of enrichment of the H3K4Me1 histone mark across the genome as\ determined by a ChIP-seq assay. The H3K4me1 histone mark is the mono-methylation of lysine 4\ of the H3 histone protein, and it is associated with enhancers and with DNA regions downstream of\ transcription starts. Additional histone marks and other chromatin associated ChIP-seq data is\ available at the\ Broad Histone page.\
\ \\ By default, this track uses a transparent overlay method of displaying data from a number of cell\ lines in the same vertical space. Each of the cell lines in this track\ is associated with a particular color, and these colors are relatively light and saturated so\ as to work best with the transparent overlay. The color of these tracks\ match their versions from their lifted source on the hg19 assembly. The colors are consistent with the\ other hg19 lifted tracks located in the ENCODE Regulation\ supertrack, with the exception being the DNase tracks, as they were not lifted from hg19 and are\ colored to reflect similarity of cell types.\
\ \\ This track shows data from the Bernstein Lab at the Broad Institute, as part of\ the ENCODE Consortium.\
\ \\ Primary ENCODE data produced during the 2007-2012 production phase were subject to a restriction\ period. However, the data here are past those restrictions and are freely available.\ The full data release policy for ENCODE is available\ here.\
\ regulation 1 aggregate transparentOverlay\ allButtonPair on\ container multiWig\ dragAndDrop subtracks\ longLabel H3K4Me1 Mark (Often Found Near Regulatory Elements) on 7 cell lines from ENCODE\ maxHeightPixels 100:30:11\ noInherit on\ origAssembly hg19\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 1.2\ shortLabel Layered H3K4Me1\ showSubtrackColorOnUi on\ superTrack wgEncodeReg hide\ track wgEncodeRegMarkH3k4me1\ type bigWig 0 10000\ viewLimits 0:50\ visibility hide\ spliceAIsnvsMasked SpliceAI SNVs (masked) bigBed 9 + SpliceAI SNVs (masked) 1 1.2 0 0 0 127 127 127 0 0 0\ SpliceAI is an open-source deep\ learning algorithm that predicts splicing probability for nucleotides and \ as a result can score DNA variants for splicing impact.\ Such variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a \ lookup tool \ is provided by the Broad institute.\
\ \The spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added and the new sequence is\ scored again. The "wildtype" container track shows the scores for the genome\ sequence itself and the "variants" container track shows the impact of all\ possible variants close to known splice sites. The "wildtype" subtracks are\ useful when looking at new transcript models, to evaluate how likely exon\ boundaries are. The "variants" subtracks are used to evaluate the impact of\ variants onto splicing, typically in medical diagnostics.\
\ \\ SpliceAI only annotates variants close to splice sites of genes defined by the \ Gencode gene annotation track. Additionally, SpliceAI does not annotate variants if they are\ close to chromosome ends (5kb on either side), deletions of length greater than\ twice the input parameter -D, or inconsistent with the reference fasta file.\
\ \\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \\ Variants are colored according to Walker et al. 2023 splicing imact:\
\\ The scores range from 0 to 1 and can be interpreted as the \ probability of the variant being splice-altering. In the paper, a detailed characterization is \ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\
The data were downloaded from Illumina. \
The spliceAI scores are represented in the VCF INFO field as \
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain \
\ Since most of the values are 0 or almost 0, we selected only those variants \ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \ \\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are \ made available by Illumina only for academic or not-for-profit research only. \ By accessing the SpliceAI data, you acknowledge and agree that you may only \ use this data for your own personal academic or not-for-profit research only, \ and not for any other purposes. You may not use this data for any for-profit, \ clinical, or other commercial purpose without obtaining a commercial license \ from Illumina, Inc.\
\ \\ Thanks to Illumina for making the data available. Thanks to Michael Hiller, Francois Lecoquierre and\ Jean-Madeleine de Sainte Agathe for making available and suggesting the SpliceAI wildtype tracks.\
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/spliceAIsnvsMasked.bb\ filter.AIscore 0.02\ filterLabel.spliceType Splice type\ filterLimits.AIscore 0.02:1\ filterValues.spliceType donor_gain|Donor gain,donor_loss|Donor loss,acceptor_gain|Acceptor gain,acceptor_loss|Acceptor loss\ html spliceAI\ itemRgb on\ longLabel SpliceAI SNVs (masked)\ mouseOver Change: $name\ The FANTOM5 track shows mapped transcription start sites (TSS) and their usage in primary cells,\ cell lines, and tissues to produce a comprehensive overview of gene expression across the human\ body by using single molecule sequencing.\
\ \Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \Individual biological states are profiled by HeliScopeCAGE, which is a variation of the CAGE\ (Cap Analysis Gene Expression) protocol based on a single molecule sequencer. The standard protocol\ requiring 5 µg of total RNA as a starting material is referred to as hCAGE, and an\ optimized version for a lower quantity (~ 100 ng) is referred to as LQhCAGE (Kanamori-Katyama\ et al. 2011).\
Transcription start sites (TSSs) were mapped and their usage in human and mouse primary cells,\ cell lines, and tissues was to produce a comprehensive overview of mammalian gene expression across the\ human body. 5′-end of the mapped CAGE reads are counted at a single base pair resolution\ (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the\ sample. Individual samples shown in "TSS activity" tracks are grouped as below.\
TSS (CAGE) peaks across the panel of the biological states (samples) are identified by DPI\ (decomposition based peak identification, Forrest et al. 2014), where each of the peaks consists of\ neighboring and related TSSs. The peaks are used as anchors to define promoters and units of\ promoter-level expression analysis. Two subsets of the peaks are defined based on evidence of read\ counts, depending on scopes of subsequent analyses, and the first subset (referred as a\ robust set of the peaks, thresholded for expression analysis is shown as TSS peaks. They are\ named "p#@GENE_SYMBOL" if associated with 5'-end of known genes, or "p@CHROM:START..END,STRAND"\ otherwise. The summary tracks consist of the TSS (CAGE) peaks and summary profiles of TSS\ activities (total and maximum values). The summary track consists of the following tracks.\
\ 5′-end of the mapped CAGE reads are counted at a single base pair resolution (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the sample. The read counts tracks indicate raw counts of CAGE reads, and the TPM tracks indicate normalized counts as TPM (tags per million).\
\ \\ FANTOM5 data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ The FANTOM5 reprocessed data can be found and downloaded on the FANTOM website.
\ \\ Thanks to the FANTOM5 consortium,\ the Large Scale Data Managing Unit and Preventive Medicine and\ Applied Genomics Unit, the Center for Integrative Medical Sciences (IMS), and\ RIKEN for providing this data\ and its analysis.
\ \\ FANTOM Consortium and the RIKEN PMI and CLST (DGT), Forrest AR, Kawaji H, Rehli M, Baillie JK, de\ Hoon MJ, Haberle V, Lassmann T, Kulakovskiy IV, Lizio M et al.\ \ A promoter-level mammalian expression atlas.\ Nature. 2014 Mar 27;507(7493):462-70.\ PMID: 24670764; PMC: PMC4529748\
\ \\ Kanamori-Katayama M, Itoh M, Kawaji H, Lassmann T, Katayama S, Kojima M, Bertin N, Kaiho A, Ninomiya\ N, Daub CO et al.\ \ Unamplified cap analysis of gene expression on a single-molecule sequencer.\ Genome Res. 2011 Jul;21(7):1150-9.\ PMID: 21596820; PMC: PMC3129257\
\ \\ Lizio M, Harshbarger J, Shimoji H, Severin J, Kasukawa T, Sahin S, Abugessaisa I, Fukuda S, Hori F,\ Ishikawa-Kato S et al.\ \ Gateways to the FANTOM5 promoter level mammalian expression atlas.\ Genome Biol. 2015 Jan 5;16(1):22.\ PMID: 25723102; PMC: PMC4310165\
\ regulation 1 bigDataUrl /gbdb/hg38/fantom5/hg38.cage_peak.bb\ boxedCfg on\ colorByStrand 255,0,0 0,0,255\ dataVersion FANTOM5 reprocessed7\ exonArrows on\ html fantom5.html\ itemRgb on\ longLabel FANTOM5: DPI peak, robust set\ priority 1.2\ searchIndex name\ searchTrix hg38.cage_peak.bb.ix\ shortLabel TSS peaks\ showSubtrackColorOnUi on\ subGroups group=peaks\ superTrack fantom5 dense\ track robustPeaks\ type bigBed 8 +\ visibility dense\ wgEncodeReg4MarkH3k4me3 H3K4me3 (Layered) bigWig H3K4me3 signal marking active and poised promoters, averaged by organ/tissue 0 1.3 0 0 0 127 127 127 0 0 0\ Chemical modifications (e.g., methylation and acetylation) to the histone proteins present in\ chromatin influence gene expression by changing how accessible the chromatin is to transcription.\ A specific modification of a specific histone protein is called a histone mark. This track\ displays genome-wide enrichment levels of the H3K4me3 histone mark, as determined by ENCODE\ ChIP-seq data across all phases of the project. H3K4me3 refers to the tri-methylation of\ lysine 4 on the H3 histone protein and is associated with active or poised promoters. The data\ are processed following the\ ENCODE Histone\ ChIP-seq pipeline. Additional histone marks and other chromatin-associated ChIP-seq datasets\ are available at the\ ENCODE portal.
\ \\ For each organ, this track provides up to two subtracks averaging H3K4me3 signal:
\\ Whether one or two subtracks appear for an organ depends on which kinds of biosamples have\ been assayed:
\| Organ/Tissue | \Tissue and Primary Cell Subtrack | \All Biosamples Subtrack | \
|---|---|---|
| adipose | – | ✓ |
| adrenal gland | – | ✓ |
| blood | ✓ | ✓ |
| blood vessel | – | ✓ |
| bone | – | ✓ |
| bone marrow | ✓ | ✓ |
| brain | ✓ | ✓ |
| breast | ✓ | ✓ |
| connective tissue | ✓ | ✓ |
| embryo | ✓ | ✓ |
| epithelium | – | ✓ |
| esophagus | – | ✓ |
| eye | ✓ | ✓ |
| heart | ✓ | ✓ |
| kidney | ✓ | ✓ |
| large intestine | ✓ | ✓ |
| liver | ✓ | ✓ |
| lung | ✓ | ✓ |
| mouth | – | ✓ |
| muscle | ✓ | ✓ |
| nerve | – | ✓ |
| ovary | – | ✓ |
| pancreas | ✓ | ✓ |
| parathyroid gland | – | ✓ |
| penis | ✓ | ✓ |
| placenta | – | ✓ |
| prostate | ✓ | ✓ |
| skin | ✓ | ✓ |
| small intestine | – | ✓ |
| spinal cord | – | ✓ |
| spleen | – | ✓ |
| stomach | – | ✓ |
| testis | ✓ | ✓ |
| thymus | – | ✓ |
| thyroid | – | ✓ |
| urinary bladder | – | ✓ |
| uterus | ✓ | ✓ |
| vagina | – | ✓ |
\ By default, this track uses a transparent overlay to visualize data from multiple organs or tissues within\ the same vertical space. For each organ or tissue, signals from all associated experiments were\ averaged to generate the displayed track. Each organ or tissue is assigned a distinct\ color following the\ ENCODE color\ mapping convention,\ selected to be light and saturated to maintain clarity when overlaid. Initially, each layered\ track displays an overlay of five representative organs: blood, brain, kidney, liver, and\ muscle. Clicking on the track opens a details page where you can view and select organs or\ tissues.
\ \ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/regulation/organAve/adiposeH3K4me3.bw stdout
\ Data were generated by the ENCODE Consortium. We thank the production labs for generating the\ data: Drs. Bing Ren (UCSD), Bradley Bernstein (Broad),\ John Stamatoyannopoulos (UW), Joseph Costello (UCSF), Michael Snyder (Stanford),\ and Peggy Farnham (USC). The data were further processed for visualization through a collaborative effort between the\ Weng lab and the\ Moore lab\ at UMass Chan Medical School (funded by NIH grant HG012343). Integration and visualization\ were developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 aggregate transparentOverlay\ allButtonPair on\ autoScale on\ container multiWig\ dragAndDrop subtracks\ html wgEncodeReg4MarkH3k4me3.html\ longLabel H3K4me3 signal marking active and poised promoters, averaged by organ/tissue\ maxHeightPixels 100:50:11\ noInherit on\ priority 1.3\ shortLabel H3K4me3 (Layered)\ showSubtrackColorOnUi on\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4MarkH3k4me3\ type bigWig\ viewLimits 0:100\ visibility hide\ wgEncodeRegMarkH3k4me3 Layered H3K4Me3 bigWig 0 10000 H3K4Me3 Mark (Often Found Near Promoters) on 7 cell lines from ENCODE 0 1.3 0 0 0 127 127 127 0 0 0\ Chemical modifications (e.g., methylation and acetylation) to the histone proteins\ present in chromatin influence gene expression by changing how\ accessible the chromatin is to transcription. A specific modification of\ a specific histone protein is called a histone mark.\ This track shows the levels of enrichment of the H3K4Me3 histone mark across the genome as\ determined by a ChIP-seq assay. The H3K4Me3 histone mark is the tri-methylation of lysine 4 of the\ H3 histone protein, and it is associated with promoters that are active or poised to be\ activated. Additional histone marks and other chromatin associated ChIP-seq data is available at \ the Broad Histone\ page.\
\ \\ By default, this track uses a transparent overlay method of displaying data from a number of cell\ lines in the same vertical space. Each of the cell lines in this track\ is associated with a particular color, and these colors are relatively light and saturated so\ as to work best with the transparent overlay. The color of these tracks\ match their versions from their lifted source on the hg19 assembly. The colors are consistent with the\ other hg19 lifted tracks located in the ENCODE Regulation\ supertrack, with the exception being the DNase tracks, as they were not lifted from hg19 and are\ colored to reflect similarity of cell types.\
\ \\ This track shows data from the Bernstein Lab at the Broad Institute, as part of\ the ENCODE Consortium.\
\ \\ Primary ENCODE data produced during the 2007-2012 production phase were subject to a restriction\ period. However, the data here are past those restrictions and are freely available.\ The full data release policy for ENCODE is available\ here.\
\ regulation 1 aggregate transparentOverlay\ allButtonPair on\ container multiWig\ dragAndDrop subtracks\ longLabel H3K4Me3 Mark (Often Found Near Promoters) on 7 cell lines from ENCODE\ maxHeightPixels 100:30:11\ noInherit on\ origAssembly hg19\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 1.3\ shortLabel Layered H3K4Me3\ showSubtrackColorOnUi on\ superTrack wgEncodeReg hide\ track wgEncodeRegMarkH3k4me3\ type bigWig 0 10000\ viewLimits 0:150\ visibility hide\ spliceAIindelsMasked SpliceAI indels (masked) bigBed 9 + SpliceAI Indels (masked) 1 1.3 0 0 0 127 127 127 0 0 0\ SpliceAI is an open-source deep\ learning algorithm that predicts splicing probability for nucleotides and \ as a result can score DNA variants for splicing impact.\ Such variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a \ lookup tool \ is provided by the Broad institute.\
\ \The spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added and the new sequence is\ scored again. The "wildtype" container track shows the scores for the genome\ sequence itself and the "variants" container track shows the impact of all\ possible variants close to known splice sites. The "wildtype" subtracks are\ useful when looking at new transcript models, to evaluate how likely exon\ boundaries are. The "variants" subtracks are used to evaluate the impact of\ variants onto splicing, typically in medical diagnostics.\
\ \\ SpliceAI only annotates variants close to splice sites of genes defined by the \ Gencode gene annotation track. Additionally, SpliceAI does not annotate variants if they are\ close to chromosome ends (5kb on either side), deletions of length greater than\ twice the input parameter -D, or inconsistent with the reference fasta file.\
\ \\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \\ Variants are colored according to Walker et al. 2023 splicing imact:\
\\ The scores range from 0 to 1 and can be interpreted as the \ probability of the variant being splice-altering. In the paper, a detailed characterization is \ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\
The data were downloaded from Illumina. \
The spliceAI scores are represented in the VCF INFO field as \
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain \
\ Since most of the values are 0 or almost 0, we selected only those variants \ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \ \\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are \ made available by Illumina only for academic or not-for-profit research only. \ By accessing the SpliceAI data, you acknowledge and agree that you may only \ use this data for your own personal academic or not-for-profit research only, \ and not for any other purposes. You may not use this data for any for-profit, \ clinical, or other commercial purpose without obtaining a commercial license \ from Illumina, Inc.\
\ \\ Thanks to Illumina for making the data available. Thanks to Michael Hiller, Francois Lecoquierre and\ Jean-Madeleine de Sainte Agathe for making available and suggesting the SpliceAI wildtype tracks.\
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/spliceAIindelsMasked.bb\ filter.AIscore 0.02\ filterLabel.spliceType Splice type\ filterLimits.AIscore 0.02:1\ filterValues.spliceType donor_gain|Donor gain,donor_loss|Donor loss,acceptor_gain|Acceptor gain,acceptor_loss|Acceptor loss\ html spliceAI\ itemRgb on\ longLabel SpliceAI Indels (masked)\ mouseOver Change: $name\ The FANTOM5 track shows mapped transcription start sites (TSS) and their usage in primary cells,\ cell lines, and tissues to produce a comprehensive overview of gene expression across the human\ body by using single molecule sequencing.\
\ \Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \Individual biological states are profiled by HeliScopeCAGE, which is a variation of the CAGE\ (Cap Analysis Gene Expression) protocol based on a single molecule sequencer. The standard protocol\ requiring 5 µg of total RNA as a starting material is referred to as hCAGE, and an\ optimized version for a lower quantity (~ 100 ng) is referred to as LQhCAGE (Kanamori-Katyama\ et al. 2011).\
Transcription start sites (TSSs) were mapped and their usage in human and mouse primary cells,\ cell lines, and tissues was to produce a comprehensive overview of mammalian gene expression across the\ human body. 5′-end of the mapped CAGE reads are counted at a single base pair resolution\ (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the\ sample. Individual samples shown in "TSS activity" tracks are grouped as below.\
TSS (CAGE) peaks across the panel of the biological states (samples) are identified by DPI\ (decomposition based peak identification, Forrest et al. 2014), where each of the peaks consists of\ neighboring and related TSSs. The peaks are used as anchors to define promoters and units of\ promoter-level expression analysis. Two subsets of the peaks are defined based on evidence of read\ counts, depending on scopes of subsequent analyses, and the first subset (referred as a\ robust set of the peaks, thresholded for expression analysis is shown as TSS peaks. They are\ named "p#@GENE_SYMBOL" if associated with 5'-end of known genes, or "p@CHROM:START..END,STRAND"\ otherwise. The summary tracks consist of the TSS (CAGE) peaks and summary profiles of TSS\ activities (total and maximum values). The summary track consists of the following tracks.\
\ 5′-end of the mapped CAGE reads are counted at a single base pair resolution (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the sample. The read counts tracks indicate raw counts of CAGE reads, and the TPM tracks indicate normalized counts as TPM (tags per million).\
\ \\ FANTOM5 data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ The FANTOM5 reprocessed data can be found and downloaded on the FANTOM website.
\ \\ Thanks to the FANTOM5 consortium,\ the Large Scale Data Managing Unit and Preventive Medicine and\ Applied Genomics Unit, the Center for Integrative Medical Sciences (IMS), and\ RIKEN for providing this data\ and its analysis.
\ \\ FANTOM Consortium and the RIKEN PMI and CLST (DGT), Forrest AR, Kawaji H, Rehli M, Baillie JK, de\ Hoon MJ, Haberle V, Lassmann T, Kulakovskiy IV, Lizio M et al.\ \ A promoter-level mammalian expression atlas.\ Nature. 2014 Mar 27;507(7493):462-70.\ PMID: 24670764; PMC: PMC4529748\
\ \\ Kanamori-Katayama M, Itoh M, Kawaji H, Lassmann T, Katayama S, Kojima M, Bertin N, Kaiho A, Ninomiya\ N, Daub CO et al.\ \ Unamplified cap analysis of gene expression on a single-molecule sequencer.\ Genome Res. 2011 Jul;21(7):1150-9.\ PMID: 21596820; PMC: PMC3129257\
\ \\ Lizio M, Harshbarger J, Shimoji H, Severin J, Kasukawa T, Sahin S, Abugessaisa I, Fukuda S, Hori F,\ Ishikawa-Kato S et al.\ \ Gateways to the FANTOM5 promoter level mammalian expression atlas.\ Genome Biol. 2015 Jan 5;16(1):22.\ PMID: 25723102; PMC: PMC4310165\
\ regulation 0 aggregate transparentOverlay\ autoScale off\ configurable on\ container multiWig\ dataVersion FANTOM5 reprocessed7\ dragAndDrop subTracks\ html fantom5.html\ longLabel FANTOM5: Total counts of CAGE reads\ maxHeightPixels 64:64:11\ priority 1.3\ shortLabel Total counts of CAGE reads\ showSubtrackColorOnUi on\ subGroups group=counts\ superTrack fantom5 full\ track Total_counts_multiwig\ type bigWig 0 100\ viewLimits 0:100\ visibility full\ wgEncodeReg4MarkH3k27ac H3K27ac (Layered) bigWig H3K27ac signal marking active enhancers and promoters, averaged by organ/tissue 2 1.4 0 0 0 127 127 127 0 0 0\ Chemical modifications (e.g., methylation and acetylation) to the histone proteins present in\ chromatin influence gene expression by changing how accessible the chromatin is to transcription.\ A specific modification of a specific histone protein is called a histone mark. This track\ displays genome-wide enrichment levels of the H3K27ac histone mark, as determined by ENCODE\ ChIP-seq data across all phases of the project. H3K27ac refers to the acetylation of lysine 27\ on the H3 histone protein and is associated with active enhancers and promoters. The data are\ processed following the\ ENCODE Histone\ ChIP-seq pipeline. Additional histone marks and other chromatin-associated ChIP-seq datasets\ are available at the\ ENCODE portal.
\ \\ For each organ, this track provides up to two subtracks averaging H3K27ac signal:
\\ Whether one or two subtracks appear for an organ depends on which kinds of biosamples have\ been assayed:
\| Organ/Tissue | \Tissue and Primary Cell Subtrack | \All Biosamples Subtrack | \
|---|---|---|
| adipose | – | ✓ |
| adrenal gland | – | ✓ |
| blood | ✓ | ✓ |
| blood vessel | – | ✓ |
| bone | – | ✓ |
| bone marrow | ✓ | ✓ |
| brain | ✓ | ✓ |
| breast | ✓ | ✓ |
| connective tissue | ✓ | ✓ |
| embryo | ✓ | ✓ |
| epithelium | – | ✓ |
| esophagus | – | ✓ |
| eye | – | ✓ |
| heart | ✓ | ✓ |
| kidney | ✓ | ✓ |
| large intestine | ✓ | ✓ |
| liver | ✓ | ✓ |
| lung | ✓ | ✓ |
| mouth | – | ✓ |
| muscle | ✓ | ✓ |
| nerve | – | ✓ |
| ovary | – | ✓ |
| pancreas | ✓ | ✓ |
| parathyroid gland | – | ✓ |
| penis | ✓ | ✓ |
| placenta | – | ✓ |
| prostate | ✓ | ✓ |
| skin | ✓ | ✓ |
| small intestine | – | ✓ |
| spinal cord | – | ✓ |
| spleen | – | ✓ |
| stomach | – | ✓ |
| testis | – | ✓ |
| thymus | – | ✓ |
| thyroid | – | ✓ |
| urinary bladder | – | ✓ |
| uterus | ✓ | ✓ |
| vagina | – | ✓ |
\ By default, this track uses a transparent overlay to visualize data from multiple organs or tissues within\ the same vertical space. For each organ or tissue, signals from all associated experiments were\ averaged to generate the displayed track. Each organ or tissue is assigned a distinct\ color following the\ ENCODE color\ mapping convention,\ selected to be light and saturated to maintain clarity when overlaid. Initially, each layered\ track displays an overlay of five representative organs: blood, brain, kidney, liver, and\ muscle. Clicking on the track opens a details page where you can view and select organs or\ tissues.
\ \ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/regulation/organAve/adiposeH3K27ac.bw stdout
\ Data were generated by the ENCODE Consortium. We thank the production labs for generating the\ data: Drs. Bing Ren (UCSD), Bradley Bernstein (Broad),\ John Stamatoyannopoulos (UW), Joseph Costello (UCSF), Michael Snyder (Stanford),\ and Peggy Farnham (USC). The data were further processed for visualization through a collaborative effort between the\ Weng lab and the\ Moore lab\ at UMass Chan Medical School (funded by NIH grant HG012343). Integration and visualization\ were developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 aggregate transparentOverlay\ allButtonPair on\ autoScale on\ container multiWig\ dragAndDrop subtracks\ html wgEncodeReg4MarkH3k27ac.html\ longLabel H3K27ac signal marking active enhancers and promoters, averaged by organ/tissue\ maxHeightPixels 100:50:11\ noInherit on\ priority 1.4\ shortLabel H3K27ac (Layered)\ showSubtrackColorOnUi on\ superTrack wgEncodeReg4 full\ track wgEncodeReg4MarkH3k27ac\ type bigWig\ viewLimits 0:100\ visibility full\ wgEncodeRegMarkH3k27ac Layered H3K27Ac bigWig 0 10000 H3K27Ac Mark (Often Found Near Regulatory Elements) on 7 cell lines from ENCODE 2 1.4 0 0 0 127 127 127 0 0 0\ Chemical modifications (e.g., methylation and acetylation) to the histone proteins\ present in chromatin influence gene expression by changing how\ accessible the chromatin is to transcription. A specific modification of\ a specific histone protein is called a histone mark.\ This track shows the levels of enrichment of the H3K27Ac histone mark across the genome as\ determined by a ChIP-seq assay. The H3K27Ac histone mark is the acetylation of lysine 27 of the H3\ histone protein, and it is thought to enhance transcription possibly by blocking the\ spread of the repressive histone mark H3K27Me3. Additional histone marks and other chromatin \ associated ChIP-seq data is available at the \ Broad Histone page.\
\ \\ By default, this track uses a transparent overlay method of displaying data from a number of cell\ lines in the same vertical space. Each of the cell lines in this track\ is associated with a particular color, and these colors are relatively light and saturated so\ as to work best with the transparent overlay. The color of these tracks\ match their versions from their lifted source on the hg19 assembly. The colors are consistent with the \ other hg19 lifted tracks located in the ENCODE Regulation\ supertrack, with the exception being the DNase tracks, as they were not lifted from hg19 and are\ colored to reflect similarity of cell types. \
\ \\ This track shows data from the Bernstein Lab at the Broad Institute, as part of\ the ENCODE Consortium.\
\ \\ Primary ENCODE data produced during the 2007-2012 production phase were subject to a restriction\ period. However, the data here are past those restrictions and are freely available.\ The full data release policy for ENCODE is available\ here.\
\ regulation 1 aggregate transparentOverlay\ allButtonPair on\ container multiWig\ dragAndDrop subtracks\ longLabel H3K27Ac Mark (Often Found Near Regulatory Elements) on 7 cell lines from ENCODE\ maxHeightPixels 100:30:11\ noInherit on\ origAssembly hg19\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 1.4\ shortLabel Layered H3K27Ac\ showSubtrackColorOnUi on\ superTrack wgEncodeReg full\ track wgEncodeRegMarkH3k27ac\ type bigWig 0 10000\ viewLimits 0:100\ visibility full\ Max_counts_multiwig Max counts of CAGE reads bigWig 0 100 FANTOM5: Max counts of CAGE reads 2 1.4 0 0 0 127 127 127 0 0 0\ The FANTOM5 track shows mapped transcription start sites (TSS) and their usage in primary cells,\ cell lines, and tissues to produce a comprehensive overview of gene expression across the human\ body by using single molecule sequencing.\
\ \Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \Individual biological states are profiled by HeliScopeCAGE, which is a variation of the CAGE\ (Cap Analysis Gene Expression) protocol based on a single molecule sequencer. The standard protocol\ requiring 5 µg of total RNA as a starting material is referred to as hCAGE, and an\ optimized version for a lower quantity (~ 100 ng) is referred to as LQhCAGE (Kanamori-Katyama\ et al. 2011).\
Transcription start sites (TSSs) were mapped and their usage in human and mouse primary cells,\ cell lines, and tissues was to produce a comprehensive overview of mammalian gene expression across the\ human body. 5′-end of the mapped CAGE reads are counted at a single base pair resolution\ (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the\ sample. Individual samples shown in "TSS activity" tracks are grouped as below.\
TSS (CAGE) peaks across the panel of the biological states (samples) are identified by DPI\ (decomposition based peak identification, Forrest et al. 2014), where each of the peaks consists of\ neighboring and related TSSs. The peaks are used as anchors to define promoters and units of\ promoter-level expression analysis. Two subsets of the peaks are defined based on evidence of read\ counts, depending on scopes of subsequent analyses, and the first subset (referred as a\ robust set of the peaks, thresholded for expression analysis is shown as TSS peaks. They are\ named "p#@GENE_SYMBOL" if associated with 5'-end of known genes, or "p@CHROM:START..END,STRAND"\ otherwise. The summary tracks consist of the TSS (CAGE) peaks and summary profiles of TSS\ activities (total and maximum values). The summary track consists of the following tracks.\
\ 5′-end of the mapped CAGE reads are counted at a single base pair resolution (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the sample. The read counts tracks indicate raw counts of CAGE reads, and the TPM tracks indicate normalized counts as TPM (tags per million).\
\ \\ FANTOM5 data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ The FANTOM5 reprocessed data can be found and downloaded on the FANTOM website.
\ \\ Thanks to the FANTOM5 consortium,\ the Large Scale Data Managing Unit and Preventive Medicine and\ Applied Genomics Unit, the Center for Integrative Medical Sciences (IMS), and\ RIKEN for providing this data\ and its analysis.
\ \\ FANTOM Consortium and the RIKEN PMI and CLST (DGT), Forrest AR, Kawaji H, Rehli M, Baillie JK, de\ Hoon MJ, Haberle V, Lassmann T, Kulakovskiy IV, Lizio M et al.\ \ A promoter-level mammalian expression atlas.\ Nature. 2014 Mar 27;507(7493):462-70.\ PMID: 24670764; PMC: PMC4529748\
\ \\ Kanamori-Katayama M, Itoh M, Kawaji H, Lassmann T, Katayama S, Kojima M, Bertin N, Kaiho A, Ninomiya\ N, Daub CO et al.\ \ Unamplified cap analysis of gene expression on a single-molecule sequencer.\ Genome Res. 2011 Jul;21(7):1150-9.\ PMID: 21596820; PMC: PMC3129257\
\ \\ Lizio M, Harshbarger J, Shimoji H, Severin J, Kasukawa T, Sahin S, Abugessaisa I, Fukuda S, Hori F,\ Ishikawa-Kato S et al.\ \ Gateways to the FANTOM5 promoter level mammalian expression atlas.\ Genome Biol. 2015 Jan 5;16(1):22.\ PMID: 25723102; PMC: PMC4310165\
\ regulation 0 aggregate transparentOverlay\ autoScale off\ configurable on\ container multiWig\ dataVersion FANTOM5 reprocessed7\ dragAndDrop subTracks\ html fantom5.html\ longLabel FANTOM5: Max counts of CAGE reads\ maxHeightPixels 64:64:11\ priority 1.4\ shortLabel Max counts of CAGE reads\ showSubtrackColorOnUi on\ subGroups group=counts\ superTrack fantom5 full\ track Max_counts_multiwig\ type bigWig 0 100\ viewLimits 0:100\ visibility full\ nmdEscMane NMD Escape MANE bigBed 9 + NMD escape predictions: MANE Select Plus Clinical transcripts 3 1.4 0 0 0 127 127 127 0 0 0\ The NMD escape ruleset tracks show predicted regions where a premature termination\ codon (PTC) or frameshift variant is likely to cause the transcript to\ escape nonsense-mediated decay (NMD), leading to the production of an\ aberrant truncated protein rather than degradation of the mRNA.\
\ \\ The following rules were applied to transcript annotations to define predicted\ NMD escape regions (Nagy et al, Trends Biochem Sci 1998 and Lindeboom et al, Nat Genet 2016):\
\ \\ Non-coding transcripts (where CDS start equals CDS end) are excluded.\ Overlapping regions from multiple transcripts with identical coordinates and\ the same rule are collapsed into a single item, with the contributing\ transcript IDs stored as a comma-separated list.\
\ \\ Three versions of this track are available, based on different transcript annotation sets:\
\\ NMD escape regions were predicted based on the Exon Junction Complex\ (EJC)-dependent model of NMD. During normal translation, EJCs are deposited at\ exon-exon junctions after splicing. As the ribosome translates the mRNA, it\ displaces each EJC it encounters. When a PTC causes the ribosome to stall\ prematurely, any remaining downstream EJCs recruit surveillance factors\ (notably UPF1) that trigger mRNA degradation via NMD.\
\ \\ However, PTCs located in the last coding exon or within approximately 50 bp\ upstream of the last exon-exon junction are too close to the final EJC (or\ have no downstream EJC at all) for NMD to be triggered—the transcript\ escapes degradation. Conversely, PTCs located more than 50–55 bp\ upstream of the last exon-exon junction are predicted to elicit NMD.\
\ \\ Additional escape mechanisms, supported by Lindeboom et al. 2016 and other\ studies, are captured by three further rules:\
\\ Regions from overlapping transcripts with the same coordinates are collapsed into\ a single item. The gene symbol is shown as the item name. Mouseover displays the\ NMD escape rule and the number of transcripts. The details page lists all\ contributing transcript IDs.\
\ \\ Items are colored by the NMD escape rule that applies:\
\\ The data underlying this track can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Guido Neidhardt for suggesting this track at HUGO VEPTC 2025 and Andreas Lahner\ for feedback. Thanks to the Decipher Genome Browser team for introducing the idea of a\ track.\
\ \\ Kurosaki T, Popp MW, Maquat LE.\ \ Quality and quantity control of gene expression by nonsense-mediated mRNA decay.\ Nat Rev Mol Cell Biol. 2019 Jul;20(7):406-420.\ PMID: 30992545; PMC: PMC6855384\
\ \\ Lindeboom RGH, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Nagy E, Maquat LE.\ \ A rule for termination-codon position within intron-containing genes: when nonsense affects RNA\ abundance.\ Trends Biochem Sci. 1998 Jun;23(6):198-9.\ PMID: 9644970\
\ \ \ genes 1 bigDataUrl /gbdb/hg38/nmd/nmdEscMane.bb\ dataVersion MANE 1.5\ defaultLabelFields ncbiIds\ filterLabel.ncbiIds RefSeq accession (e.g. "*NM_000546*")\ filterLabel.transcripts Gencode accession (e.g. "*ENST00000269305*")\ filterText.ncbiIds *\ filterText.transcripts *\ filterType.ncbiIds wildcard\ filterType.transcripts wildcard\ html nmdEscTranscripts\ labelFields ncbiIds,name,transcripts\ labelSeparator " / "\ longLabel NMD escape predictions: MANE Select Plus Clinical transcripts\ mouseOverField mouseover\ parent nmd on\ priority 1.4\ shortLabel NMD Escape MANE\ track nmdEscMane\ type bigBed 9 +\ visibility pack\ cCREs_view cCREs bigBed ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals 4 1.5 0 0 0 127 127 127 0 0 0 regulation 1 filterLabel.cCRE_class cCRE class\ filterType.cCRE_class multipleListOr\ filterValues.cCRE_class CA-only|Chromatin accessibility only (CA-only),CA-CTCF|Chromatin accessibility + CTCF (CA-CTCF),CA-H3K4me3|Chromatin accessibility + H3K4me3 (CA-H3K4me3),CA-TF|Chromatin accessibility + transcription factor (CA-TF),Distal enhancer|Distal enhancer,Proximal enhancer|Proximal enhancer,Promoter|Promoter,Low-DNase|Low-DNase\ longLabel ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals\ parent coreCcres\ shortLabel cCREs\ track cCREs_view\ type bigBed\ view cCREs_view\ visibility squish\ CTCF_view CTCF bigWig ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals 2 1.5 0 0 0 127 127 127 0 0 0 regulation 0 longLabel ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals\ parent coreCcres\ shortLabel CTCF\ track CTCF_view\ type bigWig\ view CTCF_view\ visibility full\ wgEncodeReg4MarkCtcf CTCF (Layered) bigWig CTCF binding signal from ChIP-seq, averaged by organ/tissue 0 1.5 0 0 0 127 127 127 0 0 0\ CTCF (CCCTC-binding factor) is a multifunctional DNA-binding protein involved in chromatin\ organization, transcriptional regulation, and insulation of regulatory elements. This track\ displays genome-wide CTCF binding signal, as determined by CTCF ChIP-seq data across all\ phases of the ENCODE project. CTCF plays a key role in establishing chromatin loops and\ boundary elements that influence gene expression and higher-order genome architecture. CTCF\ binding sites often occur at insulators and chromatin loop anchors, which help define\ topologically associating domains (TADs) and mediate enhancer-promoter interactions. The data\ are processed following the\ ENCODE\ transcription factor ChIP-seq pipeline. Additional transcription factor binding and\ chromatin accessibility datasets are available at the\ ENCODE portal.
\ \\ For each organ, this track provides up to two subtracks averaging CTCF signal:
\\ Whether one or two subtracks appear for an organ depends on which kinds of biosamples have\ been assayed:
\| Organ/Tissue | \Tissue and Primary Cell Subtrack | \All Biosamples Subtrack | \
|---|---|---|
| adipose | – | ✓ |
| adrenal gland | – | ✓ |
| blood | ✓ | ✓ |
| blood vessel | – | ✓ |
| bone | – | ✓ |
| bone marrow | – | ✓ |
| brain | ✓ | ✓ |
| breast | ✓ | ✓ |
| connective tissue | ✓ | ✓ |
| embryo | – | ✓ |
| epithelium | – | ✓ |
| esophagus | – | ✓ |
| eye | ✓ | ✓ |
| heart | ✓ | ✓ |
| kidney | ✓ | ✓ |
| large intestine | ✓ | ✓ |
| liver | ✓ | ✓ |
| lung | ✓ | ✓ |
| mouth | – | ✓ |
| muscle | ✓ | ✓ |
| nerve | – | ✓ |
| ovary | – | ✓ |
| pancreas | ✓ | ✓ |
| parathyroid gland | – | ✓ |
| penis | – | ✓ |
| placenta | – | ✓ |
| prostate | ✓ | ✓ |
| skin | ✓ | ✓ |
| small intestine | – | ✓ |
| spinal cord | – | ✓ |
| spleen | – | ✓ |
| stomach | – | ✓ |
| testis | – | ✓ |
| thyroid | – | ✓ |
| uterus | ✓ | ✓ |
| vagina | – | ✓ |
\ By default, this track uses a transparent overlay to visualize data from multiple organs or tissues within\ the same vertical space. For each organ or tissue, signals from all associated experiments were\ averaged to generate the displayed track. Each organ or tissue is assigned a distinct\ color following the\ ENCODE color\ mapping convention,\ selected to be light and saturated to maintain clarity when overlaid. Initially, each layered\ track displays an overlay of five representative organs: blood, brain, kidney, liver, and\ muscle. Clicking on the track opens a details page where you can view and select organs or\ tissues.
\ \ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/regulation/organAve/adiposeCTCF.bw stdout
\ Data were generated by the ENCODE Consortium. We thank the production labs for generating the\ data: Drs. Bradley Bernstein (Broad), John Stamatoyannopoulos (UW),\ Michael Snyder (Stanford), Richard Myers (HAIB), and Vishwanath Iyer (UTA). The data were\ further processed for visualization through a collaborative effort between the\ Weng lab and the\ Moore lab\ at UMass Chan Medical School (funded by NIH grant HG012343). Integration and visualization\ were developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 aggregate transparentOverlay\ allButtonPair on\ autoScale on\ container multiWig\ dragAndDrop subtracks\ html wgEncodeReg4MarkCtcf.html\ longLabel CTCF binding signal from ChIP-seq, averaged by organ/tissue\ maxHeightPixels 100:50:11\ noInherit on\ priority 1.5\ shortLabel CTCF (Layered)\ showSubtrackColorOnUi on\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4MarkCtcf\ type bigWig\ viewLimits 0:100\ visibility hide\ DNase_view DNase bigWig ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals 2 1.5 0 0 0 127 127 127 0 0 0 regulation 0 longLabel ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals\ parent coreCcres\ shortLabel DNase\ track DNase_view\ type bigWig\ view DNase_view\ visibility full\ coreCcres ENCODE4 Core Collection bigBed 12 ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals 0 1.5 0 0 0 127 127 127 0 0 0\ This track displays biosample-specific candidate cis-regulatory elements (cCREs) \ alongside genome-wide epigenomic signals for the ENCODE Core Collection of 170 \ ENCODE biosamples that have been fully profiled using four core assays: \ DNase-seq, ChIP-seq for the histone modifications H3K4me3 and H3K27ac, \ and ChIP-seq for CTCF binding.
\\ Each subtrack corresponds to an individual experiment in a specific biosample. \ These data form the basis for generating biosample-specific annotations of \ cCREs, which compose the fifth subtrack for the biosample.
\\ Additional epigenomic datasets are available at the ENCODE portal, and further exploration \ of cCREs and their supporting data is available through the SCREEN web tool, accessible via the \ track details page.
\ \\ Each biosample contains five subtracks (DNase, CTCF, H3K27ac, and H3K4me3 signals \ and biosample-specific cCREs). Click a specific biosample type and organ/tissue \ combination to view available datasets. Epigenomic subtracks can be further \ filtered by the signal type. Below is a graphic summarizing biosampling available:
\\

\ Each signal track is colored based on the type of signal.\ The cCREs subtrack displays each active cCRE in the corresponding biosample \ as a colored box by type \
\

\ For items in the cCREs track, mousing over shows the element ID, along with a linkout to \ the corresponding element on SCREEN, and the cCRE class.
\ \\ The DNase-seq data were processed using the ENCODE DNase-seq pipeline, the H3K4me3 and H3K27ac ChIP-seq \ data were processed using the ENCODE histone ChIP-seq pipeline, and the CTCF ChIP-seq data were \ processed using the ENCODE transcription factor ChIP-seq pipeline.
\\ In addition to the cell type-agnostic classification (described in the cCRE Registry track\ in this collection), we evaluated the biochemical activity of each cCRE in individual\ biosamples using the corresponding biosample-specific DNase, H3K4me3, H3K27ac, and CTCF data.\ This allowed us to annotate active cCREs in individual biosamples, included as the cCRE\ subtrack for each biosample. cCREs with low DNase Z-scores in individual biosamples were\ deemed to be inactive and labeled with "Low Chromatin Accessibility."\
\ \\ All data is available from the ENCODE data portal.
\ \\ The data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\ files that can be downloaded from\ our download server.
\\
The cCREs tracks in this data are found as bigBed files, and the biosignal tracks as bigWig files.\
See the Data format link besides the specific data track for a URL to the file on our download\
server. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/ccre/coreCollection/ENCFF811RQX.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/ccre/coreCollection/ENCFF013UBZ_ENCFF901QWB_ENCFF972ZHA_ENCFF500RDL.bb stdout
\ Data were generated by the ENCODE Consortium. We thank the production labs for generating\ the DNase-seq and ChIP-seq data: Bing Ren (UCSD), Bradley Bernstein (Broad), Gregory\ Crawford (Duke), John Stamatoyannopoulos (UW), Michael Snyder (Stanford), Peggy Farnham\ (USC), and Richard Myers (HAIB).
\\ The DNase-seq and ChIP-seq data were further processed for visualization through a\ collaborative effort between the Weng lab and the Moore lab at UMass Chan Medical School\ (funded by NIH grant HG012343). Integration and visualization were developed by Drs. Mingshi\ Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School, who were part of the\ ENCODE Data Analysis Center. We thank the ENCODE production labs for generating the data.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J, Kawli T,\ Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N, Fu Y et\ al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 1 compositeTrack on\ dataVersion ENCODE release 4, 2024. (ENCODE4 data includes ENCODE2, ENCODE3, and the Roadmap Epigenomics Project)\ dimensions dimX=biosampleType dimY=organ\ html coreCollection.html\ longLabel ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals\ parent cCREs\ priority 1.5\ shortLabel ENCODE4 Core Collection\ sortOrder dataType=+ organ=+ biosampleType=+ donor=+\ subGroup1 organ Organ_Tissue adrenal_gland=Adrenal_gland blood=Blood blood_vessel=Blood_vessel bone=Bone bone_marrow=Bone_marrow brain=Brain breast=Breast connective_tissue=Connective_tissue embryo=Embryo epithelium=Epithelium esophagus=Esophagus eye=Eye heart=Heart large_intestine=Large_intestine liver=Liver lung=Lung muscle=Muscle nerve=Nerve pancreas=Pancreas penis=Penis prostate=Prostate skin=Skin small_intestine=Small_intestine spleen=Spleen stomach=Stomach testis=Testis thyroid=Thyroid uterus=Uterus vagina=Vagina\ subGroup2 biosampleType Biosample_type tissue=Tissue primary_cell=Primary_cell in_vitro_differentiated_cells=In_vitro_differentiated_cells organoid=Organoid cell_line=Cell_line\ subGroup3 view Views cCREs_view=cCREs DNase_view=DNase H3K4me3_view=H3K4me3 H3K27ac_view=H3K27ac CTCF_view=CTCF\ subGroup4 simpleBiosample Biosample A673=A673 AG04450=AG04450 CD14-positive_monocyte-_female=CD14-positive_monocyte-_female Caco-2=Caco-2 DND-41=DND-41 GM12878=GM12878 GM23338=GM23338 H1=H1 H9=H9 HCT116=HCT116 HFFc6=HFFc6 HL-60=HL-60 HeLa-S3=HeLa-S3 HepG2=HepG2 IMR-90=IMR-90 K562=K562 MCF-7=MCF-7 MM_1S=MM_1S NCI-H929=NCI-H929 OCI-LY7=OCI-LY7 PC-3=PC-3 PC-9=PC-9 Panc1=Panc1 Peyers_patch-_female_adult__51_years_=Peyers_patch-_female_adult__51_years_ Peyers_patch-_female_adult__53_years_=Peyers_patch-_female_adult__53_years_ Peyers_patch-_male_adult__37_years_=Peyers_patch-_male_adult__37_years_ Peyers_patch-_male_adult__54_years_=Peyers_patch-_male_adult__54_years_ SK-N-SH=SK-N-SH WERI-Rb-1=WERI-Rb-1 adrenal_gland-_female_adult__41_years_=adrenal_gland-_female_adult__41_years_ adrenal_gland-_female_adult__51_years_=adrenal_gland-_female_adult__51_years_ adrenal_gland-_female_adult__53_years_=adrenal_gland-_female_adult__53_years_ adrenal_gland-_male_adult__37_years_=adrenal_gland-_male_adult__37_years_ adrenal_gland-_male_adult__54_years_=adrenal_gland-_male_adult__54_years_ ascending_aorta-_female_adult__51_years_=ascending_aorta-_female_adult__51_years_ ascending_aorta-_female_adult__53_years_=ascending_aorta-_female_adult__53_years_ astrocyte=astrocyte astrocyte-_male_adult__53_years_=astrocyte-_male_adult__53_years_ bipolar_neuron__treated_-_male_adult__53_years__treated_with_0_5_ug_mL_doxycycline_hyclate_for_4_days=bipolar_neuron__treated_-_male_adult__53_years__treated_with_0_5_ug_mL_doxycycline_hyclate_for_4_days body_of_pancreas-_female_adult__51_years_=body_of_pancreas-_female_adult__51_years_ body_of_pancreas-_male_adult__37_years_=body_of_pancreas-_male_adult__37_years_ body_of_pancreas-_male_adult__54_years_=body_of_pancreas-_male_adult__54_years_ brain_microvascular_endothelial_cell=brain_microvascular_endothelial_cell breast_epithelium-_female_adult__51_years_=breast_epithelium-_female_adult__51_years_ cardiac_muscle_cell-_embryo=cardiac_muscle_cell-_embryo chondrocyte-_female_embryo__5_days_=chondrocyte-_female_embryo__5_days_ colonic_mucosa-_female_adult__41_years_=colonic_mucosa-_female_adult__41_years_ coronary_artery-_female_adult__53_years_=coronary_artery-_female_adult__53_years_ endodermal_cell-_female_embryo__5_days_=endodermal_cell-_female_embryo__5_days_ endothelial_cell-_male_adult__53_years_=endothelial_cell-_male_adult__53_years_ esophagus_muscularis_mucosa-_male_adult__37_years_=esophagus_muscularis_mucosa-_male_adult__37_years_ esophagus_squamous_epithelium-_male_adult__37_years_=esophagus_squamous_epithelium-_male_adult__37_years_ gastrocnemius_medialis-_female_adult__51_years_=gastrocnemius_medialis-_female_adult__51_years_ gastrocnemius_medialis-_female_adult__53_years_=gastrocnemius_medialis-_female_adult__53_years_ gastrocnemius_medialis-_male_adult__37_years_=gastrocnemius_medialis-_male_adult__37_years_ gastrocnemius_medialis-_male_adult__54_years_=gastrocnemius_medialis-_male_adult__54_years_ glutamatergic_neuron-_male_adult__53_years__male_adult__53_years__nuclear_fraction=glutamatergic_neuron-_male_adult__53_years__male_adult__53_years__nuclear_fraction heart_left_ventricle-_female_adult__46_years_=heart_left_ventricle-_female_adult__46_years_ heart_left_ventricle-_female_adult__53_years_=heart_left_ventricle-_female_adult__53_years_ heart_left_ventricle-_female_adult__56_years_=heart_left_ventricle-_female_adult__56_years_ heart_left_ventricle-_female_adult__59_years_=heart_left_ventricle-_female_adult__59_years_ heart_left_ventricle-_male_adult__43_years_=heart_left_ventricle-_male_adult__43_years_ heart_right_ventricle-_female_adult__46_years_=heart_right_ventricle-_female_adult__46_years_ heart_right_ventricle-_female_adult__56_years_=heart_right_ventricle-_female_adult__56_years_ heart_right_ventricle-_male_adult__40_years_=heart_right_ventricle-_male_adult__40_years_ heart_right_ventricle-_male_adult__43_years_=heart_right_ventricle-_male_adult__43_years_ heart_right_ventricle-_male_adult__61_years_=heart_right_ventricle-_male_adult__61_years_ heart_right_ventricle-_male_adult__66_years_=heart_right_ventricle-_male_adult__66_years_ heart_right_ventricle-_male_adult__69_years_=heart_right_ventricle-_male_adult__69_years_ hepatocyte-_female_embryo__5_days_=hepatocyte-_female_embryo__5_days_ keratinocyte-_female=keratinocyte-_female left_lung-_female_child__16_years_=left_lung-_female_child__16_years_ left_lung-_male_adult__40_years_=left_lung-_male_adult__40_years_ left_ventricle_myocardium_inferior-_male_adult__60_years_=left_ventricle_myocardium_inferior-_male_adult__60_years_ lower_lobe_of_left_lung-_female_adult__59_years_=lower_lobe_of_left_lung-_female_adult__59_years_ lower_lobe_of_left_lung-_male_adult__60_years_=lower_lobe_of_left_lung-_male_adult__60_years_ mesothelial_cell_of_epicardium-_female_embryo__5_days_=mesothelial_cell_of_epicardium-_female_embryo__5_days_ middle_frontal_area_46__Alzheimers_disease_-_female_adult__74_years__with_Alzheimers_disease=middle_frontal_area_46__Alzheimers_disease_-_female_adult__74_years__with_Alzheimers_disease middle_frontal_area_46__Alzheimers_disease_-_female_adult__81_years__with_Alzheimers_disease=middle_frontal_area_46__Alzheimers_disease_-_female_adult__81_years__with_Alzheimers_disease middle_frontal_area_46__Alzheimers_disease_-_female_adult__85_years__with_Alzheimers_disease=middle_frontal_area_46__Alzheimers_disease_-_female_adult__85_years__with_Alzheimers_disease middle_frontal_area_46__Alzheimers_disease_-_female_adult__86_years__with_Alzheimers_disease=middle_frontal_area_46__Alzheimers_disease_-_female_adult__86_years__with_Alzheimers_disease middle_frontal_area_46__Alzheimers_disease_-_female_adult__88_years__with_Alzheimers_disease=middle_frontal_area_46__Alzheimers_disease_-_female_adult__88_years__with_Alzheimers_disease middle_frontal_area_46__Alzheimers_disease_-_female_adult__89_years__with_Alzheimers_disease=middle_frontal_area_46__Alzheimers_disease_-_female_adult__89_years__with_Alzheimers_disease middle_frontal_area_46__Alzheimers_disease_-_female_adult__90_or_above_years__with_Alzheimers_disease=middle_frontal_area_46__Alzheimers_disease_-_female_adult__90_or_above_years__with_Alzheimers_disease middle_frontal_area_46__cognitive_impairment_-_female_adult__81_years__with_Cognitive_impairment=middle_frontal_area_46__cognitive_impairment_-_female_adult__81_years__with_Cognitive_impairment middle_frontal_area_46__cognitive_impairment_-_female_adult__86_years__with_Cognitive_impairment=middle_frontal_area_46__cognitive_impairment_-_female_adult__86_years__with_Cognitive_impairment middle_frontal_area_46__cognitive_impairment_-_female_adult__90_or_above_years__with_Cognitive_impairment=middle_frontal_area_46__cognitive_impairment_-_female_adult__90_or_above_years__with_Cognitive_impairment middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__83_years__with_mild_cognitive_impairment=middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__83_years__with_mild_cognitive_impairment middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__87_years__with_mild_cognitive_impairment=middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__87_years__with_mild_cognitive_impairment middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__88_years__with_mild_cognitive_impairment=middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__88_years__with_mild_cognitive_impairment middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__90_or_above_years__with_mild_cognitive_impairment=middle_frontal_area_46__mild_cognitive_impairment_-_female_adult__90_or_above_years__with_mild_cognitive_impairment middle_frontal_area_46__mild_cognitive_impairment_-_male_adult__89_years__with_mild_cognitive_impairment=middle_frontal_area_46__mild_cognitive_impairment_-_male_adult__89_years__with_mild_cognitive_impairment middle_frontal_area_46__mild_cognitive_impairment_-_male_adult__90_or_above_years__with_mild_cognitive_impairment=middle_frontal_area_46__mild_cognitive_impairment_-_male_adult__90_or_above_years__with_mild_cognitive_impairment middle_frontal_area_46-_female_adult__78_years_=middle_frontal_area_46-_female_adult__78_years_ middle_frontal_area_46-_female_adult__79_years_=middle_frontal_area_46-_female_adult__79_years_ middle_frontal_area_46-_female_adult__82_years_=middle_frontal_area_46-_female_adult__82_years_ middle_frontal_area_46-_female_adult__83_years_=middle_frontal_area_46-_female_adult__83_years_ middle_frontal_area_46-_female_adult__84_years_=middle_frontal_area_46-_female_adult__84_years_ middle_frontal_area_46-_female_adult__87_years_=middle_frontal_area_46-_female_adult__87_years_ middle_frontal_area_46-_female_adult__88_years_=middle_frontal_area_46-_female_adult__88_years_ middle_frontal_area_46-_female_adult__89_years_=middle_frontal_area_46-_female_adult__89_years_ middle_frontal_area_46-_female_adult__90_or_above_years_=middle_frontal_area_46-_female_adult__90_or_above_years_ middle_frontal_area_46-_male_adult__71_years_=middle_frontal_area_46-_male_adult__71_years_ middle_frontal_area_46-_male_adult__78_years_=middle_frontal_area_46-_male_adult__78_years_ middle_frontal_area_46-_male_adult__82_years_=middle_frontal_area_46-_male_adult__82_years_ middle_frontal_area_46-_male_adult__83_years_=middle_frontal_area_46-_male_adult__83_years_ middle_frontal_area_46-_male_adult__84_years_=middle_frontal_area_46-_male_adult__84_years_ middle_frontal_area_46-_male_adult__86_years_=middle_frontal_area_46-_male_adult__86_years_ middle_frontal_area_46-_male_adult__87_years_=middle_frontal_area_46-_male_adult__87_years_ neural_crest_cell-_female_embryo__5_days_=neural_crest_cell-_female_embryo__5_days_ neural_progenitor_cell-_female_embryo__5_days_=neural_progenitor_cell-_female_embryo__5_days_ osteocyte-_female_embryo__5_days_=osteocyte-_female_embryo__5_days_ pancreas-_female_adult__41_years_=pancreas-_female_adult__41_years_ pancreas-_female_adult__59_years_=pancreas-_female_adult__59_years_ pancreas-_female_adult__61_years_=pancreas-_female_adult__61_years_ pancreas-_female_child__16_years_=pancreas-_female_child__16_years_ progenitor_cell_of_endocrine_pancreas-_female_embryo__5_days_=progenitor_cell_of_endocrine_pancreas-_female_embryo__5_days_ prostate_gland-_male_adult__37_years_=prostate_gland-_male_adult__37_years_ right_atrium_auricular_region-_female_adult__51_years_=right_atrium_auricular_region-_female_adult__51_years_ right_atrium_auricular_region-_female_adult__53_years_=right_atrium_auricular_region-_female_adult__53_years_ right_lobe_of_liver-_female_adult__53_years_=right_lobe_of_liver-_female_adult__53_years_ sigmoid_colon-_female_adult__53_years_=sigmoid_colon-_female_adult__53_years_ sigmoid_colon-_male_adult__54_years_=sigmoid_colon-_male_adult__54_years_ spleen-_female_adult__41_years_=spleen-_female_adult__41_years_ spleen-_female_adult__53_years_=spleen-_female_adult__53_years_ spleen-_female_adult__59_years_=spleen-_female_adult__59_years_ spleen-_female_adult__61_years_=spleen-_female_adult__61_years_ stomach-_female_adult__51_years_=stomach-_female_adult__51_years_ stomach-_female_adult__53_years_=stomach-_female_adult__53_years_ stomach-_male_adult__37_years_=stomach-_male_adult__37_years_ stomach-_male_adult__54_years_=stomach-_male_adult__54_years_ testis-_male_adult__37_years_=testis-_male_adult__37_years_ testis-_male_adult__54_years_=testis-_male_adult__54_years_ thoracic_aorta-_male_adult__37_years_=thoracic_aorta-_male_adult__37_years_ thyroid_gland-_female_adult__51_years_=thyroid_gland-_female_adult__51_years_ thyroid_gland-_female_adult__53_years_=thyroid_gland-_female_adult__53_years_ thyroid_gland-_male_adult__37_years_=thyroid_gland-_male_adult__37_years_ thyroid_gland-_male_adult__54_years_=thyroid_gland-_male_adult__54_years_ tibial_artery-_male_adult__37_years_=tibial_artery-_male_adult__37_years_ tibial_nerve-_female_adult__51_years_=tibial_nerve-_female_adult__51_years_ tibial_nerve-_male_adult__37_years_=tibial_nerve-_male_adult__37_years_ tibial_nerve-_male_adult__54_years_=tibial_nerve-_male_adult__54_years_ transverse_colon-_female_adult__51_years_=transverse_colon-_female_adult__51_years_ transverse_colon-_female_adult__53_years_=transverse_colon-_female_adult__53_years_ transverse_colon-_male_adult__37_years_=transverse_colon-_male_adult__37_years_ transverse_colon-_male_adult__54_years_=transverse_colon-_male_adult__54_years_ type_B_pancreatic_cell-_female_embryo__5_days_=type_B_pancreatic_cell-_female_embryo__5_days_ upper_lobe_of_left_lung-_female_adult__51_years_=upper_lobe_of_left_lung-_female_adult__51_years_ upper_lobe_of_left_lung-_female_adult__53_years_=upper_lobe_of_left_lung-_female_adult__53_years_ upper_lobe_of_left_lung-_male_adult__37_years_=upper_lobe_of_left_lung-_male_adult__37_years_ upper_lobe_of_left_lung-_male_adult__54_years_=upper_lobe_of_left_lung-_male_adult__54_years_ uterus-_female_adult__53_years_=uterus-_female_adult__53_years_ vagina-_female_adult__51_years_=vagina-_female_adult__51_years_ vagina-_female_adult__53_years_=vagina-_female_adult__53_years_\ subGroup5 donor Donor_ID ENCDO000AAB=ENCDO000AAB ENCDO000AAC=ENCDO000AAC ENCDO000AAD=ENCDO000AAD ENCDO000AAE=ENCDO000AAE ENCDO000AAK=ENCDO000AAK ENCDO000AAM=ENCDO000AAM ENCDO000AAW=ENCDO000AAW ENCDO000AAX=ENCDO000AAX ENCDO000ABB=ENCDO000ABB ENCDO000ABD=ENCDO000ABD ENCDO000ABE=ENCDO000ABE ENCDO000ACR=ENCDO000ACR ENCDO000ADT=ENCDO000ADT ENCDO001AAA=ENCDO001AAA ENCDO006DAA=ENCDO006DAA ENCDO027VXA=ENCDO027VXA ENCDO033BMB=ENCDO033BMB ENCDO070VNS=ENCDO070VNS ENCDO077CCP=ENCDO077CCP ENCDO080EZF=ENCDO080EZF ENCDO097MEH=ENCDO097MEH ENCDO101GPB=ENCDO101GPB ENCDO151OJB=ENCDO151OJB ENCDO153NUY=ENCDO153NUY ENCDO183AAA=ENCDO183AAA ENCDO186XRB=ENCDO186XRB ENCDO201EUI=ENCDO201EUI ENCDO203ASI=ENCDO203ASI ENCDO218FFZ=ENCDO218FFZ ENCDO220OYR=ENCDO220OYR ENCDO222AAA=ENCDO222AAA ENCDO227AAA=ENCDO227AAA ENCDO236YSH=ENCDO236YSH ENCDO250PFZ=ENCDO250PFZ ENCDO265AAA=ENCDO265AAA ENCDO268AAA=ENCDO268AAA ENCDO271OUW=ENCDO271OUW ENCDO290OPS=ENCDO290OPS ENCDO336AAA=ENCDO336AAA ENCDO349AAA=ENCDO349AAA ENCDO351AAA=ENCDO351AAA ENCDO354SJE=ENCDO354SJE ENCDO359XWR=ENCDO359XWR ENCDO392CRK=ENCDO392CRK ENCDO407UTA=ENCDO407UTA ENCDO411EVD=ENCDO411EVD ENCDO423GGP=ENCDO423GGP ENCDO448YMQ=ENCDO448YMQ ENCDO448ZXP=ENCDO448ZXP ENCDO451RUA=ENCDO451RUA ENCDO461DJY=ENCDO461DJY ENCDO471EKG=ENCDO471EKG ENCDO477WED=ENCDO477WED ENCDO520EJG=ENCDO520EJG ENCDO570AKP=ENCDO570AKP ENCDO575EGL=ENCDO575EGL ENCDO575WHY=ENCDO575WHY ENCDO592ZWW=ENCDO592ZWW ENCDO609ZOG=ENCDO609ZOG ENCDO623FPG=ENCDO623FPG ENCDO634UMA=ENCDO634UMA ENCDO637GUS=ENCDO637GUS ENCDO640RUC=ENCDO640RUC ENCDO647UHQ=ENCDO647UHQ ENCDO660TGP=ENCDO660TGP ENCDO666UNK=ENCDO666UNK ENCDO669IVL=ENCDO669IVL ENCDO672KST=ENCDO672KST ENCDO697GBW=ENCDO697GBW ENCDO697SWU=ENCDO697SWU ENCDO707TUE=ENCDO707TUE ENCDO736YJH=ENCDO736YJH ENCDO737WWC=ENCDO737WWC ENCDO739EFE=ENCDO739EFE ENCDO793LXB=ENCDO793LXB ENCDO808ASZ=ENCDO808ASZ ENCDO830KFO=ENCDO830KFO ENCDO832DBZ=ENCDO832DBZ ENCDO845GYA=ENCDO845GYA ENCDO845WKR=ENCDO845WKR ENCDO847KYQ=ENCDO847KYQ ENCDO853VGZ=ENCDO853VGZ ENCDO856ZOJ=ENCDO856ZOJ ENCDO877NVF=ENCDO877NVF ENCDO907CMO=ENCDO907CMO ENCDO907YUG=ENCDO907YUG ENCDO915WZE=ENCDO915WZE ENCDO916IIE=ENCDO916IIE ENCDO924HBJ=ENCDO924HBJ ENCDO926KEV=ENCDO926KEV ENCDO967KID=ENCDO967KID ENCDO997SGX=ENCDO997SGX ENCDO999WDR=ENCDO999WDR\ subGroup6 dataType Data_type typeCcres=cCREs typeDNase=DNase typeH3k4me3=H3K4me3 typeH3k27ac=H3K27ac typeCtcf=CTCF\ track coreCcres\ type bigBed 12\ visibility hide\ H3K27ac_view H3K27ac bigWig ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals 2 1.5 0 0 0 127 127 127 0 0 0 regulation 0 longLabel ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals\ parent coreCcres\ shortLabel H3K27ac\ track H3K27ac_view\ type bigWig\ view H3K27ac_view\ visibility full\ H3K4me3_view H3K4me3 bigWig ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals 2 1.5 0 0 0 127 127 127 0 0 0 regulation 0 longLabel ENCODE4 Core Collection of 170 biosamples with sample-specific cCRE annotations & epigenomic signals\ parent coreCcres\ shortLabel H3K4me3\ track H3K4me3_view\ type bigWig\ view H3K4me3_view\ visibility full\ jaspar2024 JASPAR 2024 TFBS bigBed 6 + JASPAR CORE 2024 - Predicted Transcription Factor Binding Sites 3 1.5 0 0 0 127 127 127 1 0 0 http://jaspar.genereg.net/search?q=$$&collection=all&tax_group=all&tax_id=all&type=all&class=all&family=all&version=all regulation 1 bigDataUrl /gbdb/hg38/jaspar/JASPAR2024.bb\ filter.score 400\ filterByRange.score 0:1000\ filterValues.TFName Ahr::Arnt,Alx1,ALX3,Alx4,Ar,ARGFX,Arid3a,Arid3b,Arid5a,Arnt,ARNT2,ARNT::HIF1A,Arntl,Arx,ASCL1,ASCL1,Ascl2,Atf1,ATF2,Atf3,ATF3,ATF4,ATF6,ATF7,Atoh1,Atoh1,ATOH7,BACH1,Bach1::Mafk,BACH2,BACH2,BARHL1,BARHL2,BARX1,BARX2,BATF,BATF3,BATF::JUN,BCL11A,Bcl11B,BCL6,BCL6B,Bhlha15,BHLHA15,BHLHE22,BHLHE22,BHLHE23,BHLHE40,BHLHE41,BNC2,BSX,CDX1,CDX2,CDX4,CEBPA,CEBPB,CEBPD,CEBPE,CEBPG,CEBPG,CLOCK,CREB1,CREB3,CREB3L1,Creb3l2,CREB3L4,CREB3L4,Creb5,CREM,Crx,CTCF,CTCF,CTCF,CTCFL,CUX1,CUX2,DBP,Ddit3::Cebpa,DLX1,Dlx2,Dlx3,Dlx4,Dlx5,DLX6,Dmbx1,Dmrt1,DMRT3,DMRTA1,DMRTA2,DMRTC2,DPRX,DRGX,Dux,DUX4,DUXA,E2F1,E2F2,E2F3,E2F4,E2F6,E2F7,E2F8,EBF1,Ebf2,EBF3,Ebf4,EGR1,EGR2,EGR3,EGR4,EHF,ELF1,ELF2,ELF3,ELF4,Elf5,ELK1,ELK1::HOXA1,ELK1::HOXB13,ELK1::SREBF2,ELK3,ELK4,EMX1,EMX2,EN1,EN2,EOMES,EPAS1,ERF,ERF::FIGLA,ERF::FOXI1,ERF::FOXO1,ERF::HOXB13,ERF::NHLH1,ERF::SREBF2,Erg,ESR1,ESR2,ESRRA,ESRRB,Esrrg,ESX1,ETS1,ETS2,ETV1,ETV2,ETV2::DRGX,ETV2::FIGLA,ETV2::FOXI1,ETV2::HOXB13,ETV3,ETV4,ETV5,ETV5::DRGX,ETV5::FIGLA,ETV5::FOXI1,ETV5::FOXO1,ETV5::HOXA2,ETV6,ETV7,EVX1,EVX2,EWSR1-FLI1,FERD3L,FEV,FEZF2,FIGLA,FLI1,FLI1::DRGX,FLI1::FOXI1,FOS,FOS,FOSB::JUN,FOSB::JUNB,FOSB::JUNB,FOS::JUN,FOS::JUN,FOS::JUNB,FOS::JUND,FOSL1,FOSL1::JUN,FOSL1::JUN,FOSL1::JUNB,FOSL1::JUND,FOSL1::JUND,FOSL2,FOSL2::JUN,FOSL2::JUN,FOSL2::JUNB,FOSL2::JUNB,FOSL2::JUND,FOSL2::JUND,FOXA1,FOXA2,FOXA3,FOXB1,FOXC1,FOXC2,FOXD1,FOXD2,FOXD3,FOXE1,Foxf1,FOXF2,FOXG1,FOXH1,FOXI1,Foxj2,FOXJ2::ELF1,Foxj3,FOXK1,FOXK2,FOXL1,Foxl2,Foxn1,FOXN3,Foxo1,FOXO1::ELF1,FOXO1::ELK1,FOXO1::ELK3,FOXO1::FLI1,Foxo3,FOXO4,FOXO6,FOXP1,FOXP2,FOXP3,FOXP4,Foxq1,FOXS1,GABPA,GATA1,GATA1::TAL1,GATA2,Gata3,GATA4,GATA5,GATA6,GBX1,GBX2,GCM1,GCM2,GFI1,Gfi1B,Gli1,Gli2,GLI3,GLIS1,GLIS2,GLIS3,Gmeb1,GMEB2,GRHL1,GRHL2,GSC,GSC2,GSX1,GSX2,Hand1,Hand1::Tcf3,HAND2,HES1,HES2,HES5,HES6,HES7,HESX1,HEY1,HEY2,Hic1,HIC2,HIF1A,HINFP,HLF,HMBOX1,Hmga1,Hmx1,Hmx2,Hmx3,Hnf1A,HNF1A,HNF1B,HNF4A,HNF4A,HNF4G,HOXA1,HOXA10,Hoxa11,Hoxa13,HOXA2,HOXA3,HOXA4,HOXA5,HOXA6,HOXA7,HOXA9,HOXB1,HOXB13,HOXB2,HOXB2::ELK1,HOXB3,HOXB4,HOXB5,HOXB6,HOXB7,HOXB8,HOXB9,HOXC10,HOXC11,HOXC12,HOXC13,HOXC4,HOXC8,HOXC9,HOXD10,HOXD11,HOXD12,HOXD12::ELK1,Hoxd13,HOXD3,HOXD4,HOXD8,HOXD9,HSF1,HSF2,HSF4,IKZF1,IKZF2,Ikzf3,INSM1,Irf1,IRF2,IRF3,IRF4,IRF5,IRF6,IRF7,IRF8,IRF9,Isl1,ISL2,ISX,JDP2,JDP2,Jun,JUN,JUNB,JUNB,JUND,JUND,JUN::JUNB,JUN::JUNB,KLF1,KLF10,KLF11,KLF12,KLF13,KLF14,KLF15,KLF16,KLF17,KLF2,KLF3,KLF4,KLF5,KLF6,KLF7,KLF9,LBX1,LBX2,Lef1,Lhx1,LHX2,Lhx3,Lhx4,LHX5,LHX6,Lhx8,LHX9,LIN54,LMX1A,LMX1B,MAF,MAFA,Mafb,MAFF,Mafg,MAFG::NFE2L1,MAFK,MAF::NFE2,MAX,MAX::MYC,MAZ,Mecom,MEF2A,MEF2B,MEF2C,MEF2D,MEIS1,MEIS1,MEIS2,MEIS2,MEIS3,MEOX1,MEOX2,MGA,MGA::EVX1,MITF,mix-a,MIXL1,MLX,Mlxip,MLXIPL,MNT,MNX1,MSANTD3,MSC,Msgn1,MSX1,MSX2,Msx3,MTF1,MXI1,MYB,MYBL1,MYBL2,MYC,MYCN,MYF5,MYF6,MYOD1,MYOG,MZF1,Nanog,NEUROD1,Neurod2,Neurod2,NEUROG1,NEUROG2,NEUROG2,Nfat5,Nfatc1,Nfatc2,NFATC3,NFATC4,NFE2,Nfe2l2,NFIA,NFIB,NFIC,NFIC,NFIC::TLX1,NFIL3,NFIX,NFIX,NFKB1,NFKB2,NFYA,NFYB,NFYC,NHLH1,NHLH2,Nkx2-1,NKX2-2,NKX2-3,NKX2-4,NKX2-5,NKX2-8,Nkx3-1,Nkx3-2,NKX6-1,NKX6-2,NKX6-3,Nobox,NOTO,Npas2,Npas4,NR1D1,NR1D2,Nr1H2,NR1H2::RXRA,Nr1h3,Nr1h3::Rxra,Nr1H4,NR1H4::RXRA,NR1I2,NR1I3,NR2C1,NR2C2,NR2C2,Nr2e1,Nr2e3,NR2F1,NR2F1,NR2F1,NR2F2,Nr2f6,Nr2F6,NR2F6,NR3C1,NR3C2,NR4A1,NR4A2,NR4A2::RXRA,NR5A1,Nr5A2,NR6A1,Nrf1,NRL,OLIG1,Olig2,OLIG2,OLIG3,ONECUT1,ONECUT2,ONECUT3,OSR1,OSR2,OTX1,OTX2,OVOL1,OVOL2,PATZ1,PAX1,PAX2,PAX3,PAX3,PAX4,PAX5,PAX6,Pax7,PAX8,PAX9,PBX1,PBX2,PBX3,PDX1,Pgr,PGR,PHOX2A,PHOX2B,PITX1,PITX2,PITX3,PKNOX1,PKNOX2,PLAG1,Plagl1,PLAGL2,POU1F1,POU2F1,POU2F1::SOX2,POU2F2,POU2F3,POU3F1,POU3F2,POU3F3,POU3F4,POU4F1,POU4F2,POU4F3,POU5F1,POU5F1B,Pou5f1::Sox2,POU6F1,POU6F1,POU6F2,Ppara,PPARA::RXRA,PPARD,PPARG,Pparg::Rxra,PRDM1,Prdm14,Prdm15,Prdm4,Prdm5,PRDM9,PROP1,PROX1,PRRX1,PRRX2,Ptf1A,Ptf1A,Ptf1A,RARA,RARA,RARA::RXRA,RARA::RXRG,Rarb,Rarb,RARB,Rarg,Rarg,RARG,RAX,RAX2,RBPJ,REL,RELA,RELB,REST,RFX1,RFX2,RFX3,RFX4,RFX5,Rfx6,RFX7,Rhox11,RHOXF1,RORA,RORA,RORB,RORC,RREB1,Runx1,RUNX2,RUNX3,Rxra,RXRA::VDR,RXRB,RXRB,RXRG,RXRG,SATB1,SCRT1,SCRT2,SHOX,Shox2,SIX1,SIX2,Six3,Six4,SMAD2,SMAD3,Smad4,SMAD5,SNAI1,SNAI2,SNAI3,SOHLH2,Sox1,SOX10,Sox11,SOX12,SOX13,SOX14,SOX15,Sox17,SOX18,SOX2,SOX21,Sox3,SOX4,Sox5,Sox6,Sox7,SOX8,SOX9,SP1,SP2,SP3,SP4,SP5,SP8,SP9,SPDEF,Spi1,SPIB,SPIC,Spz1,SREBF1,SREBF1,SREBF2,SREBF2,SRF,SRY,STAT1,STAT1::STAT2,Stat2,STAT3,Stat4,Stat5a,Stat5a::Stat5b,Stat5b,Stat6,TAL1::TCF3,TBP,TBR1,TBX1,TBX15,TBX18,TBX19,TBX2,TBX20,TBX21,TBX3,TBX4,TBX5,Tbx6,TBXT,Tcf12,TCF12,Tcf21,TCF21,TCF3,TCF4,TCF7,TCF7L1,TCF7L2,TCFL5,TEAD1,TEAD2,TEAD3,TEAD4,TEF,TFAP2A,TFAP2A,TFAP2A,TFAP2B,TFAP2B,TFAP2B,TFAP2C,TFAP2C,TFAP2C,TFAP2E,TFAP4,TFAP4,TFAP4::ETV1,TFAP4::FLI1,TFCP2,Tfcp2l1,TFDP1,TFE3,TFEB,TFEC,TGIF1,TGIF2,TGIF2LX,TGIF2LY,THAP1,Thap11,THRA,THRB,THRB,THRB,TLX2,TP53,TP63,TP73,TRPS1,TWIST1,Twist2,UNCX,USF1,USF2,VAX1,VAX2,Vdr,VENTX,VEZF1,VSX1,VSX2,Wt1,XBP1,Yy1,YY2,ZBED1,ZBED2,ZBED4,ZBTB11,ZBTB12,ZBTB14,ZBTB17,ZBTB18,Zbtb2,ZBTB24,ZBTB26,ZBTB32,ZBTB33,ZBTB6,ZBTB7A,ZBTB7B,ZBTB7C,ZEB1,ZFP14,Zfp335,ZFP42,ZFP57,Zfp809,Zfp961,Zfx,ZIC1,Zic1::Zic2,Zic2,Zic3,ZIC4,ZIC5,ZIM3,ZKSCAN1,ZKSCAN3,ZKSCAN5,ZNF135,ZNF136,ZNF140,ZNF143,ZNF148,ZNF157,ZNF16,ZNF175,ZNF184,ZNF189,ZNF211,ZNF213,ZNF214,ZNF24,ZNF257,ZNF263,ZNF274,ZNF281,ZNF282,ZNF317,ZNF320,ZNF324,ZNF331,ZNF341,ZNF343,ZNF35,ZNF354A,ZNF354C,ZNF382,ZNF384,ZNF410,ZNF416,ZNF417,ZNF418,Znf423,ZNF449,ZNF454,ZNF460,ZNF524,ZNF528,ZNF530,ZNF547,ZNF549,ZNF558,ZNF574,ZNF582,ZNF610,ZNF652,ZNF667,ZNF669,ZNF675,ZNF677,ZNF680,ZNF682,ZNF684,ZNF692,ZNF701,ZNF707,ZNF708,ZNF740,ZNF75A,ZNF75D,ZNF76,ZNF766,ZNF768,ZNF770,ZNF784,ZNF8,ZNF816,ZNF85,ZNF93,ZSCAN16,ZSCAN21,ZSCAN29,ZSCAN31,ZSCAN4\ labelFields TFName\ longLabel JASPAR CORE 2024 - Predicted Transcription Factor Binding Sites\ maxItems 100000\ motifPwmTable hgFixed.jasparCore2024\ parent jaspar off\ priority 1.5\ shortLabel JASPAR 2024 TFBS\ track jaspar2024\ type bigBed 6 +\ visibility pack\ nmdEscGencode NMD Escape Gencode bigBed 9 + NMD escape predictions: Gencode transcripts 1 1.5 0 0 0 127 127 127 0 0 0\ The NMD escape ruleset tracks show predicted regions where a premature termination\ codon (PTC) or frameshift variant is likely to cause the transcript to\ escape nonsense-mediated decay (NMD), leading to the production of an\ aberrant truncated protein rather than degradation of the mRNA.\
\ \\ The following rules were applied to transcript annotations to define predicted\ NMD escape regions (Nagy et al, Trends Biochem Sci 1998 and Lindeboom et al, Nat Genet 2016):\
\ \\ Non-coding transcripts (where CDS start equals CDS end) are excluded.\ Overlapping regions from multiple transcripts with identical coordinates and\ the same rule are collapsed into a single item, with the contributing\ transcript IDs stored as a comma-separated list.\
\ \\ Three versions of this track are available, based on different transcript annotation sets:\
\\ NMD escape regions were predicted based on the Exon Junction Complex\ (EJC)-dependent model of NMD. During normal translation, EJCs are deposited at\ exon-exon junctions after splicing. As the ribosome translates the mRNA, it\ displaces each EJC it encounters. When a PTC causes the ribosome to stall\ prematurely, any remaining downstream EJCs recruit surveillance factors\ (notably UPF1) that trigger mRNA degradation via NMD.\
\ \\ However, PTCs located in the last coding exon or within approximately 50 bp\ upstream of the last exon-exon junction are too close to the final EJC (or\ have no downstream EJC at all) for NMD to be triggered—the transcript\ escapes degradation. Conversely, PTCs located more than 50–55 bp\ upstream of the last exon-exon junction are predicted to elicit NMD.\
\ \\ Additional escape mechanisms, supported by Lindeboom et al. 2016 and other\ studies, are captured by three further rules:\
\\ Regions from overlapping transcripts with the same coordinates are collapsed into\ a single item. The gene symbol is shown as the item name. Mouseover displays the\ NMD escape rule and the number of transcripts. The details page lists all\ contributing transcript IDs.\
\ \\ Items are colored by the NMD escape rule that applies:\
\\ The data underlying this track can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Guido Neidhardt for suggesting this track at HUGO VEPTC 2025 and Andreas Lahner\ for feedback. Thanks to the Decipher Genome Browser team for introducing the idea of a\ track.\
\ \\ Kurosaki T, Popp MW, Maquat LE.\ \ Quality and quantity control of gene expression by nonsense-mediated mRNA decay.\ Nat Rev Mol Cell Biol. 2019 Jul;20(7):406-420.\ PMID: 30992545; PMC: PMC6855384\
\ \\ Lindeboom RGH, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Nagy E, Maquat LE.\ \ A rule for termination-codon position within intron-containing genes: when nonsense affects RNA\ abundance.\ Trends Biochem Sci. 1998 Jun;23(6):198-9.\ PMID: 9644970\
\ \ \ genes 1 bigDataUrl /gbdb/hg38/nmd/nmdEscRegions.bb\ dataVersion Gencode V49\ filterLabel.transcripts Filter on transcript ID (e.g. "*ENST00000269305*")\ filterText.transcripts *\ filterType.transcripts wildcard\ html nmdEscTranscripts\ longLabel NMD escape predictions: Gencode transcripts\ mouseOverField mouseover\ parent nmd on\ priority 1.5\ shortLabel NMD Escape Gencode\ track nmdEscGencode\ type bigBed 9 +\ visibility dense\ wgEncodeRegDnaseClustered DNase Clusters bed 5 . DNase I Hypersensitivity Peak Clusters from ENCODE (95 cell types) 0 1.6 0 0 0 127 127 127 1 0 0\ This track shows clusters of DNaseI hypersensitivity derived from assays in 95 cell types\ by the\ John Stamatoyannapoulos lab\ at the University of Washington from September 2007 to January 2011, as part of the\ ENCODE project first production phase.\ Regulatory regions in general, and promoters in particular, tend to be DNase-sensitive. \
\ \\ Additional views of this data sites are displayed from the\ DNaseI HS track.\ The peaks in that track are the basis for the clusters shown here, \ which combine data from peaks from the different cell lines.\ Please note that track colors for the DNase tracks are based on similiarity of cell types,\ while there is different coloring for cell types on the ENCODE hg38\ Transcription track,\ Layered H3K4Me1 track,\ Layered H3K4Me3 track, and\ Layered H3K27Ac track,\ which match the coloring used in their previous versions lifted from the hg19 assembly.\
\ \ \\ A gray box indicates the extent of the hypersensitive region. \ The darkness is proportional to the maximum signal strength observed in any cell line. \ The number to the left of the box shows how many cell lines are hypersensitive in the region. \ The track can be configured to restrict the display to elements above a specified score \ in the range 1-1000 (where score is based on signal strength).\
\ \\ Raw sequence data files were processed by the UCSC ENCODE DNase analysis pipeline (July 2014\ specification), diagrammed here:\
\ \
\
Credit: Qian Alvin Qin, X. Liu lab\
\ Briefly, sequence files were aligned to the hg38 (GRCh38) genome assembly augmented with 'sponge'\ sequence (ref). Multi-mapped reads were removed, as were reads that aligned to 'sponge' or\ mitochondiral sequence. Results from all replicates were pooled, and further processed by\ the Hotspot program to call peaks.\
\ \\ Peaks of DNaseI hypersensitivity from the ENCODE DNase Analysis Pipeline at UCSC\ were assigned normalized scores (by UCSC regClusterMakeTableOfTables) in the range 0-1000 based\ on the \ narrowPeak\ signalValue and then clustered on score (by UCSC regCluster) to generate singly-linked clusters. \ Additional documentation on the methods used to identify hypersensitive sites are \ available from the\ DNaseI HS track.\
\ \\ This track is based on sequence data from the University of Washington ENCODE group, \ with subsequent processing by UCSC.\ For additional credits and references, see the\ DNaseI HS track.\
\ regulation 1 controlledVocabulary cellType=wgEncodeCell treatment=wgEncodeTreatment\ group regulation\ html wgEncodeRegDnaseClustered\ inputTableFieldDisplay cellType treatment\ inputTrackTable wgEncodeRegDnaseClusteredInputs\ longLabel DNase I Hypersensitivity Peak Clusters from ENCODE (95 cell types)\ priority 1.6\ scoreFilter 200\ scoreFilterLimits 1:1000\ shortLabel DNase Clusters\ sourceTable wgEncodeRegDnaseClusteredSources\ spectrum on\ superTrack wgEncodeReg hide\ track wgEncodeRegDnaseClustered\ type bed 5 .\ nmdEscNcbiRefSeq NMD Escape RefSeq bigBed 9 + NMD escape predictions: NCBI RefSeq Curated transcripts 0 1.6 0 0 0 127 127 127 0 0 0\ The NMD escape ruleset tracks show predicted regions where a premature termination\ codon (PTC) or frameshift variant is likely to cause the transcript to\ escape nonsense-mediated decay (NMD), leading to the production of an\ aberrant truncated protein rather than degradation of the mRNA.\
\ \\ The following rules were applied to transcript annotations to define predicted\ NMD escape regions (Nagy et al, Trends Biochem Sci 1998 and Lindeboom et al, Nat Genet 2016):\
\ \\ Non-coding transcripts (where CDS start equals CDS end) are excluded.\ Overlapping regions from multiple transcripts with identical coordinates and\ the same rule are collapsed into a single item, with the contributing\ transcript IDs stored as a comma-separated list.\
\ \\ Three versions of this track are available, based on different transcript annotation sets:\
\\ NMD escape regions were predicted based on the Exon Junction Complex\ (EJC)-dependent model of NMD. During normal translation, EJCs are deposited at\ exon-exon junctions after splicing. As the ribosome translates the mRNA, it\ displaces each EJC it encounters. When a PTC causes the ribosome to stall\ prematurely, any remaining downstream EJCs recruit surveillance factors\ (notably UPF1) that trigger mRNA degradation via NMD.\
\ \\ However, PTCs located in the last coding exon or within approximately 50 bp\ upstream of the last exon-exon junction are too close to the final EJC (or\ have no downstream EJC at all) for NMD to be triggered—the transcript\ escapes degradation. Conversely, PTCs located more than 50–55 bp\ upstream of the last exon-exon junction are predicted to elicit NMD.\
\ \\ Additional escape mechanisms, supported by Lindeboom et al. 2016 and other\ studies, are captured by three further rules:\
\\ Regions from overlapping transcripts with the same coordinates are collapsed into\ a single item. The gene symbol is shown as the item name. Mouseover displays the\ NMD escape rule and the number of transcripts. The details page lists all\ contributing transcript IDs.\
\ \\ Items are colored by the NMD escape rule that applies:\
\\ The data underlying this track can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Guido Neidhardt for suggesting this track at HUGO VEPTC 2025 and Andreas Lahner\ for feedback. Thanks to the Decipher Genome Browser team for introducing the idea of a\ track.\
\ \\ Kurosaki T, Popp MW, Maquat LE.\ \ Quality and quantity control of gene expression by nonsense-mediated mRNA decay.\ Nat Rev Mol Cell Biol. 2019 Jul;20(7):406-420.\ PMID: 30992545; PMC: PMC6855384\
\ \\ Lindeboom RGH, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Nagy E, Maquat LE.\ \ A rule for termination-codon position within intron-containing genes: when nonsense affects RNA\ abundance.\ Trends Biochem Sci. 1998 Jun;23(6):198-9.\ PMID: 9644970\
\ \ \ genes 1 bigDataUrl /gbdb/hg38/nmd/nmdEscNcbiRefSeq.bb\ dataVersion GCF_000001405.40-RS_2025_08\ filterLabel.transcripts Filter on transcript ID (e.g. "NM_005228*")\ filterText.transcripts *\ filterType.transcripts wildcard\ html nmdEscTranscripts\ longLabel NMD escape predictions: NCBI RefSeq Curated transcripts\ mouseOverField mouseover\ parent nmd off\ priority 1.6\ shortLabel NMD Escape RefSeq\ track nmdEscNcbiRefSeq\ type bigBed 9 +\ visibility hide\ wgEncodeReg4Txn Transcription (Layered) bigWig Strand-specific transcription signal measured by total RNA-seq, averaged by organ/tissue 0 1.6 0 0 0 127 127 127 0 0 0\ For each organ, this track provides up to two subtracks averaging total RNA-seq signal per\ strand:
\\ Each subtrack provides separate signal tracks for the plus and minus genomic strands.\ Whether one or two pairs of strand subtracks appear for an organ depends on which kinds of\ biosamples have been assayed:
\| Organ/Tissue | \Tissue and Primary Cell Subtrack | \All Biosamples Subtrack | \
|---|---|---|
| adipose | ✓ | – |
| adrenal gland | ✓ | – |
| blood | ✓ | ✓ |
| blood vessel | ✓ | – |
| brain | ✓ | ✓ |
| breast | ✓ | ✓ |
| connective tissue | ✓ | ✓ |
| embryo | ✓ | ✓ |
| epithelium | ✓ | ✓ |
| esophagus | ✓ | – |
| eye | ✓ | – |
| gallbladder | ✓ | – |
| heart | ✓ | ✓ |
| kidney | ✓ | ✓ |
| large intestine | ✓ | ✓ |
| liver | ✓ | ✓ |
| lung | ✓ | ✓ |
| mouth | ✓ | – |
| muscle | ✓ | ✓ |
| nerve | ✓ | – |
| nose | ✓ | – |
| ovary | ✓ | – |
| pancreas | ✓ | ✓ |
| placenta | ✓ | – |
| prostate | ✓ | ✓ |
| skin | ✓ | ✓ |
| small intestine | ✓ | – |
| spinal cord | ✓ | – |
| spleen | ✓ | – |
| stomach | ✓ | – |
| testis | ✓ | – |
| thyroid | ✓ | – |
| trachea | ✓ | – |
| urinary bladder | ✓ | – |
| uterus | ✓ | – |
| vagina | ✓ | – |
\ By default, this track uses a transparent overlay to visualize data from multiple organs or tissues within\ the same vertical space. For each organ or tissue, signals from all associated experiments were\ averaged to generate the displayed track. Each organ or tissue is assigned a distinct\ color following the\ ENCODE color\ mapping convention,\ selected to be light and saturated to maintain clarity when overlaid. Initially, each layered\ track displays an overlay of five representative organs: blood, brain, kidney, liver, and\ muscle. Clicking on the track opens a details page where you can view and select organs or\ tissues. Subtracks can be further filtered by strandedness (+ strand or - strand) and life\ stage of the biosample.
\ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/regulation/organAve/adiposePlus.bw stdout
\ Data were generated by the ENCODE Consortium. We thank the production labs for generating the\ data: Drs. Barbara Wold (Caltech) and Thomas Gingeras (CSHL). The data were further processed\ for visualization through a collaborative effort between the\ Weng lab and the\ Moore lab\ at UMass Chan Medical School (funded by NIH grant HG012343). Integration and visualization\ were developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 aggregate transparentOverlay\ allButtonPair on\ autoScale on\ container multiWig\ dragAndDrop subtracks\ html wgEncodeReg4Txn.html\ longLabel Strand-specific transcription signal measured by total RNA-seq, averaged by organ/tissue\ maxHeightPixels 100:50:11\ noInherit on\ priority 1.6\ shortLabel Transcription (Layered)\ showSubtrackColorOnUi on\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4Txn\ type bigWig\ visibility hide\ nmdDetectiveAi NMDetective-AI bigWig NMDetective-AI: Deep-learning NMD efficiency prediction per position (MANE Select only) 0 1.7 128 0 128 191 127 191 0 0 0\ The NMDetective-AI tracks display deep-learning predictions of\ nonsense-mediated mRNA decay (NMD) efficiency for every possible stop-gain\ single-nucleotide variant in MANE Select transcripts. The model was trained on\ ~14,000 somatic premature termination codons (PTCs) measured by allele-specific\ expression in large human cohorts (TCGA) and was tested on ~1,800 held-out\ germline PTCs (TCGA germline and GTEx) (Veiner et al.).\
\ \\ Predictions are continuous: higher values indicate that a PTC at that codon is\ predicted to trigger NMD (the mRNA is degraded); lower values indicate that the\ PTC is predicted to evade NMD (the truncated mRNA may be translated into an\ aberrant protein). The output is normalized against canonical controls so that\ +0.5 corresponds to full NMD efficiency at a PTC and −0.5\ corresponds to no NMD efficiency (a last-exon PTC). The scale is not strictly\ bounded: due to measurement and prediction noise, observed values fall in\ roughly −1.1 to +1.5, with the bulk of items inside the nominal\ −0.5 to +0.5 interval.\
\ \| Track | Description |
|---|---|
| NMDetective-AI | \Signal track (bigWig) showing the position-averaged prediction across\ all stop-gain SNVs at each codon. Useful for browsing efficiency along a\ transcript at a glance. |
| NMDetective-AI variants | \Per-stop-gain track (bigBed) with one item per (transcript, codon,\ mutant codon) combination. Each item is colored by its prediction and\ carries the reference and mutant codon, amino-acid position, transcript\ accession, and a pre-rendered mouseover summary. |
\ The NMDetective-AI signal track is drawn with a default y-axis range of\ −1.1 to +1.5. Positions with positive values (predicted NMD-triggering)\ are shown above the baseline; positions with negative values (predicted NMD\ escape) are shown below.\
\ \\ The NMDetective-AI variants track colors each item along a continuous\ diverging Okabe-Ito palette running from blue (most NMD-evading) through grey\ (near zero) to vermillion (most NMD-triggering). The mouseover verdict groups\ items into three categories using the binarization thresholds derived in the\ Veiner et al. Methods (Gaussian mixture model fit to gnomAD\ predictions):\
\\ Mouseover for each variant shows the codon change, the prediction value with\ its NMD verdict, and the MANE Select transcript accession. Click an item to\ see the full set of fields on the details page.\
\ \\ NMDetective-AI is a fine-tuned version of the Orthrus mRNA foundation model\ (Mamba architecture, ~10M parameters), trained on full-length transcript\ sequences encoded as a six-track representation (four nucleotide channels,\ one CDS-start channel, one splice-site channel). The model integrates\ allele-specific PTC expression from large-scale genomic data with mRNA\ language-model embeddings and high-throughput deep mutational scanning, and\ predicts NMD efficiency for every possible stop-gain mutation in every codon\ of a MANE Select transcript.\
\ \\ The training set comprised 14,337 somatic PTCs from TCGA, with chromosomes 1\ and 20 held out as a validation set. The held-out test set comprised 1,065\ germline PTCs from TCGA and 763 germline PTCs from GTEx. The authors report\ that the model's accuracy on the somatic validation set approaches the\ empirical reproducibility ceiling of the underlying allele-specific\ expression measurements.\
\ \\ The publicly released predictions cover MANE Select transcripts at Gencode\ v46. Predictions for transcripts outside the MANE Select set are not yet\ available; broader coverage is planned by the authors after peer review.\
\ \\ Source files were obtained from the\ Vejni/NMDetectiveAI\ GitHub repository (supplementary files\ NMDetectiveAI_MANE.bw.gz and NMDetectiveAI_MANE.bed.gz) and\ processed at UCSC: the bigWig is used as supplied; the BED was recolored with\ the diverging Okabe-Ito palette described above, rescored into the\ 0–1000 BED range, and augmented with a pre-rendered mouseover column\ before conversion to bigBed.\
\ \\ Note: the manuscript is currently a bioRxiv preprint and has not yet\ completed peer review. Predictions may be refreshed when the final version\ of the data is released.\
\ \\ The data underlying these tracks can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Marcell Veiner and Fran Supek for sharing the NMDetective-AI\ predictions ahead of publication, and to the wider Veiner et al.\ author group for developing the model.\
\ \\ Veiner M, Toledano I, Palou-Márquez G, Lehner B, Supek F.\ \ Quantitative prediction of nonsense-mediated mRNA decay across human genes by\ genomic language model and large-scale mutational scanning.\ bioRxiv. 2026 Mar 26.\ doi: 10.64898/2026.03.24.714003.\ Supplementary prediction files at\ github.com/Vejni/NMDetectiveAI.\
\ genes 0 autoScale off\ bigDataUrl /gbdb/hg38/nmd/nmdDetectAi.bw\ color 128,0,128\ html nmdDetectiveAi\ longLabel NMDetective-AI: Deep-learning NMD efficiency prediction per position (MANE Select only)\ maxHeightPixels 128:120:8\ parent nmd off\ priority 1.7\ shortLabel NMDetective-AI\ track nmdDetectiveAi\ type bigWig\ viewLimits -1.1:1.5\ visibility hide\ wgEncodeReg4TfPeaks TF rPeaks bigBed 12 + Transcription factor representative peak (rPeak) clusters from ENCODE 4 0 1.7 0 0 0 127 127 127 1 0 0This track displays representative ChIP-seq peaks (rPeaks) and detected DNA motif\ sites for regulatory regions in the human genome, identified using ENCODE ChIP-seq data\ across all phases of the project. The regions are bound by DNA-associated proteins\ involved in transcriptional regulation, including RNA polymerase, transcription factors\ (TFs), and chromatin remodeling proteins. Sequence-specific TFs bind directly to short\ DNA motifs via their DNA-binding domains, while other proteins associate indirectly\ through interactions with sequence-specific TFs. Chromatin immunoprecipitation followed\ by sequencing (ChIP-seq) is a high-throughput method used to map genome-wide protein-DNA\ interactions. Regions with high ChIP-seq signal (peaks) frequently contain binding sites\ for the assayed protein. For each DNA-associated protein, ChIP-seq peaks from all ENCODE\ biosamples were integrated to define a set of representative peaks (rPeaks).\ For detailed information on individual factors and their motifs, see\ Factorbook.org.
\ \Each rPeak is colored in grayscale by maximum ChIP-seq signal across contributing\ biosamples (darker = higher signal, score 0 to 1,000):
\| Color | \Score | \
|---|---|
| \ | 1000 (highest signal) | \
| \ | 750 | \
| \ | 500 | \
| \ | 250 | \
| \ | 1 (lowest signal) | \
Low-scoring peaks appear in very light gray by default; the\ Shade of lowest-scoring items setting can darken them for easier visibility.\ Note: Increasing the shade reduces the visible contrast between low and\ high scoring peaks.
\ \If the rPeak overlaps a cognate TF motif site from the collection in Andrews\ et al., 2023, the motif site is colored green using decorators.
\ \Clicking on an rPeak provides detailed information about the biosamples where the\ rPeak was detected, including the count of biosamples with contributing ChIP-seq peaks,\ the total number of biosamples assayed for the protein, and a per-biosample table\ listing each contributing experiment with its ENCODE accession. The protein name links\ to Factorbook, and overlapping\ ENCODE candidate cis-regulatory elements (cCREs) link to\ SCREEN.
\ \By default, rPeaks for all 912 DNA-associated proteins with ENCODE ChIP-seq data are\ displayed. A filter is available to select specific proteins.
\ \2,509 ENCODE ChIP-seq experiments were integrated from 912 DNA-associated proteins across\ 1,152 unique biosamples to produce representative peaks (rPeaks) for each protein. The\ processing steps were as follows:
\ \\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigBed\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.
\ \\
These files may also be locally explored using our tool bigBedToBed,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain features confined to a given range, e.g.,\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 https://hgdownload.soe.ucsc.edu/gbdb/hg38/encode4/regulation/tfRpeak/TFrPeakClusters.bb stdout
This track was made possible by the efforts of the ENCODE Consortium, ENCODE\ ChIP-seq production laboratories, and the ENCODE Data Coordination Center for generating\ and processing the ChIP-seq datasets. The ENCODE accession numbers for the constituent\ datasets are accessible from the peak details page. The data were generated by the\ following production labs: Drs. Bradley Bernstein (Broad),\ John Stamatoyannopoulos (UW), Kevin Struhl (HMS), Kevin White (UChicago),\ Michael Snyder (Stanford), Peggy Farnham (USC), Richard Myers (HAIB),\ Sherman Weissman (Yale), Tim Reddy (Duke), Vishwanath Iyer (UTA),\ and Xiang-Dong Fu (UCSD).
\ \The data were further processed for visualization through a collaborative effort between\ the Weng lab and the\ Moore lab at UMass\ Chan Medical School (funded by NIH grant HG012343). Special thanks to Drs. Mingshi Gao,\ Greg Andrews, Jill Moore, and Zhiping Weng at UMass Chan Medical School, who were members\ of the ENCODE Data Analysis Center, for developing this track, including providing the rPeak\ and motif datasets and associated metadata and building the track.
\ \\ Andrews G, Fan K, Pratt HE, Phalke N, Zoonomia Consortium, Karlsson EK, Lindblad-Toh K,\ Weng Z.\ \ Mammalian evolution of human cis-regulatory elements and transcription factor binding\ sites.\ Science. 2023;380(6643):eabn7930.\ PMID: 37104580\
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\ \\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ \\ Pratt HE, Andrews GR, Phalke N, Huey JD, Purcaro MJ, van der Velde A, Moore JE, Weng Z.\ \ Factorbook: an updated catalog of transcription factor motifs and candidate regulatory\ motif sites.\ Nucleic Acids Research. 2022;50(D1):D141-D149.\ PMID: 34747468\
\ regulation 1 bigDataUrl /gbdb/hg38/encode4/regulation/tfRpeak/TFrPeakClusters.bb\ decorator.default.bigDataUrl /gbdb/hg38/encode4/regulation/tfRpeak/TFrPeakClustersDecorator.bb\ detailsDynamicTable json_table|Experiments supporting this rPeak\ filterType.factor multipleListOr\ filterValues.factor ADNP,AFF1,AFF4,AGO1,AGO2,AHDC1,AHR,AKAP8,AKNA,ARHGAP35,ARID1B,ARID2,ARID3A,ARID4A,ARID4B,ARID5B,ARNTL,ARNT,AR,ASH1L,ASH2L,ATF1,ATF2,ATF3,ATF4,ATF5,ATF6,ATF7,ATOH8,BACH1,BATF,BCL11A,BCL11B,BCL3,BCL6B,BCL6,BCLAF1,BCOR,BDP1,BHLHE40,BMI1,BNC2,BRCA1,BRCA2,BRD4,BRD9,C11orf30,CAMTA2,CBFA2T2,CBFA2T3,CBFB,CBX1,CBX2,CBX3,CBX5,CBX8,CC2D1A,CCAR2,CCNT2,CDC5L,CEBPA,CEBPB,CEBPD,CEBPG,CEBPZ,CERS6,CGGBP1,CHAMP1,CHD1,CHD2,CHD4,CHD7,CLOCK,CREB1,CREB3L1,CREB3,CREB5,CREM,CSDC2,CSDE1,CSRNP3,CTBP1,CTBP2,CTCFL,CTCF,CUX1,DACH1,DBP,DDIT3,DDX20,DEAF1,DEK,DIDO1,DLX6,DMAP1,DMTF1,DNMT1,DNMT3B,DPF2,DR1,DRAP1,DZIP1,E2F1,E2F2,E2F3,E2F4,E2F5,E2F6,E2F7,E2F8,E4F1,EBF1,EEA1,EED,EGR1,EGR2,EHF,EHMT2,ELF1,ELF2,ELF3,ELF4,ELK1,ELK4,EMX1,EP300,EP400,ERF,ERG,ESR1,ESRRA,ESRRG,ETS1,ETV1,ETV4,ETV5,ETV6,EZH2phosphoT487,EZH2,FEZF1,FIP1L1,FOSB,FOSL1,FOSL2,FOS,FOXA1,FOXA2,FOXA3,FOXC1,FOXF2,FOXJ2,FOXJ3,FOXK1,FOXK2,FOXM1,FOXO1,FOXO4,FOXP1,FOXP2,FOXP4,FOXS1,FUBP1,FUBP3,FUS,GABPA,GABPB1,GATA1,GATA2,GATA3,GATA4,GATAD1,GATAD2A,GATAD2B,GFI1B,GFI1,GLI2,GLI4,GLIS1,GLIS2,GLIS3,GLYR1,GMEB1,GMEB2,GPN1,GTF2A2,GTF2B,GTF2E2,GTF2F1,GTF2I,GTF3A,GTF3C2,GZF1,HBP1,HCFC1,HDAC1,HDAC2,HDAC3,HDAC6,HDAC8,HDGF,HES1,HES2,HES4,HEYL,HHEX,HIC1,HIC2,HINFP,HIVEP1,HLF,HLTF,HMBOX1,HMG20A,HMG20B,HMGA2,HMGN3,HMGXB3,HMGXB4,HNF1A,HNF1B,HNF4A,HNF4G,HNRNPH1,HNRNPK,HNRNPLL,HNRNPL,HNRNPUL1,HOMEZ,HOXA3,HOXA5,HOXA7,HOXB13,HOXB5,HOXD1,HSF1,HSF2,HSF4,ID3,IKZF1,IKZF2,IKZF3,IKZF4,IKZF5,ILF3,INSM2,IRF1,IRF2,IRF3,IRF4,IRF5,IRF9,ISL1,ISL2,ISX,JRK,JUNB,JUND,JUN,KAT2B,KAT7,KAT8,KDM1A,KDM2A,KDM2B,KDM3A,KDM4A,KDM4B,KDM5A,KDM5B,KDM6A,KHSRP,KIAA2018,KLF10,KLF11,KLF12,KLF13,KLF16,KLF17,KLF1,KLF4,KLF5,KLF6,KLF7,KLF8,KLF9,KMT2A,KMT2B,L3MBTL2,LARP7,LBX2,LCORL,LCOR,LEF1,LIN54,MAF1,MAFF,MAFG,MAFK,MAX,MAZ,MBD1,MBD2,MBD4,MED13,MED1,MEF2A,MEF2B,MEF2C,MEF2D,MEIS1,MEIS2,MGA,MIER1,MIER2,MIER3,MITF,MIXL1,MLLT1,MLXIP,MLX,MNT,MNX1,MSX2,MTA1,MTA2,MTA3,MTF1,MTF2,MXD1,MXD3,MXD4,MXI1,MYBL2,MYB,MYC,MYNN,MYRF,MZF1,NAIF1,NANOG,NBN,NCOA1,NCOA2,NCOA3,NCOA6,NCOR1,NEUROD1,NFAT5,NFATC1,NFATC3,NFATC4,NFE2L1,NFE2L2,NFE2,NFIA,NFIB,NFIC,NFIL3,NFIX,NFKB2,NFKBIZ,NFRKB,NFXL1,NFYA,NFYB,NFYC,NKRF,NKX3-1,NONO,NR0B2,NR1H2,NR2C1,NR2C2,NR2E3,NR2F1,NR2F2,NR2F6,NR3C1,NR4A1,NR5A1,NR5A2,NRF1,NRL,NUFIP1,ONECUT1,ONECUT2,OSR2,OTX2,OVOL1,OVOL3,PAF1,PATZ1,PAWR,PAX5,PAX8,PAXIP1,PBX1,PBX2,PBX3,PCBP1,PCBP2,PHB2,PHB,PHF20,PHF21A,PHF5A,PHF8,PITX1,PKNOX1,PLSCR1,PML,POGZ,POLR2AphosphoS2,POLR2AphosphoS5,POLR2A,POLR2B,POLR2G,POLR2H,POLR3A,POU2F2,POU5F1,POU6F1,PPARG,PRDM10,PRDM15,PRDM1,PRDM4,PRDM6,PRPF4,PRRX2,PTBP1,PTRF,PTTG1,PYGO2,RAD21,RAD51,RARA,RARB,RARG,RB1,RBAK,RBBP5,RBFOX2,RBM14,RBM22,RBM25,RBM39,RBPJ,RCOR1,RCOR2,RELA,RELB,REPIN1,RERE,REST,RFX1,RFX3,RFX5,RFXANK,RFXAP,RLF,RNF219,RNF2,RORA,RREB1,RUNX1,RUNX3,RXRA,RXRB,SAFB2,SAFB,SAP130,SAP30,SATB2,SCRT1,SCRT2,SETDB1,SFPQ,SHOX2,SIN3A,SIN3B,SIRT6,SIX1,SIX4,SIX5,SKIL,SKI,SMAD1,SMAD3,SMAD4,SMAD5,SMAD9,SMARCA4,SMARCA5,SMARCB1,SMARCC1,SMARCC2,SMARCE1,SMC3,SNAI1,SNAI2,SNAPC4,SNIP1,SOX13,SOX18,SOX5,SOX6,SP110,SP140L,SP1,SP2,SP3,SP4,SP5,SP7,SPDEF,SPEN,SPI1,SREBF1,SREBF2,SRF,SRSF1,SRSF3,SRSF4,SSRP1,STAG1,STAT1,STAT3,STAT5A,STAT5B,STAT6,SUPT5H,SUZ12,TAF15,TAF1,TAF7,TAF9B,TAL1,TARDBP,TBL1XR1,TBPL1,TBP,TBX18,TBX21,TBX2,TBX3,TCF12,TCF15,TCF3,TCF4,TCF7L2,TCF7,TEAD1,TEAD2,TEAD3,TEAD4,TEF,TFAP4,TFCP2L1,TFCP2,TFDP1,TFDP2,TFE3,TGIF2,THAP11,THAP12,THAP1,THAP7,THAP8,THAP9,THRAP3,THRA,THRB,TIGD3,TIGD6,TOE1,TOX2,TOX,TP53,TP63,TRAFD1,TRIM22,TRIM24,TRIM25,TRIM28,TSC22D2,TSHZ1,TSHZ2,U2AF1,U2AF2,UBTF,USF1,USF2,VEZF1,WIZ,WRNIP1,WT1,XBP1,XRCC5,YBX1,YEATS2,YEATS4,YY1,YY2,ZBED1,ZBED4,ZBED5,ZBTB10,ZBTB11,ZBTB12,ZBTB14,ZBTB17,ZBTB1,ZBTB20,ZBTB21,ZBTB24,ZBTB25,ZBTB26,ZBTB2,ZBTB33,ZBTB37,ZBTB38,ZBTB39,ZBTB3,ZBTB40,ZBTB42,ZBTB43,ZBTB44,ZBTB46,ZBTB48,ZBTB49,ZBTB4,ZBTB5,ZBTB6,ZBTB7A,ZBTB7B,ZBTB8A,ZBTB9,ZC3H10,ZC3H11A,ZC3H4,ZC3H8,ZCCHC11,ZEB1,ZEB2,ZFHX2,ZFP14,ZFP1,ZFP36L2,ZFP36,ZFP37,ZFP3,ZFP41,ZFP64,ZFP69B,ZFP82,ZFP91,ZFX,ZFY,ZGPAT,ZHX1,ZHX2,ZHX3,ZIC2,ZIK1,ZKSCAN1,ZKSCAN5,ZKSCAN8,ZMAT3,ZMIZ1,ZMYM2,ZMYM3,ZNF101,ZNF10,ZNF121,ZNF124,ZNF12,ZNF133,ZNF134,ZNF138,ZNF140,ZNF142,ZNF143,ZNF146,ZNF148,ZNF157,ZNF165,ZNF16,ZNF175,ZNF17,ZNF180,ZNF184,ZNF189,ZNF18,ZNF197,ZNF202,ZNF205,ZNF207,ZNF20,ZNF215,ZNF217,ZNF219,ZNF221,ZNF223,ZNF224,ZNF225,ZNF230,ZNF232,ZNF234,ZNF239,ZNF24,ZNF251,ZNF256,ZNF257,ZNF25,ZNF263,ZNF264,ZNF266,ZNF26,ZNF274,ZNF276,ZNF280A,ZNF280B,ZNF280D,ZNF281,ZNF282,ZNF296,ZNF2,ZNF30,ZNF311,ZNF316,ZNF317,ZNF318,ZNF319,ZNF324,ZNF329,ZNF331,ZNF333,ZNF335,ZNF337,ZNF33A,ZNF33B,ZNF341,ZNF343,ZNF34,ZNF350,ZNF354B,ZNF354C,ZNF362,ZNF366,ZNF367,ZNF383,ZNF384,ZNF391,ZNF394,ZNF395,ZNF397,ZNF398,ZNF3,ZNF407,ZNF414,ZNF416,ZNF41,ZNF423,ZNF426,ZNF430,ZNF431,ZNF433,ZNF441,ZNF444,ZNF445,ZNF446,ZNF449,ZNF44,ZNF451,ZNF460,ZNF462,ZNF483,ZNF484,ZNF485,ZNF488,ZNF48,ZNF490,ZNF501,ZNF503,ZNF507,ZNF510,ZNF511,ZNF512B,ZNF512,ZNF513,ZNF518A,ZNF526,ZNF529,ZNF530,ZNF532,ZNF543,ZNF547,ZNF548,ZNF549,ZNF550,ZNF552,ZNF555,ZNF556,ZNF557,ZNF558,ZNF561,ZNF569,ZNF570,ZNF572,ZNF574,ZNF576,ZNF579,ZNF580,ZNF583,ZNF584,ZNF585B,ZNF589,ZNF592,ZNF596,ZNF598,ZNF600,ZNF605,ZNF607,ZNF608,ZNF609,ZNF610,ZNF614,ZNF615,ZNF616,ZNF619,ZNF623,ZNF624,ZNF629,ZNF639,ZNF644,ZNF646,ZNF652,ZNF654,ZNF660,ZNF664,ZNF671,ZNF674,ZNF677,ZNF678,ZNF680,ZNF687,ZNF691,ZNF692,ZNF697,ZNF700,ZNF703,ZNF707,ZNF709,ZNF70,ZNF710,ZNF713,ZNF737,ZNF740,ZNF75A,ZNF761,ZNF764,ZNF766,ZNF768,ZNF76,ZNF770,ZNF772,ZNF773,ZNF775,ZNF776,ZNF777,ZNF778,ZNF781,ZNF782,ZNF784,ZNF785,ZNF786,ZNF788,ZNF791,ZNF792,ZNF79,ZNF7,ZNF800,ZNF816,ZNF830,ZNF839,ZNF83,ZNF843,ZNF84,ZNF850,ZNF865,ZNF883,ZNF891,ZNF8,ZSCAN12,ZSCAN16,ZSCAN18,ZSCAN20,ZSCAN21,ZSCAN22,ZSCAN23,ZSCAN29,ZSCAN30,ZSCAN31,ZSCAN32,ZSCAN4,ZSCAN5A,ZSCAN5C,ZSCAN9,ZXDB,ZXDC,ZZZ3\ html wgEncodeReg4TfPeaks.html\ itemRgb on\ labelFields factor\ longLabel Transcription factor representative peak (rPeak) clusters from ENCODE 4\ mouseOverField factor\ priority 1.7\ scoreMax 1000\ scoreMin 1\ shortLabel TF rPeaks\ spectrum on\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4TfPeaks\ type bigBed 12 +\ urls exp="https://www.encodeproject.org/experiments/$$/" cCRE="https://screen.wenglab.org/search/?q=$$&assembly=GRCh38" factor="https://www.factorbook.org/tf/human/$$/function"\ visibility hide\ wgEncodeRegDnaseWig DNase Signal bigWig 0 10000 DNase I Hypersensitivity Signal Colored by Similarity from ENCODE 0 1.8 0 0 0 127 127 127 0 0 0\ This track provides an integrated display of DNase hypersensitivity in multiple\ cell types using overlapping colored graphs of signal density with graph colors\ assigned to cell types based on similarity of signal. The track is based on\ results of experiments performed by the John Stamatoyannapoulos lab at the\ University of Washington from September 2007 to January 2011 as part of the\ ENCODE project first production phase.
\\ The signal graphs displayed here are also included in the comprehensive\ DNaseI HS track,\ which also provides peak and region calls and uses the same coloring based on\ similiarity of cell types (please note there is different coloring on the ENCODE hg38\ Transcription track,\ Layered H3K4Me1 track,\ Layered H3K4Me3 track, and\ Layered H3K27Ac track,\ which match the coloring used in their previous versions lifted from the hg19 assembly). \
\ \\ Raw sequence data files were processed by the UCSC ENCODE DNase analysis pipeline\ described in the \ DNaseI HS\ track description.\ Signal graphs were normalized so the average value genome-wide is 1.\ Colors for the signal graphs were assigned by the UCSC BigWigCluster tool.\ \
\ The cell types were clustered into a binary tree, a rainbow was cast to the leaf nodes providing coloring based on similarity. \
\
\
Credit: Chris Eisenhart, J. Kent lab \
\ The processed data for this track were generated at UCSC.\ Credits for the primary data underlying this track are included in the\ DNaseI HS\ track description.\
\ \\ Miga KH, Eisenhart C, Kent WJ.\ \ Utilizing mapping targets of sequences underrepresented in the reference assembly to reduce false\ positive alignments.\ Nucleic Acids Res. 2015 Nov 16;43(20):e133.\ PMID: 26163063\
\\ Thurman RE, Rynes E, Humbert R, Vierstra J, Maurano MT, Haugen E, Sheffield NC, Stergachis AB, Wang\ H, Vernot B et al.\ \ The accessible chromatin landscape of the human genome.\ Nature. 2012 Sep 6;489(7414):75-82.\ PMID: 22955617; PMC: PMC3721348\
\ \\ See also the references in the\ DNaseI HS\ track.\
\ regulation 1 aggregate transparentOverlay\ configurable on\ container multiWig\ controlledVocabulary cellType=wgEncodeCell\ dimensions dimA=tissue dimB=cancer\ filterComposite dimA dimB\ group regulation\ html wgEncodeRegDnaseSignal\ longLabel DNase I Hypersensitivity Signal Colored by Similarity from ENCODE\ maxHeightPixels 128:64:11\ priority 1.8\ shortLabel DNase Signal\ showSubtrackColorOnUi on\ sortOrder subtrackColor=+ cellType=+ tissue=+\ subGroup1 cellType Cell_Type A549=A549 AG04449=AG04449 AG04450=AG04450 AG09309=AG09309 AG09319=AG09319 AG10803=AG10803 AoAF=AoAF bone_marrow_MSC=bone_marrow_MSC BE2_C=BE2_C BJ=BJ CD20_RO01778=CD20+_RO01778 Caco-2=Caco-2 GM04503=GM04503 GM04504=GM04504 GM06990=GM06990 GM12865=GM12865 GM12878=GM12878 H7-hESC=H7-hESC HA-h=HA-h HA-sp=HA-sp HAEpiC=HAEpiC HAc=HAc HBMEC=HBMEC HBVSMC=HBVSMC HCF=HCF HCFaa=HCFaa HCM=HCM HCPEpiC=HCPEpiC HCT-116=HCT-116 HConF=HConF HEEpiC=HEEpiC HFF-Myc=HFF-Myc HFF=HFF HGF=HGF HIPEpiC=HIPEpiC HL-60=HL-60 HMEC=HMEC HMF=HMF HMVEC-LBl=HMVEC-LBl HMVEC-LLy=HMVEC-LLy HMVEC-dAd=HMVEC-dAd HMVEC-dBl-Ad=HMVEC-dBl-Ad HMVEC-dBl-Neo=HMVEC-dBl-Neo HMVEC-dLy-Ad=HMVEC-dLy-Ad HMVEC-dLy-Neo=HMVEC-dLy-Neo HMVEC-dNeo=HMVEC-dNeo HNPCEpiC=HNPCEpiC HPAF=HPAF HPF=HPF HPdLF=HPdLF HRCEpiC=HRCEpiC HRE=HRE HRGEC=HRGEC HRPEpiC=HRPEpiC HSMM=HSMM HSMMtube=HSMMtube HUVEC=HUVEC HVMF=HVMF HeLa-S3=HeLa-S3 HepG2=HepG2 Jurkat=Jurkat K562=K562 LHCN-M2=LHCN-M2 LNCaP=LNCaP M059J=M059J MCF-7=MCF-7 Monocytes_CD14_RO01746=Monocytes-CD14+_RO01746 NB4=NB4 NH-A=NH-A NHBE_RA=NHBE_RA NHDF-Ad=NHDF-Ad NHDF-neo=NHDF-neo NHEK=NHEK NHLF=NHLF NT2-D1=NT2-D1 PANC-1=PANC-1 PrEC=PrEC RPMI-7951=RPMI-7951 RPTEC=RPTEC SAEC=SAEC SK-N-MC=SK-N-MC SK-N-SH_RA=SK-N-SH_RA SKMC=SKMC T-47D=T-47D Th1=Th1 Th1_Wb54553204=Th1_Wb54553204 Th2=Th2 WERI-Rb-1=WERI-Rb-1 WI-38=WI-38\ subGroup2 treatment Treatment diffProtA_5d=diffProtA_5d diffProtA_14d=diffProtA_14d DIFF_4d=DIFF_4d n_a=n/a 4OHTAM_20nM_72hr=4OHTAM_20nM_72hr Estradiol_ctrl_0hr=Estradiol_ctrl_0hr Estradiol_100nM_1hr=Estradiol_100nM_1hr\ subGroup3 tissue Tissue blood=blood blood_vessel=blood_vessel bone_marrow=bone_marrow brain=brain breast=breast cervix=cervix colon=colon embryo=embryo esophagus=esophagus eye=eye heart=heart kidney=kidney liver=liver lung=lung muscle=muscle pancreas=pancreas periodontium=periodontium periodontium=periodontium placenta=placenta prostate=prostate skin=skin spinal_cord=spinal_cord testis=testis\ subGroup4 cancer Cancer cancer=cancer normal=normal unknown=unknown\ subGroup5 subtrackColor Similarity\ superTrack wgEncodeReg hide\ track wgEncodeRegDnaseWig\ type bigWig 0 10000\ viewLimits 0:200\ visibility hide\ knownGeneV48 GENCODE V48 bigGenePred knownGenePep knownGeneMrna GENCODE V48 3 1.8 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 48, April 2025) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The v48 release was derived from the GTF file that contains annotations only on the main\ chromosomes. Statistics for this build and information on how they were generated can be found on\ the GENCODE site.
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is mostly based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \ \\ Within a gene using the pack display mode, transcripts below a specified rank will be\ condensed into a view similar to squish mode. The transcript ranking approach is\ preliminary and will change in future releases. The transcripts rankings are defined by the\ following criteria for protein-coding and non-coding genes:
\ Protein_coding genes\\
The GENCODE v48 track was built from the GENCODE downloads file \
gencode.v48.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources\
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney. This version of the track was\ generated by Jonathan Casper.
\ \\ Mudge JM, Carbonell-Sala S, Diekhans M, Martinez JG, Hunt T, Jungreis I, Loveland JE, Arnan C,\ Barnes I, Bennett R et al.\ \ GENCODE 2025: reference gene annotation for human and mouse.\ Nucleic Acids Res. 2025 Jan 6;53(D1):D966-D975.\ PMID: 39565199; PMC: PMC11701607\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV48.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV48\ group genes\ hgsid on\ html knownGeneV48\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V48\ maxItems 50000\ parent knownGeneArchive\ priority 1.8\ searchIndex name\ shortLabel GENCODE V48\ squishyPackField rank\ squishyPackLabel Number of transcripts shown at full height (ranked by GENCODE transcript ranking)\ squishyPackPoint 1\ track knownGeneV48\ type bigGenePred knownGenePep knownGeneMrna\ visibility pack\ nmdDetectiveAiBed NMDetective-AI variants bigBed 9 + NMDetective-AI: Per-stop-gain predictions for every codon (MANE Select only) 0 1.8 0 0 0 127 127 127 0 0 0\ The NMDetective-AI tracks display deep-learning predictions of\ nonsense-mediated mRNA decay (NMD) efficiency for every possible stop-gain\ single-nucleotide variant in MANE Select transcripts. The model was trained on\ ~14,000 somatic premature termination codons (PTCs) measured by allele-specific\ expression in large human cohorts (TCGA) and was tested on ~1,800 held-out\ germline PTCs (TCGA germline and GTEx) (Veiner et al.).\
\ \\ Predictions are continuous: higher values indicate that a PTC at that codon is\ predicted to trigger NMD (the mRNA is degraded); lower values indicate that the\ PTC is predicted to evade NMD (the truncated mRNA may be translated into an\ aberrant protein). The output is normalized against canonical controls so that\ +0.5 corresponds to full NMD efficiency at a PTC and −0.5\ corresponds to no NMD efficiency (a last-exon PTC). The scale is not strictly\ bounded: due to measurement and prediction noise, observed values fall in\ roughly −1.1 to +1.5, with the bulk of items inside the nominal\ −0.5 to +0.5 interval.\
\ \| Track | Description |
|---|---|
| NMDetective-AI | \Signal track (bigWig) showing the position-averaged prediction across\ all stop-gain SNVs at each codon. Useful for browsing efficiency along a\ transcript at a glance. |
| NMDetective-AI variants | \Per-stop-gain track (bigBed) with one item per (transcript, codon,\ mutant codon) combination. Each item is colored by its prediction and\ carries the reference and mutant codon, amino-acid position, transcript\ accession, and a pre-rendered mouseover summary. |
\ The NMDetective-AI signal track is drawn with a default y-axis range of\ −1.1 to +1.5. Positions with positive values (predicted NMD-triggering)\ are shown above the baseline; positions with negative values (predicted NMD\ escape) are shown below.\
\ \\ The NMDetective-AI variants track colors each item along a continuous\ diverging Okabe-Ito palette running from blue (most NMD-evading) through grey\ (near zero) to vermillion (most NMD-triggering). The mouseover verdict groups\ items into three categories using the binarization thresholds derived in the\ Veiner et al. Methods (Gaussian mixture model fit to gnomAD\ predictions):\
\\ Mouseover for each variant shows the codon change, the prediction value with\ its NMD verdict, and the MANE Select transcript accession. Click an item to\ see the full set of fields on the details page.\
\ \\ NMDetective-AI is a fine-tuned version of the Orthrus mRNA foundation model\ (Mamba architecture, ~10M parameters), trained on full-length transcript\ sequences encoded as a six-track representation (four nucleotide channels,\ one CDS-start channel, one splice-site channel). The model integrates\ allele-specific PTC expression from large-scale genomic data with mRNA\ language-model embeddings and high-throughput deep mutational scanning, and\ predicts NMD efficiency for every possible stop-gain mutation in every codon\ of a MANE Select transcript.\
\ \\ The training set comprised 14,337 somatic PTCs from TCGA, with chromosomes 1\ and 20 held out as a validation set. The held-out test set comprised 1,065\ germline PTCs from TCGA and 763 germline PTCs from GTEx. The authors report\ that the model's accuracy on the somatic validation set approaches the\ empirical reproducibility ceiling of the underlying allele-specific\ expression measurements.\
\ \\ The publicly released predictions cover MANE Select transcripts at Gencode\ v46. Predictions for transcripts outside the MANE Select set are not yet\ available; broader coverage is planned by the authors after peer review.\
\ \\ Source files were obtained from the\ Vejni/NMDetectiveAI\ GitHub repository (supplementary files\ NMDetectiveAI_MANE.bw.gz and NMDetectiveAI_MANE.bed.gz) and\ processed at UCSC: the bigWig is used as supplied; the BED was recolored with\ the diverging Okabe-Ito palette described above, rescored into the\ 0–1000 BED range, and augmented with a pre-rendered mouseover column\ before conversion to bigBed.\
\ \\ Note: the manuscript is currently a bioRxiv preprint and has not yet\ completed peer review. Predictions may be refreshed when the final version\ of the data is released.\
\ \\ The data underlying these tracks can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Marcell Veiner and Fran Supek for sharing the NMDetective-AI\ predictions ahead of publication, and to the wider Veiner et al.\ author group for developing the model.\
\ \\ Veiner M, Toledano I, Palou-Márquez G, Lehner B, Supek F.\ \ Quantitative prediction of nonsense-mediated mRNA decay across human genes by\ genomic language model and large-scale mutational scanning.\ bioRxiv. 2026 Mar 26.\ doi: 10.64898/2026.03.24.714003.\ Supplementary prediction files at\ github.com/Vejni/NMDetectiveAI.\
\ genes 1 bigDataUrl /gbdb/hg38/nmd/nmdDetectAi.bb\ html nmdDetectiveAi\ itemRgb on\ longLabel NMDetective-AI: Per-stop-gain predictions for every codon (MANE Select only)\ mouseOverField mouseOver\ parent nmd off\ priority 1.8\ shortLabel NMDetective-AI variants\ track nmdDetectiveAiBed\ type bigBed 9 +\ visibility hide\ wgEncodeRegDnase DNase HS bed 3 + DNase I Hypersensitivity in 95 cell types from ENCODE 0 1.9 0 0 0 127 127 127 0 0 0\ These tracks contain the results of DNase I hypersensitivity experiments performed by the\ John Stamatoyannapoulos lab\ at the University of Washington from September 2007 to January 2011, as part of the\ ENCODE project first production phase.\ Colors were assigned to cell types based on similarity of signal.\
\ \\ Other views of this data (along with additional documentation) are available from the hg19\ ENCODE UW DNaseI HS track.\
\ \\ This track is a composite annotation track containing multiple subtracks, one for each cell type.\ The display mode and filtering of each subtrack can be individually controlled. \ For more information about track configuration, see\ Configuring Multi-View Tracks.\
\ \\ Raw sequence data files were processed by the UCSC ENCODE DNase analysis pipeline (July 2014 specification), diagrammed here:\
\
Credit: Qian Alvin Qin, X. Liu lab\
\ Briefly, sequence files were aligned to the hg38 (GRCh38) genome assembly augmented with 'sponge'\ sequence (ref). Multi-mapped reads were removed, as were reads that aligned to 'sponge' or\ mitochondrial sequence. Results from all replicates were pooled, and further processed by\ the Hotspot program to call peaks as well as broader regions of activity ('hotspots'), and to\ create signal density graphs.\ Signal graphs were normalized so the average value genome-wide is 1.\
\\ The cell types were clustered into a binary tree, a rainbow was cast to the leaf nodes providing coloring based on similarity.\
\
\
Credit: Chris Eisenhart, J. Kent lab \
\ The processed data for this track were produced by UCSC. Credits for the primary data \ underlying this track are included in the\ ENCODE UW DNaseI HS track\ description.\
\ \\ Miga KH, Eisenhart C, Kent WJ.\ \ Utilizing mapping targets of sequences underrepresented in the reference assembly to reduce false\ positive alignments.\ Nucleic Acids Res. 2015 Nov 16;43(20):e133.\ PMID: 26163063\
\\ Thurman RE, Rynes E, Humbert R, Vierstra J, Maurano MT, Haugen E, Sheffield NC, Stergachis AB, Wang\ H, Vernot B et al.\ \ The accessible chromatin landscape of the human genome.\ Nature. 2012 Sep 6;489(7414):75-82.\ PMID: 22955617; PMC: PMC3721348\
\ \\ See also the references in the\ ENCODE UW DNaseI HS\ track.\
\ regulation 1 compositeTrack on\ controlledVocabulary cellType=wgEncodeCell\ dimensions dimA=cellType dimB=tissue dimC=cancer\ dragAndDrop subTracks\ filterComposite dimA dimB dimC\ group regulation\ html wgEncodeRegDnase\ longLabel DNase I Hypersensitivity in 95 cell types from ENCODE\ noInherit on\ priority 1.9\ shortLabel DNase HS\ showSubtrackColorOnUi on\ sortOrder view=+ subtrackColor=+ cellType=+ tissue=+\ subGroup1 view Views a_Peaks=Peaks b_Hot=Hotspots c_Signal=Signal\ subGroup2 cellType Cell_Type A549=A549 AG04449=AG04449 AG04450=AG04450 AG09309=AG09309 AG09319=AG09319 AG10803=AG10803 AoAF=AoAF bone_marrow_MSC=bone_marrow_MSC BE2_C=BE2_C BJ=BJ CD20_RO01778=CD20+_RO01778 Caco-2=Caco-2 GM04503=GM04503 GM04504=GM04504 GM06990=GM06990 GM12865=GM12865 GM12878=GM12878 H7-hESC=H7-hESC HA-h=HA-h HA-sp=HA-sp HAEpiC=HAEpiC HAc=HAc HBMEC=HBMEC HBVSMC=HBVSMC HCF=HCF HCFaa=HCFaa HCM=HCM HCPEpiC=HCPEpiC HCT-116=HCT-116 HConF=HConF HEEpiC=HEEpiC HFF-Myc=HFF-Myc HFF=HFF HGF=HGF HIPEpiC=HIPEpiC HL-60=HL-60 HMEC=HMEC HMF=HMF HMVEC-LBl=HMVEC-LBl HMVEC-LLy=HMVEC-LLy HMVEC-dAd=HMVEC-dAd HMVEC-dBl-Ad=HMVEC-dBl-Ad HMVEC-dBl-Neo=HMVEC-dBl-Neo HMVEC-dLy-Ad=HMVEC-dLy-Ad HMVEC-dLy-Neo=HMVEC-dLy-Neo HMVEC-dNeo=HMVEC-dNeo HNPCEpiC=HNPCEpiC HPAF=HPAF HPF=HPF HPdLF=HPdLF HRCEpiC=HRCEpiC HRE=HRE HRGEC=HRGEC HRPEpiC=HRPEpiC HSMM=HSMM HSMMtube=HSMMtube HUVEC=HUVEC HVMF=HVMF HeLa-S3=HeLa-S3 HepG2=HepG2 Jurkat=Jurkat K562=K562 LHCN-M2=LHCN-M2 LNCaP=LNCaP M059J=M059J MCF-7=MCF-7 Monocytes_CD14_RO01746=Monocytes-CD14+_RO01746 NB4=NB4 NH-A=NH-A NHBE_RA=NHBE_RA NHDF-Ad=NHDF-Ad NHDF-neo=NHDF-neo NHEK=NHEK NHLF=NHLF NT2-D1=NT2-D1 PANC-1=PANC-1 PrEC=PrEC RPMI-7951=RPMI-7951 RPTEC=RPTEC SAEC=SAEC SK-N-MC=SK-N-MC SK-N-SH_RA=SK-N-SH_RA SKMC=SKMC T-47D=T-47D Th1=Th1 Th1_Wb54553204=Th1_Wb54553204 Th2=Th2 WERI-Rb-1=WERI-Rb-1 WI-38=WI-38\ subGroup3 treatment Treatment diffProtA_5d=diffProtA_5d diffProtA_14d=diffProtA_14d DIFF_4d=DIFF_4d n_a=n/a OHTAM_20nM_72hr=4OHTAM_20nM_72hr Estradiol_ctrl_0hr=Estradiol_ctrl_0hr Estradiol_100nM_1hr=Estradiol_100nM_1hr\ subGroup4 tissue Tissue blood=blood blood_vessel=blood_vessel bone_marrow=bone_marrow brain=brain breast=breast cervix=cervix colon=colon embryo=embryo esophagus=esophagus eye=eye heart=heart kidney=kidney liver=liver lung=lung muscle=muscle pancreas=pancreas periodontium=periodontium periodontium=periodontium placenta=placenta prostate=prostate skin=skin spinal_cord=spinal_cord testis=testis\ subGroup5 cancer Cancer cancer=cancer normal=normal unknown=unknown\ subGroup6 subtrackColor Similarity\ superTrack wgEncodeReg hide\ track wgEncodeRegDnase\ type bed 3 +\ knownGeneV47 GENCODE V47 bigGenePred knownGenePep knownGeneMrna GENCODE V47 3 1.9 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 47, October 2024) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The v47 release was derived from the GTF file that contains annotations only on the main\ chromosomes. Statistics for this build and information on how they were generated can be found on\ the GENCODE site.
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is mostly based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \ \\ Within a gene using the pack display mode, transcripts below a specified rank will be\ condensed into a view similar to squish mode. The transcript ranking approach is\ preliminary and will change in future releases. The transcripts rankings are defined by the\ following criteria for protein-coding and non-coding genes:
\ Protein_coding genes\\
The GENCODE v47 track was built from the GENCODE downloads file \
gencode.v47.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources\
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney. This version of the track was\ generated by Jonathan Casper.
\ \\ Frankish A, Carbonell-Sala S, Diekhans M, Jungreis I, Loveland JE, Mudge JM, Sisu C, Wright JC,\ Arnan C, Barnes I et al.\ \ GENCODE: reference annotation for the human and mouse genomes in 2023.\ Nucleic Acids Res. 2023 Jan 6;51(D1):D942-D949.\ PMID: 36420896; PMC: PMC9825462\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV47.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV47\ group genes\ html knownGeneV47\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V47\ maxItems 50000\ parent knownGeneArchive\ priority 1.9\ searchIndex name\ shortLabel GENCODE V47\ squishyPackField rank\ squishyPackLabel Number of transcripts shown at full height (ranked by GENCODE transcript ranking)\ squishyPackPoint 1\ track knownGeneV47\ type bigGenePred knownGenePep knownGeneMrna\ visibility pack\ wgEncodeRegDnaseHotspot Hotspots bed 3 + Hotspot5 hotspot calls on BWA. Dupe, sponge and mitochondria filtered 0 1.9 0 0 0 127 127 127 1 0 0 regulation 1 longLabel Hotspot5 hotspot calls on BWA. Dupe, sponge and mitochondria filtered\ minGrayLevel 2\ parent wgEncodeRegDnase\ scoreFilter 0\ scoreFilterLimits 0:1000\ shortLabel Hotspots\ spectrum on\ track wgEncodeRegDnaseHotspot\ view b_Hot\ visibility hide\ wgEncodeRegDnasePeak Peaks narrowPeak HotSpot5 peak calls on BWA. Dupe, sponge and mitochondria filtered 1 1.9 0 0 0 127 127 127 1 0 0 regulation 1 longLabel HotSpot5 peak calls on BWA. Dupe, sponge and mitochondria filtered\ minGrayLevel 2\ parent wgEncodeRegDnase\ scoreFilter 0\ scoreFilterLimits 0:1000\ shortLabel Peaks\ spectrum on\ track wgEncodeRegDnasePeak\ type narrowPeak\ view a_Peaks\ visibility dense\ wgEncodeRegDnaseSignal Signal bed 3 + HotSpot5 signal on BWA. Dupe, sponge and mitochondria filtered 0 1.9 0 0 0 127 127 127 0 0 0\ This track provides an integrated display of DNase hypersensitivity in multiple\ cell types using overlapping colored graphs of signal density with graph colors\ assigned to cell types based on similarity of signal. The track is based on\ results of experiments performed by the John Stamatoyannapoulos lab at the\ University of Washington from September 2007 to January 2011 as part of the\ ENCODE project first production phase.
\\ The signal graphs displayed here are also included in the comprehensive\ DNaseI HS track,\ which also provides peak and region calls and uses the same coloring based on\ similiarity of cell types (please note there is different coloring on the ENCODE hg38\ Transcription track,\ Layered H3K4Me1 track,\ Layered H3K4Me3 track, and\ Layered H3K27Ac track,\ which match the coloring used in their previous versions lifted from the hg19 assembly). \
\ \\ Raw sequence data files were processed by the UCSC ENCODE DNase analysis pipeline\ described in the \ DNaseI HS\ track description.\ Signal graphs were normalized so the average value genome-wide is 1.\ Colors for the signal graphs were assigned by the UCSC BigWigCluster tool.\ \
\ The cell types were clustered into a binary tree, a rainbow was cast to the leaf nodes providing coloring based on similarity. \
\
\
Credit: Chris Eisenhart, J. Kent lab \
\ The processed data for this track were generated at UCSC.\ Credits for the primary data underlying this track are included in the\ DNaseI HS\ track description.\
\ \\ Miga KH, Eisenhart C, Kent WJ.\ \ Utilizing mapping targets of sequences underrepresented in the reference assembly to reduce false\ positive alignments.\ Nucleic Acids Res. 2015 Nov 16;43(20):e133.\ PMID: 26163063\
\\ Thurman RE, Rynes E, Humbert R, Vierstra J, Maurano MT, Haugen E, Sheffield NC, Stergachis AB, Wang\ H, Vernot B et al.\ \ The accessible chromatin landscape of the human genome.\ Nature. 2012 Sep 6;489(7414):75-82.\ PMID: 22955617; PMC: PMC3721348\
\ \\ See also the references in the\ DNaseI HS\ track.\
\ regulation 1 autoScale off\ longLabel HotSpot5 signal on BWA. Dupe, sponge and mitochondria filtered\ maxHeightPixels 100:32:16\ maxLimit 100000\ minLimit 0\ parent wgEncodeRegDnase\ shortLabel Signal\ track wgEncodeRegDnaseSignal\ view c_Signal\ viewLimits 0:100\ visibility hide\ windowingFunction mean+whiskers\ encRegTfbsClustered TF Clusters factorSource Transcription Factor ChIP-seq Clusters (340 factors, 129 cell types) from ENCODE 3 0 1.9 0 0 0 127 127 127 1 0 0 http://www.factorbook.org/mediawiki/index.php/$$\ This track shows regions of transcription factor binding derived from a large collection\ of ChIP-seq experiments performed by the ENCODE project between February 2011 and November 2018,\ spanning the first production phase of ENCODE ("ENCODE 2") through the second full production\ phase ("ENCODE 3").\
\\ Transcription factors (TFs) are proteins that bind to DNA and interact with RNA polymerases to\ regulate gene expression. Some TFs contain a DNA binding domain and can bind directly to \ specific short DNA sequences ('motifs');\ others bind to DNA indirectly through interactions with TFs containing a DNA binding domain.\ High-throughput antibody capture and sequencing methods (e.g. chromatin immunoprecipitation\ followed by sequencing, or 'ChIP-seq') can be used to identify regions of\ TF binding genome-wide. These regions are commonly called ChIP-seq peaks.
\\ ENCODE TF ChIP-seq data were processed using the \ ENCODE Transcription Factor ChIP-seq Processing Pipeline to generate peaks of TF binding.\ Peaks from 1264 experiments (1256 in hg38) representing 338 transcription factors \ (340 in hg38) in 130 cell types (129 in hg38) are combined here into clusters to produce a \ summary display showing occupancy regions for each factor.\ The underlying ChIP-seq peak data are available from the\ ENCODE 3 TF ChIP Peaks tracks (\ hg19,\ hg38)
\ \\ A gray box encloses each peak cluster of transcription factor occupancy, with the\ darkness of the box being proportional to the maximum signal strength observed in any cell type\ contributing to the cluster. The HGNC gene name for the transcription factor is shown \ to the left of each cluster.
\
\ To the right of the cluster a configurable label can optionally display information about the\ cell types contributing to the cluster and how many cell types were assayed for the factor\ (count where detected / count where assayed).\ For brevity in the display, each cell type is abbreviated to a single letter.\ The darkness of the letter is proportional to the signal strength observed in the cell line. \ Abbreviations starting with capital letters designate\ ENCODE cell types initially identified for intensive study, \ while those starting with lowercase letters designate cell lines added later in the project.
\\ Click on a peak cluster to see more information about the TF/cell assays contributing to the\ cluster and the cell line abbreviation table.\
\ \\ Peaks of transcription factor occupancy ("optimal peak set") from ENCODE ChIP-seq datasets\ were clustered using the UCSC hgBedsToBedExps tool. \ Scores were assigned to peaks by multiplying the input signal values by a normalization\ factor calculated as the ratio of the maximum score value (1000) to the signal value at one\ standard deviation from the mean, with values exceeding 1000 capped at 1000. This has the\ effect of distributing scores up to mean plus one 1 standard deviation across the score range,\ but assigning all above to the maximum score.\ The cluster score is the highest score for any peak contributing to the cluster.
\ \\ The raw data for the ENCODE3 TF Clusters track can be accessed from the\ \ Table Browser or combined with other datasets through the \ Data Integrator. This data is stored internally as a BED5+3 MySQL table with additional \ metadata tables. For automated analysis and download, the \ encRegTfbsClusteredWithCells.hg38.bed.gz track data file can be downloaded from \ our \ downloads server, which has 5 fields of BED data followed by a comma-separated list of cell types. \ The data can also be queried using the \ JSON API or the\ Public SQL server.
\ \\ Thanks to the ENCODE Consortium, the ENCODE ChIP-seq production laboratories, and the\ ENCODE Data Coordination Center for generating and processing the TF ChIP-seq datasets used here.\ The ENCODE accession numbers of the constituent datasets are available from the peak details page.\ Special thanks to Henry Pratt, Jill Moore, Michael Purcaro, and Zhiping Weng, PI, at the \ ENCODE Data Analysis Center\ (ZLab at UMass Medical Center) for providing the peak datasets, metadata,\ and guidance developing this track. Please check the\ ZLab ENCODE Public Hubs\ for the most updated data.\
\\ The integrative view presented here was developed by Jim Kent at UCSC.
\ \ENCODE Project Consortium.\ \ A user's guide to the encyclopedia of DNA elements (ENCODE).\ PLoS Biol. 2011 Apr;9(4):e1001046. PMID: 21526222; PMCID: PMC3079585\
\ \ENCODE Project Consortium.\ \ An integrated encyclopedia of DNA elements in the human genome.\ Nature. 2012 Sep 6;489(7414):57-74. PMID: 22955616; PMCID: PMC3439153\
\\ Sloan CA, Chan ET, Davidson JM, Malladi VS, Strattan JS, Hitz BC, Gabdank I, Narayanan AK, Ho M, Lee\ BT et al.\ \ ENCODE data at the ENCODE portal.\ Nucleic Acids Res. 2016 Jan 4;44(D1):D726-32.\ PMID: 26527727; PMC: PMC4702836\
\\ Gerstein MB, Kundaje A, Hariharan M, Landt SG, Yan KK, Cheng C, Mu XJ, Khurana E, Rozowsky J,\ Alexander R et al.\ \ Architecture of the human regulatory network derived from ENCODE data.\ Nature. 2012 Sep 6;489(7414):91-100.\ PMID: 22955619\
\\ Wang J, Zhuang J, Iyer S, Lin X, Whitfield TW, Greven MC, Pierce BG, Dong X, Kundaje A, Cheng Y\ et al.\ \ Sequence features and chromatin structure around the genomic regions bound by 119 human\ transcription factors.\ Genome Res. 2012 Sep;22(9):1798-812.\ PMID: 22955990; PMC: PMC3431495\
\\ Wang J, Zhuang J, Iyer S, Lin XY, Greven MC, Kim BH, Moore J, Pierce BG, Dong X, Virgil D et\ al.\ \ Factorbook.org: a Wiki-based database for transcription factor-binding data generated by the ENCODE\ consortium.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D171-6.\ PMID: 23203885; PMC: PMC3531197\
\ \Users may freely download, analyze and publish results based on any ENCODE data without \ restrictions.\ Researchers using unpublished ENCODE data are encouraged to contact the data producers to discuss possible coordinated publications; however, this is optional.
\ Users of ENCODE datasets are requested to cite the ENCODE Consortium and ENCODE\ production laboratory(s) that generated the datasets used, as described in\ Citing ENCODE.\ regulation 1 dataVersion ENCODE 3 Nov 2018\ filterBy name:factor=AFF1,AGO1,AGO2,ARHGAP35,ARID1B,ARID2,ARID3A,ARNT,ASH1L,ASH2L,ATF2,ATF3,ATF4,ATF7,ATM,BACH1,BATF,BCL11A,BCL3,BCOR,BHLHE40,BMI1,BRCA1,BRD4,BRD9,C11orf30,CBFA2T2,CBFA2T3,CBFB,CBX1,CBX2,CBX3,CBX5,CBX8,CC2D1A,CCAR2,CDC5L,CEBPB,CHAMP1,CHD1,CHD4,CHD7,CLOCK,COPS2,CREB1,CREB3L1,CREBBP,CREM,CTBP1,CTCF,CUX1,DACH1,DEAF1,DNMT1,DPF2,E2F1,E2F4,E2F6,E2F7,E2F8,E4F1,EBF1,EED,EGR1,EHMT2,ELF1,ELF4,ELK1,EP300,EP400,ESR1,ESRRA,ETS1,ETV4,ETV6,EWSR1,EZH2,FIP1L1,FOS,FOSL1,FOSL2,FOXA1,FOXA2,FOXK2,FOXM1,FOXP1,FUS,GABPA,GABPB1,GATA1,GATA2,GATA3,GATA4,GATAD2A,GATAD2B,GMEB1,HCFC1,HDAC1,HDAC2,HDAC3,HDAC6,HES1,HMBOX1,HNF1A,HNF4A,HNF4G,HNRNPH1,HNRNPK,HNRNPL,HNRNPLL,HNRNPUL1,HSF1,IKZF1,IKZF2,IRF1,IRF2,IRF3,IRF4,IRF5,JUN,JUNB,JUND,KAT2A,KAT2B,KAT8,KDM1A,KDM4A,KDM4B,KDM5A,KDM5B,KLF16,KLF5,L3MBTL2,LCORL,LEF1,MAFF,MAFK,MAX,MBD2,MCM2,MCM3,MCM5,MCM7,MEF2A,MEF2B,MEF2C,MEIS2,MGA,MIER1,MITF,MLLT1,MNT,MTA1,MTA2,MTA3,MXI1,MYB,MYBL2,MYC,MYNN,NANOG,NBN,NCOA1,NCOA2,NCOA3,NCOA4,NCOA6,NCOR1,NEUROD1,NFATC1,NFATC3,NFE2,NFE2L2,NFIB,NFIC,NFRKB,NFXL1,NFYA,NFYB,NR0B1,NR2C1,NR2C2,NR2F1,NR2F2,NR2F6,NR3C1,NRF1,NUFIP1,PAX5,PAX8,PBX3,PCBP1,PCBP2,PHB2,PHF20,PHF21A,PHF8,PKNOX1,PLRG1,PML,POLR2A,POLR2G,POU2F2,PRDM10,PRPF4,PTBP1,PYGO2,RAD21,RAD51,RB1,RBBP5,RBFOX2,RBM14,RBM15,RBM17,RBM22,RBM25,RBM34,RBM39,RCOR1,RELB,REST,RFX1,RFX5,RLF,RNF2,RUNX1,RUNX3,RXRA,SAFB,SAFB2,SAP30,SETDB1,SIN3A,SIN3B,SIRT6,SIX4,SIX5,SKI,SKIL,SMAD1,SMAD2,SMAD5,SMARCA4,SMARCA5,SMARCB1,SMARCC2,SMARCE1,SMC3,SNRNP70,SOX13,SOX6,SP1,SPI1,SREBF1,SREBF2,SRF,SRSF4,SRSF7,SRSF9,STAT1,STAT2,STAT3,STAT5A,SUPT20H,SUZ12,TAF1,TAF15,TAF7,TAF9B,TAL1,TBL1XR1,TBP,TBX21,TBX3,TCF12,TCF7,TCF7L2,TEAD4,TFAP4,THAP1,THRA,TRIM22,TRIM24,TRIM28,TRIP13,U2AF1,U2AF2,UBTF,USF1,USF2,WHSC1,WRNIP1,XRCC3,XRCC5,YY1,ZBED1,ZBTB1,ZBTB11,ZBTB2,ZBTB33,ZBTB40,ZBTB5,ZBTB7A,ZBTB7B,ZBTB8A,ZEB1,ZEB2,ZFP91,ZFX,ZHX1,ZHX2,ZKSCAN1,ZMIZ1,ZMYM3,ZNF143,ZNF184,ZNF207,ZNF217,ZNF24,ZNF263,ZNF274,ZNF280A,ZNF282,ZNF316,ZNF318,ZNF384,ZNF407,ZNF444,ZNF507,ZNF512B,ZNF574,ZNF579,ZNF592,ZNF639,ZNF687,ZNF8,ZNF830,ZSCAN29,ZZZ3\ idInUrlSql select value from factorbookGeneAlias where name='%s'\ inputTableFieldDisplay cellType factor experiment lab\ inputTableFieldUrls experiment="https://www.encodeproject.org/experiments/$$"\ inputTrackTable encRegTfbsClusteredInputs\ longLabel Transcription Factor ChIP-seq Clusters (340 factors, 129 cell types) from ENCODE 3\ maxWindowToDraw 10000000\ parent wgEncodeReg\ priority 1.90\ shortLabel TF Clusters\ sourceTable encRegTfbsClusteredSources\ track encRegTfbsClustered\ type factorSource\ url http://www.factorbook.org/mediawiki/index.php/$$\ urlLabel Factorbook Link:\ useScore 1\ visibility hide\ encTfChipPk TF ChIP narrowPeak Transcription Factor ChIP-seq Peaks (340 factors in 129 cell types) from ENCODE 3 0 1.91 0 0 0 127 127 127 0 0 0\ This track represents a comprehensive set of human transcription factor binding sites based on \ ChIP-seq experiments generated by production groups in the ENCODE Consortium between \ February 2011 and November 2018.
\\ Transcription factors (TFs) are proteins that bind to DNA and interact with RNA polymerases to\ regulate gene expression. Some TFs contain a DNA binding domain and can bind directly to \ specific short DNA sequences ('motifs');\ others bind to DNA indirectly through interactions with TFs containing a DNA binding domain.\ High-throughput antibody capture and sequencing methods (e.g. chromatin immunoprecipitation\ followed by sequencing, or 'ChIP-seq') can be used to identify regions of\ TF binding genome-wide. These regions are commonly called ChIP-seq peaks.
\ \ The related\ Transcription Factor ChIP-seq Clusters tracks \ (hg19,\ hg38)\ provide summary views of this data.\ \\ \
\ The display for this track shows site location with the point-source of the peak marked with a \ colored vertical bar and the level of enrichment at the site indicated by the darkness of the item.\ The subtracks are colored by UCSC ENCODE 2 cell type color conventions on the hg19 assembly, \ and by similarity of cell types in DNaseI hypersensitivity assays (as in the\ DNase Signal)\ track in the hg38 assembly.
\ \ The display can be filtered to higher valued items, using the \ Score range: configuration item.\ The score values were computed at UCSC based on signal values assigned by the ENCODE\ pipeline.\ The input signal values were multiplied by a normalization factor calculated as the ratio\ of the maximum score value (1000) to the signal value at 1 standard deviation from the mean,\ with values exceeding 1000 capped at 1000. This has the effect of distributing scores up to \ mean + 1std across the score range, but assigning all above to the maximum score.\ \
\ The ChIP-seq peaks in this track were\ generated by the\ the ENCODE Transcription Factor ChIP-seq Processing Pipeline.\ Methods documentation and full metadata for each track can be found at the \ ENCODE project portal, using\ The ENCODE file accession (ENCFF*) listed in the track label.\
\ \\ Thanks to the ENCODE Consortium, the ENCODE ChIP-seq production laboratories, and the\ ENCODE Data Coordination Center for generating and processing the datasets used here.\ Special thanks to Henry Pratt, Jill Moore, Michael Purcaro, and Zhiping Weng, PI, at the \ ENCODE Data Analysis Center\ (ZLab at UMass Medical Center) for providing the peak datasets, metadata,\ and guidance developing this track. Please check the\ ZLab ENCODE Public Hubs\ for the most updated data.\
\ \ENCODE Project Consortium.\ \ A user's guide to the encyclopedia of DNA elements (ENCODE).\ PLoS Biol. 2011 Apr;9(4):e1001046. PMID: 21526222; PMCID: PMC3079585\
\ \ENCODE Project Consortium.\ \ An integrated encyclopedia of DNA elements in the human genome.\ Nature. 2012 Sep 6;489(7414):57-74. PMID: 22955616; PMCID: PMC3439153\
\\ Sloan CA, Chan ET, Davidson JM, Malladi VS, Strattan JS, Hitz BC, Gabdank I, Narayanan AK, Ho M, Lee\ BT et al.\ \ ENCODE data at the ENCODE portal.\ Nucleic Acids Res. 2016 Jan 4;44(D1):D726-32.\ PMID: 26527727; PMC: PMC4702836\
\\ Gerstein MB, Kundaje A, Hariharan M, Landt SG, Yan KK, Cheng C, Mu XJ, Khurana E, Rozowsky J,\ Alexander R et al.\ \ Architecture of the human regulatory network derived from ENCODE data.\ Nature. 2012 Sep 6;489(7414):91-100.\ PMID: 22955619\
\\ Wang J, Zhuang J, Iyer S, Lin X, Whitfield TW, Greven MC, Pierce BG, Dong X, Kundaje A, Cheng Y\ et al.\ \ Sequence features and chromatin structure around the genomic regions bound by 119 human\ transcription factors.\ Genome Res. 2012 Sep;22(9):1798-812.\ PMID: 22955990; PMC: PMC3431495\
\\ Wang J, Zhuang J, Iyer S, Lin XY, Greven MC, Kim BH, Moore J, Pierce BG, Dong X, Virgil D et\ al.\ \ Factorbook.org: a Wiki-based database for transcription factor-binding data generated by the ENCODE\ consortium.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D171-6.\ PMID: 23203885; PMC: PMC3531197\
\ \Users may freely download, analyze and publish results based on any ENCODE data without \ restrictions.\ Researchers using unpublished ENCODE data are encouraged to contact the data producers to discuss possible coordinated publications; however, this is optional.
\\ Users of ENCODE datasets are requested to cite the ENCODE Consortium and ENCODE \ production laboratory(s) that generated the datasets used, as described in\ Citing ENCODE.
\ \ regulation 1 compositeTrack on\ darkerLabels on\ dataVersion ENCODE 3 Nov 2018\ dimensions dimX=cellType dimY=factor\ dragAndDrop subTracks\ group regulation\ longLabel Transcription Factor ChIP-seq Peaks (340 factors in 129 cell types) from ENCODE 3\ parent wgEncodeReg\ priority 1.91\ scoreFilter 0\ scoreFilterLimits 0:1000\ shortLabel TF ChIP\ sortOrder cellType=+ factor=+\ subGroup1 cellType Cell_Type X22Rv1=22Rv1 A549=A549 A673=A673 AG04449=AG04449 AG04450=AG04450 AG09309=AG09309 AG09319=AG09319 AG10803=AG10803 BE2C=BE2C BJ=BJ B_cell=B_cell C4-2B=C4-2B CD14-positive_monocyte=CD14-positive_monocyte Caco-2=Caco-2 DOHH2=DOHH2 GM06990=GM06990 GM08714=GM08714 GM10266=GM10266 GM12864=GM12864 GM12865=GM12865 GM12873=GM12873 GM12874=GM12874 GM12878=GM12878 GM12891=GM12891 GM12892=GM12892 GM13977=GM13977 GM20000=GM20000 GM23248=GM23248 GM23338=GM23338 H1-hESC=H1-hESC H54=H54 HCT116=HCT116 HEK293=HEK293 HEK293T=HEK293T HFF-Myc=HFF-Myc HL-60=HL-60 HeLa-S3=HeLa-S3 HepG2=HepG2 IMR-90=IMR-90 Ishikawa=Ishikawa K562=K562 KMS-11=KMS-11 LNCAP=LNCAP LNCaP_clone_FGC=LNCaP_clone_FGC Loucy=Loucy MCF-7=MCF-7 MCF_10A=MCF_10A MM_1S=MM.1S NB4=NB4 NCI-H929=NCI-H929 NT2_D1=NT2/D1 OCI-LY1=OCI-LY1 OCI-LY3=OCI-LY3 OCI-LY7=OCI-LY7 PC-3=PC-3 PC-9=PC-9 PFSK-1=PFSK-1 Panc1=Panc1 Parathyroid_adenoma=Parathyroid_adenoma Peyers_patch=Peyer's_patch RWPE1=RWPE1 RWPE2=RWPE2 Raji=Raji SH-SY5Y=SH-SY5Y SK-N-MC=SK-N-MC SK-N-SH=SK-N-SH SU-DHL-6=SU-DHL-6 T47D=T47D VCaP=VCaP WERI-Rb-1=WERI-Rb-1 WI38=WI38 adrenal_gland=adrenal_gland ascending_aorta=ascending_aorta astrocyte=astrocyte astrocyte_of_the_cerebellum=astrocyte_of_the_cerebellum astrocyte_of_the_spinal_cord=astrocyte_of_the_spinal_cord bipolar_neuron=bipolar_neuron body_of_pancreas=body_of_pancreas brain_microvascular_endothelial_cell=brain_microvascular_endothelial_cell breast_epithelium=breast_epithelium cardiac_fibroblast=cardiac_fibroblast cardiac_muscle_cell=cardiac_muscle_cell choroid_plexus_epithelial_cell=choroid_plexus_epithelial_cell endothelial_cell_of_umbilical_vein=endothelial_cell_of_umbilical_vein epithelial_cell_of_esophagus=epithelial_cell_of_esophagus epithelial_cell_of_prostate=epithelial_cell_of_prostate erythroblast=erythroblast esophagus_muscularis_mucosa=esophagus_muscularis_mucosa esophagus_squamous_epithelium=esophagus_squamous_epithelium fibroblast_of_lung=fibroblast_of_lung fibroblast_of_mammary_gland=fibroblast_of_mammary_gland fibroblast_of_pulmonary_artery=fibroblast_of_pulmonary_artery fibroblast_of_the_aortic_adventitia=fibroblast_of_the_aortic_adventitia fibroblast_of_villous_mesenchyme=fibroblast_of_villous_mesenchyme foreskin_fibroblast=foreskin_fibroblast foreskin_keratinocyte=foreskin_keratinocyte gastrocnemius_medialis=gastrocnemius_medialis gastroesophageal_sphincter=gastroesophageal_sphincter heart_left_ventricle=heart_left_ventricle hepatocyte=hepatocyte keratinocyte=keratinocyte kidney_epithelial_cell=kidney_epithelial_cell liver=liver lower_leg_skin=lower_leg_skin mammary_epithelial_cell=mammary_epithelial_cell medulloblastoma=medulloblastoma myotube=myotube neural_cell=neural_cell neural_progenitor_cell=neural_progenitor_cell neutrophil=neutrophil omental_fat_pad=omental_fat_pad ovary=ovary prostate_gland=prostate_gland retinal_pigment_epithelial_cell=retinal_pigment_epithelial_cell right_lobe_of_liver=right_lobe_of_liver sigmoid_colon=sigmoid_colon smooth_muscle_cell=smooth_muscle_cell spleen=spleen stomach=stomach subcutaneous_adipose_tissue=subcutaneous_adipose_tissue suprapubic_skin=suprapubic_skin testis=testis thyroid_gland=thyroid_gland tibial_artery=tibial_artery tibial_nerve=tibial_nerve transverse_colon=transverse_colon upper_lobe_of_left_lung=upper_lobe_of_left_lung uterus=uterus vagina=vagina\ subGroup2 factor Factor AFF1=AFF1 AGO1=AGO1 AGO2=AGO2 ARHGAP35=ARHGAP35 ARID1B=ARID1B ARID2=ARID2 ARID3A=ARID3A ARNT=ARNT ASH1L=ASH1L ASH2L=ASH2L ATF2=ATF2 ATF3=ATF3 ATF4=ATF4 ATF7=ATF7 ATM=ATM BACH1=BACH1 BATF=BATF BCL11A=BCL11A BCL3=BCL3 BCOR=BCOR BHLHE40=BHLHE40 BMI1=BMI1 BRCA1=BRCA1 BRD4=BRD4 BRD9=BRD9 C11orf30=C11orf30 CBFA2T2=CBFA2T2 CBFA2T3=CBFA2T3 CBFB=CBFB CBX1=CBX1 CBX2=CBX2 CBX3=CBX3 CBX5=CBX5 CBX8=CBX8 CC2D1A=CC2D1A CCAR2=CCAR2 CDC5L=CDC5L CEBPB=CEBPB CHAMP1=CHAMP1 CHD1=CHD1 CHD4=CHD4 CHD7=CHD7 CLOCK=CLOCK COPS2=COPS2 CREB1=CREB1 CREB3L1=CREB3L1 CREBBP=CREBBP CREM=CREM CTBP1=CTBP1 CTCF=CTCF CUX1=CUX1 DACH1=DACH1 DEAF1=DEAF1 DNMT1=DNMT1 DPF2=DPF2 E2F1=E2F1 E2F4=E2F4 E2F6=E2F6 E2F7=E2F7 E2F8=E2F8 E4F1=E4F1 EBF1=EBF1 EED=EED EGR1=EGR1 EHMT2=EHMT2 ELF1=ELF1 ELF4=ELF4 ELK1=ELK1 EP300=EP300 EP400=EP400 ESR1=ESR1 ESRRA=ESRRA ETS1=ETS1 ETV4=ETV4 ETV6=ETV6 EWSR1=EWSR1 EZH2=EZH2 FIP1L1=FIP1L1 FOS=FOS FOSL1=FOSL1 FOSL2=FOSL2 FOXA1=FOXA1 FOXA2=FOXA2 FOXK2=FOXK2 FOXM1=FOXM1 FOXP1=FOXP1 FUS=FUS GABPA=GABPA GABPB1=GABPB1 GATA1=GATA1 GATA2=GATA2 GATA3=GATA3 GATA4=GATA4 GATAD2A=GATAD2A GATAD2B=GATAD2B GMEB1=GMEB1 HCFC1=HCFC1 HDAC1=HDAC1 HDAC2=HDAC2 HDAC3=HDAC3 HDAC6=HDAC6 HES1=HES1 HMBOX1=HMBOX1 HNF1A=HNF1A HNF4A=HNF4A HNF4G=HNF4G HNRNPH1=HNRNPH1 HNRNPK=HNRNPK HNRNPL=HNRNPL HNRNPLL=HNRNPLL HNRNPUL1=HNRNPUL1 HSF1=HSF1 IKZF1=IKZF1 IKZF2=IKZF2 IRF1=IRF1 IRF2=IRF2 IRF3=IRF3 IRF4=IRF4 IRF5=IRF5 JUN=JUN JUNB=JUNB JUND=JUND KAT2A=KAT2A KAT2B=KAT2B KAT8=KAT8 KDM1A=KDM1A KDM4A=KDM4A KDM4B=KDM4B KDM5A=KDM5A KDM5B=KDM5B KLF16=KLF16 KLF5=KLF5 L3MBTL2=L3MBTL2 LCORL=LCORL LEF1=LEF1 MAFF=MAFF MAFK=MAFK MAX=MAX MBD2=MBD2 MCM2=MCM2 MCM3=MCM3 MCM5=MCM5 MCM7=MCM7 MEF2A=MEF2A MEF2B=MEF2B MEF2C=MEF2C MEIS2=MEIS2 MGA=MGA MIER1=MIER1 MITF=MITF MLLT1=MLLT1 MNT=MNT MTA1=MTA1 MTA2=MTA2 MTA3=MTA3 MXI1=MXI1 MYB=MYB MYBL2=MYBL2 MYC=MYC MYNN=MYNN NANOG=NANOG NBN=NBN NCOA1=NCOA1 NCOA2=NCOA2 NCOA3=NCOA3 NCOA4=NCOA4 NCOA6=NCOA6 NCOR1=NCOR1 NEUROD1=NEUROD1 NFATC1=NFATC1 NFATC3=NFATC3 NFE2=NFE2 NFE2L2=NFE2L2 NFIB=NFIB NFIC=NFIC NFRKB=NFRKB NFXL1=NFXL1 NFYA=NFYA NFYB=NFYB NR0B1=NR0B1 NR2C1=NR2C1 NR2C2=NR2C2 NR2F1=NR2F1 NR2F2=NR2F2 NR2F6=NR2F6 NR3C1=NR3C1 NRF1=NRF1 NUFIP1=NUFIP1 PAX5=PAX5 PAX8=PAX8 PBX3=PBX3 PCBP1=PCBP1 PCBP2=PCBP2 PHB2=PHB2 PHF20=PHF20 PHF21A=PHF21A PHF8=PHF8 PKNOX1=PKNOX1 PLRG1=PLRG1 PML=PML POLR2A=POLR2A POLR2G=POLR2G POU2F2=POU2F2 PRDM10=PRDM10 PRPF4=PRPF4 PTBP1=PTBP1 PYGO2=PYGO2 RAD21=RAD21 RAD51=RAD51 RB1=RB1 RBBP5=RBBP5 RBFOX2=RBFOX2 RBM14=RBM14 RBM15=RBM15 RBM17=RBM17 RBM22=RBM22 RBM25=RBM25 RBM34=RBM34 RBM39=RBM39 RCOR1=RCOR1 RELB=RELB REST=REST RFX1=RFX1 RFX5=RFX5 RLF=RLF RNF2=RNF2 RUNX1=RUNX1 RUNX3=RUNX3 RXRA=RXRA SAFB=SAFB SAFB2=SAFB2 SAP30=SAP30 SETDB1=SETDB1 SIN3A=SIN3A SIN3B=SIN3B SIRT6=SIRT6 SIX4=SIX4 SIX5=SIX5 SKI=SKI SKIL=SKIL SMAD1=SMAD1 SMAD2=SMAD2 SMAD5=SMAD5 SMARCA4=SMARCA4 SMARCA5=SMARCA5 SMARCB1=SMARCB1 SMARCC2=SMARCC2 SMARCE1=SMARCE1 SMC3=SMC3 SNRNP70=SNRNP70 SOX13=SOX13 SOX6=SOX6 SP1=SP1 SPI1=SPI1 SREBF1=SREBF1 SREBF2=SREBF2 SRF=SRF SRSF4=SRSF4 SRSF7=SRSF7 SRSF9=SRSF9 STAT1=STAT1 STAT2=STAT2 STAT3=STAT3 STAT5A=STAT5A SUPT20H=SUPT20H SUZ12=SUZ12 TAF1=TAF1 TAF15=TAF15 TAF7=TAF7 TAF9B=TAF9B TAL1=TAL1 TBL1XR1=TBL1XR1 TBP=TBP TBX21=TBX21 TBX3=TBX3 TCF12=TCF12 TCF7=TCF7 TCF7L2=TCF7L2 TEAD4=TEAD4 TFAP4=TFAP4 THAP1=THAP1 THRA=THRA TRIM22=TRIM22 TRIM24=TRIM24 TRIM28=TRIM28 TRIP13=TRIP13 U2AF1=U2AF1 U2AF2=U2AF2 UBTF=UBTF USF1=USF1 USF2=USF2 WHSC1=WHSC1 WRNIP1=WRNIP1 XRCC3=XRCC3 XRCC5=XRCC5 YY1=YY1 ZBED1=ZBED1 ZBTB1=ZBTB1 ZBTB11=ZBTB11 ZBTB2=ZBTB2 ZBTB33=ZBTB33 ZBTB40=ZBTB40 ZBTB5=ZBTB5 ZBTB7A=ZBTB7A ZBTB7B=ZBTB7B ZBTB8A=ZBTB8A ZEB1=ZEB1 ZEB2=ZEB2 ZFP91=ZFP91 ZFX=ZFX ZHX1=ZHX1 ZHX2=ZHX2 ZKSCAN1=ZKSCAN1 ZMIZ1=ZMIZ1 ZMYM3=ZMYM3 ZNF143=ZNF143 ZNF184=ZNF184 ZNF207=ZNF207 ZNF217=ZNF217 ZNF24=ZNF24 ZNF263=ZNF263 ZNF274=ZNF274 ZNF280A=ZNF280A ZNF282=ZNF282 ZNF316=ZNF316 ZNF318=ZNF318 ZNF384=ZNF384 ZNF407=ZNF407 ZNF444=ZNF444 ZNF507=ZNF507 ZNF512B=ZNF512B ZNF574=ZNF574 ZNF579=ZNF579 ZNF592=ZNF592 ZNF639=ZNF639 ZNF687=ZNF687 ZNF8=ZNF8 ZNF830=ZNF830 ZSCAN29=ZSCAN29 ZZZ3=ZZZ3\ track encTfChipPk\ type narrowPeak\ visibility hide\ netCriGriChoV2 Chinese hamster Net netAlign criGriChoV2 chainCriGriChoV2 Chinese hamster (Jun. 2017 (CHOK1S_HZDv1/criGriChoV2)) Alignment Net 1 2 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Chinese hamster (Jun. 2017 (CHOK1S_HZDv1/criGriChoV2)) Alignment Net\ otherDb criGriChoV2\ parent placentalChainNetViewnet off\ shortLabel Chinese hamster Net\ subGroups view=net species=s004b clade=c00\ track netCriGriChoV2\ type netAlign criGriChoV2 chainCriGriChoV2\ netMonDom5 Opossum Net netAlign monDom5 chainMonDom5 Opossum (Oct. 2006 (Broad/monDom5)) Alignment Net 1 2 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Opossum (Oct. 2006 (Broad/monDom5)) Alignment Net\ otherDb monDom5\ parent vertebrateChainNetViewnet on\ shortLabel Opossum Net\ subGroups view=net species=s003 clade=c00\ track netMonDom5\ type netAlign monDom5 chainMonDom5\ netPanTro6 Chimp Net netAlign panTro6 chainPanTro6 Chimp (Jan. 2018 (Clint_PTRv2/panTro6)) Alignment Net 1 2 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Chimp (Jan. 2018 (Clint_PTRv2/panTro6)) Alignment Net\ otherDb panTro6\ parent primateChainNetViewnet off\ shortLabel Chimp Net\ subGroups view=net species=s0025 clade=c00\ track netPanTro6\ type netAlign panTro6 chainPanTro6\ lrSv1kgOnt 1KG ONT 1019 SVs bigBed 9 + Structural Variants from 1,019 Diverse Humans (Vienna ONT, Schloissnig et al. 2025) 0 2 0 0 0 127 127 127 0 0 0\ This track shows structural variants (SVs) identified by Oxford Nanopore long-read\ sequencing of 1,019 individuals from the 1000 Genomes Project, representing 26\ populations across 5 continental regions: Africa (275 samples), East Asia (192),\ South Asia (199), Europe (189), and Americas (164). Median sequencing coverage\ was 16.9x per sample with a median N50 read length of 20.3 kb.\
\\ SVs were discovered using the SAGA framework (SV Analysis by Graph Augmentation)\ and annotated with SVAN, which classifies insertions and deletions by their\ mechanism of origin. The full release is native to the T2T-CHM13 assembly\ (hs1) and contains 161,332 annotated SVs (75,324 insertions, 66,192 deletions,\ and 19,816 complex rearrangements). For GRCh38 (hg38), coordinates were converted\ using liftOver and 148,375 records mapped successfully (73,298 insertions,\ 58,637 deletions, and 16,440 complex rearrangements).\
\\ The 1,019 samples sequenced here are distinct from those in the\ 1KG ONT 100 track (Gustafson et al. 2024);\ the two releases were produced by separate consortia (Vienna and the 1000 Genomes\ ONT Sequencing Consortium, respectively) and there is no sample overlap between\ the two.\
\ \\ Items are colored by SV class:\
\ Filters are available for SV type, insertion/deletion type, transposon family,\ and SV length. For insertions, the item is placed at the insertion site with a\ width of 1 bp; for deletions, the item spans the deleted region.\
\\ The detail page for each item shows SVAN annotation fields including:\
\ Schloissnig et al. 2025 generated intermediate-coverage Oxford Nanopore\ long-read sequencing of 1,019 samples from the 1000 Genomes Project on\ PromethION 48 instruments with R9.4.1 (FLO-PRO002) flow cells (SQK-LSK110\ libraries, 24-h runs with flow-cell wash and reload). SVs were discovered\ with the SAGA framework (SV Analysis by Graph Augmentation), which combines\ linear-reference callers (Sniffles and DELLY, run against both GRCh38 and\ T2T-CHM13), graph-aware discovery with SVarp (local long-read assembly of\ SV-supporting graph-aligned reads) and graph-based joint genotyping with\ Giggles across a pangenome graph. Insertions and deletions were then\ annotated with \ SVAN v1.3, which classifies SVs by mechanism of origin. The release\ contains 161,332 SVAN-annotated SVs: 75,324 insertions, 66,192 deletions\ and 19,816 complex rearrangements. The original VCF is on T2T-CHM13 contig\ coordinates; for the hg38 version of this track, SVs were lifted with\ liftOver (148,375 of 161,332 records mapped), while the hs1 version uses\ the native coordinates.\
\\ The SVAN-annotated unphased VCF (final-vcf.unphased.SVAN_1.3.vcf.gz)\ was downloaded from\ \ the IGSR 1KG_ONT_VIENNA v1.1 SVAN-annotation directory; allele counts\ were added from the companion shapeit5-phased-callset\ (shapeit5-phased-callset_final-vcf.phased.vcf.gz) in the same\ release tree.\
\\ The step-by-step build commands (download, liftOver, format conversion,\ bigBed build) are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ Source data is available from the\ 1000 Genomes ONT Vienna data collection at IGSR.\
\ \\ Thanks to the 1000 Genomes ONT Vienna consortium for making their structural\ variant calls and SVAN annotations publicly available.\
\ \\ Schloissnig S, Pani S, Ebler J, Hain C, Tsapalou V, Söylev A, Hüther P, Ashraf H, Prodanov T,\ Asparuhova M et al.\ \ Structural variation in 1,019 diverse humans based on long-read sequencing.\ Nature. 2025 Aug;644(8076):442-452.\ PMID: 40702182; PMC: PMC12350158\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/1kgOnt.bb\ dataVersion 1.1\ filter.AC 0:1816\ filter.insLen 0:48091\ filter.svLen 0:49171\ filterByRange.AC on\ filterByRange.alleleFreq on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Allele Count\ filterLabel.alleleFreq Allele Frequency\ filterLabel.family Transposon Family\ filterLabel.insLen Insertion Length\ filterLabel.insType Insertion/Deletion Type\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.alleleFreq 0:1\ filterType.family multipleListOr\ filterType.insType multipleListOr\ filterType.svType multipleListOr\ filterValues.family Alu,HERVK,L1,LTR5_Hs,SVA\ filterValues.insType COMPLEX_DUP,DUP,DUP_INTERSPERSED,INV_DUP,NUMT,PSD,VNTR,chimera,orphan,partnered,solo\ filterValues.svType DEL,INS,CPX\ itemRgb on\ longLabel Structural Variants from 1,019 Diverse Humans (Vienna ONT, Schloissnig et al. 2025)\ mouseOver Var: $name ($svType)\ The All of Us Research Program is a\ large-scale biomedical research initiative launched by the U.S. National Institutes of Health (NIH)\ in 2018. Its goal is to build one of the most diverse health databases, enrolling over one\ million participants who reflect the full diversity of the United States, including groups that\ have been historically underrepresented in biomedical research. Participants contribute health\ surveys, electronic health records (EHR), physical measurements, and biosamples for genomic\ analysis.\
\ \\ This track shows allele frequencies from the v7 short-read whole-genome sequencing (srWGS)\ release of 245,388 participants. A minimum allele count filter of ≥20 was applied.\ Frequencies are provided both overall and broken down by genetic ancestry using local ancestry\ inference: European (EUR), East Asian (EAS), African (AFR), Indigenous American (AMR),\ Oceanian (OCE), and South Asian (SAS). Some variants are flagged with an "NW" tag\ (not in window) when the variant was not within a genomic window covered by the ancestry\ reference files; in these cases the closest available position was used for ancestry assignment.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ Variant data and individual-level data are accessible through the\ All of Us Researcher Workbench,\ which requires registration and completion of a training program. Aggregate allele frequency\ data is freely available.\
\ \\ Whole-genome sequencing was performed on the Illumina NovaSeq 6000 platform with PCR-free library\ preparation targeting 30x coverage. Reads were aligned to GRCh38 and variants were called using\ the Illumina DRAGEN (Dynamic Read Analysis for GENomics) pipeline, which performs mapping,\ alignment, sorting, duplicate marking, and variant calling (SNVs and indels) in a single\ hardware-accelerated workflow. Joint genotyping was performed across all samples. Quality control\ included sample-level filtering for contamination, sex discordance, and relatedness, and\ variant-level filtering using VQSR.\ Population-specific allele frequencies were determined using local ancestry inference at UCSC by the Ioannidis group.\ The ancestry breakdown into European, East Asian, African, Indigenous American, Oceanian,\ and South Asian components is part of a pending publication.\
\\ The conversion of all source files for the varFreqs track is documented in the makeDoc file of the track.\ For some tracks, python scripts were needed and are also available from GitHub.\
\ \\ The All of Us Research Program is supported by the National Institutes of Health. We thank the\ participants and the program for making frequency data available.\ The local ancestry inference was performed by Qudsi Aljabiri and Cole Shanks under\ Prof. Alexander Ioannidis, UC Santa Cruz.\
\ \\ All of Us Research Program Genomics Investigators.\ \ Genomic data in the All of Us Research Program.\ Nature. 2024 Mar;627(8003):340-346.\ PMID: 38374255; PMC: PMC10937371\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_allofus/allOfUs.locAncFreq.vcf.gz\ dataVersion V7\ longLabel SNV Frequencies: AllOfUs v7 - 245k WGS, local-ancestry-stratified, AC>=20\ parent varFreqs on\ priority 2\ shortLabel AllOfUs v7 245k WGS\ tableBrowser off\ track allofus\ type vcfTabix\ visibility hide\ AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_ctss_rev AorticSmsToFgf2_00hr00minBr1- bigWig Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_reverse 0 2 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr00min%2c%20biol_rep1%20%28LK1%29.CNhs13339.12642-134G5.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12642-134G5 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr00minBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_tpm_rev AorticSmsToFgf2_00hr00minBr1- bigWig Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_reverse 1 2 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr00min%2c%20biol_rep1%20%28LK1%29.CNhs13339.12642-134G5.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr00min, biol_rep1 (LK1)_CNhs13339_12642-134G5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12642-134G5 sequence_tech=hCAGE\ parent TSS_activity_TPM on\ shortLabel AorticSmsToFgf2_00hr00minBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr00minBiolRep1LK1_CNhs13339_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12642-134G5\ urlLabel FANTOM5 Details:\ cons241wayViewphyloP Basewise Conservation (phyloP) bed 4 Zoonomia Alignment - 241 Placental Mammal Genomes aligned by the Zoonomia Project with Cactus 2 2 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel Zoonomia Alignment - 241 Placental Mammal Genomes aligned by the Zoonomia Project with Cactus\ parent cons241way\ shortLabel Basewise Conservation (phyloP)\ track cons241wayViewphyloP\ view phyloP\ viewLimits -20.0:9.869\ viewLimitsMax -20:0.869\ visibility full\ iscaBenign Benign gvf ClinGen CNVs: Benign 3 2 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/?term=$$ phenDis 1 longLabel ClinGen CNVs: Benign\ parent iscaViewDetail off\ shortLabel Benign\ subGroups view=cnv class=ben level=sub\ track iscaBenign\ bismap36Pos Bismap S36 + bigBed 6 Single-read mappability with 36-mers after bisulfite conversion (forward strand) 0 2 240 70 80 247 162 167 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k36.C2T-Converted.bb\ color 240,70,80\ longLabel Single-read mappability with 36-mers after bisulfite conversion (forward strand)\ parent bismapBigBed off\ priority 2\ shortLabel Bismap S36 +\ subGroups view=SR\ track bismap36Pos\ visibility hide\ wgEncodeReg4DnaseBone Bone bigWig DNase level of 1 bone experiment (tissues and primary cells only) 0 2 121 147 150 188 201 202 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBoneDNase.bw\ color 121,147,150\ longLabel DNase level of 1 bone experiment (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 2\ shortLabel Bone\ track wgEncodeReg4DnaseBone\ type bigWig\ wgEncodeReg4MarkH3k27acBoneMarrow Bone marrow bigWig Avg. H3K27ac level of 9 bone marrow experiments (tissues and primary cells only) 2 2 184 120 120 219 187 187 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBoneMarrowH3K27ac.bw\ color 184,120,120\ longLabel Avg. H3K27ac level of 9 bone marrow experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 2\ shortLabel Bone marrow\ track wgEncodeReg4MarkH3k27acBoneMarrow\ type bigWig\ wgEncodeReg4MarkH3k4me3BoneMarrow Bone marrow bigWig Avg. H3K4me3 level of 13 bone marrow experiments (tissues and primary cells only) 0 2 184 120 120 219 187 187 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBoneMarrowH3K4me3.bw\ color 184,120,120\ longLabel Avg. H3K4me3 level of 13 bone marrow experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 2\ shortLabel Bone marrow\ track wgEncodeReg4MarkH3k4me3BoneMarrow\ type bigWig\ wgEncodeReg4MarkCtcfBrain Brain bigWig Avg. CTCF level of 54 brain experiments (tissues and primary cells only) 0 2 155 155 18 205 205 136 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBrainCTCF.bw\ color 155,155,18\ longLabel Avg. CTCF level of 54 brain experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf\ priority 2\ shortLabel Brain\ track wgEncodeReg4MarkCtcfBrain\ type bigWig\ wgEncodeReg4AtacBrain Brain bigWig Avg. ATAC level of 2 brain experiments (tissues and primary cells only) 0 2 155 155 18 205 205 136 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBrainATAC.bw\ color 155,155,18\ longLabel Avg. ATAC level of 2 brain experiments (tissues and primary cells only)\ parent wgEncodeReg4Atac\ priority 2\ shortLabel Brain\ track wgEncodeReg4AtacBrain\ type bigWig\ cons241wayViewalign Cactus Alignments bed 4 Zoonomia Alignment - 241 Placental Mammal Genomes aligned by the Zoonomia Project with Cactus 3 2 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel Zoonomia Alignment - 241 Placental Mammal Genomes aligned by the Zoonomia Project with Cactus\ parent cons241way\ shortLabel Cactus Alignments\ track cons241wayViewalign\ view align\ viewUi on\ visibility pack\ cerebNeuron0TB Cerebellum - Neuron - Z000000TB bigWig Methylation Atlas: Cerebellum - Neuron - Z000000TB 2 2 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/cerebNeuron0TB.bw\ color 138,43,226\ longLabel Methylation Atlas: Cerebellum - Neuron - Z000000TB\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 2\ shortLabel Cerebellum - Neuron - Z000000TB\ subGroups cellType=Neuron dataType=Replicate\ track cerebNeuron0TB\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ chineseTrio Chinese Trio vcfPhasedTrio Genome In a Bottle Chinese Trio 0 2 0 0 0 127 127 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/giab/ChineseTrio/merged.vcf.gz\ longLabel Genome In a Bottle Chinese Trio\ maxWindowToDraw 5000000\ parent triosView\ shortLabel Chinese Trio\ subGroups view=trios\ track chineseTrio\ type vcfPhasedTrio\ vcfChildSample HG005|son\ vcfDoFilter off\ vcfDoMaf off\ vcfDoQual off\ vcfParentSamples HG006|father,HG007|mother\ vcfUseAltSampleNames on\ clinGenTriplo ClinGen Triplosensitivity bigBed 9 + ClinGen Dosage Sensitivity Map - Triplosensitivity 3 2 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/bbi/clinGen/clinGenTriplo.bb\ dataVersion /gbdb/$D/bbi/clinGen/clinGenDosageVersion.txt\ filterLabel.triploScore Dosage Sensitivity Score\ filterValues.triploScore 0|No evidence available,1|Little evidence for dosage pathogenicity,2|Some evidence for dosage pathogenicity,3|Sufficient evidence for dosage pathogenicity,30|Gene associated with autosomal recessive phenotype,40|Dosage sensitivity unlikely\ longLabel ClinGen Dosage Sensitivity Map - Triplosensitivity\ mouseOver Gene/ISCA ID: $name\ This track set shows the results of the\ GWAS Data Release 4 (October 2020) \ from the \ \ COVID-19 Host Genetics Initiative (HGI): \ a collaborative effort to facilitate \ the generation of meta-analysis across multiple studies contributed by\ partners world-wide\ to identify the genetic determinants of SARS-CoV-2 infection susceptibility, disease severity \ and outcomes. The COVID-19 HGI also aims to provide a platform for study partners to \ share analytical results in the form of summary statistics and/or individual level data of COVID-19\ host genetics research. At the time of this release, a total of 137 studies were registered with \ this effort.\
\ \\ The specific phenotypes studied by the COVID-19 HGI are those that benefit from maximal sample \ size: primary analysis on disease severity. For the Data Release 4 the number of cases have\ increased by nearly ten-fold (more than 30,000 COVID-19 cases and 1.47 million controls) by combining\ data from 34 studies across 16 countries. \
\ \\ The four tracks here are based on data from HGI meta-analyses A2, B2, C1, and C2, described here:\
\ \| SNP | \Human GRCh37/hg19 Assembly | \Human GRCh38/hg38 Assembly | \Risk Allele | \Alternative | \Gene nearest to SNP | \
|---|---|---|---|---|---|
| rs73064425 | \chr3:45901089-45901089 | \chr3:45859597-45859597 | \T | \C | \LZTFL1 | \
| rs9380142 | \chr6:29798794-29798794 | \chr6:29831017-29831017 | \A | \G | \HLA-G | \
| rs143334143 | \chr6:31121426-31121426 | \chr6:31153649-31153649 | \A | \G | \CCHCR1 | \
| rs10735079 | \chr12:113380008-113380008 | \chr12:112942203-112942203 | \A | \G | \OAS3 | \
| rs74956615 | \chr19:10427721-10427721 | \chr19:10317045-10317045 | \A | \T | \ICAM5/TYK2 | \
| rs2109069 | \chr19:4719443-4719443 | \chr19:4719431-4719431 | \A | \G | \DPP9 | \
| rs2236757 | \chr21:34624917-34624917 | \chr21:33252612-33252612 | \A | \G | \IFNAR2 | \
\
\
\
\ Displayed items are colored by GWAS effect: red for positive (harmful) effect, \ blue for negative (protective) effect.\ The height ('lollipop stem') of the item is based on statistical significance (p-value). \ For better visualization of the data, only SNPs with p-values smaller than 1e-3 are \ displayed by default.
\\ The color saturation indicates effect size (beta coefficient): values over the median of effect \ size are brightly colored (bright red\ \ , bright blue\ \ ),\ those below the median are paler (light red\ \ , light blue\ \ ). \
\\ Each track has separate display controls and data can be filtered according to the\ number of studies, minimum -log10 p-value, and the\ effect size (beta coefficient), using the track Configure options.
\\ Mouseover on items shows the rs ID (or chrom:pos if none assigned), both the non-effect \ and effect alleles, the effect size (beta coefficient), the p-value, and the number of \ studies.\ Additional information on each variant can be found on the details page by clicking on \ the item.
\ \\ COVID-19 Host Genetics Initiative (HGI) GWAS meta-analysis round 4 (October 2020) results were \ used in this study. \ Each participating study partner submitted GWAS summary statistics for up to four \ of the COVID-19 phenotype definitions.
\\ Data were generated from genome-wide SNP array and whole exome and genome\ sequencing, leveraging the impact of both common and rare variants. The statistical analysis\ performed takes into account differences between sex, ancestry, and date of sample collection. \ Alleles were harmonized across studies and reported allele frequencies are based on gnomAD \ version 3.0 reference data. Most study partners used the SAIGE GWAS pipeline in order \ to generate summary statistics used for the COVID-19 HGI meta-analysis. The summary statistics \ of individual studies were manually examined for inflation, \ deflation, and excessive number of false positives. \ Qualifying summary statistics were filtered for \ INFO > 0.6 and MAF > 0.0001 prior to meta-analyzing the entirety of the data. \
\ The meta-analysis was performed using fixed effects inverse variance weighting.\ The meta-analysis software and workflow are available here. More information about the \ prospective studies, processing pipeline, results and data sharing can be found \ here.\ \ \\ The data underlying these tracks and summary statistics results are publicly available in COVID19-hg Release 4 (October 2020).\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. \ Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.\
\ \\ Thanks to the COVID-19 Host Genetics Initiative contributors and project leads for making these \ data available, and in particular to Rachel Liao, Juha Karjalainen, and Kumar Veerapen at the \ Broad Institute for their review and input during browser track development.\
\ \\ COVID-19 Host Genetics Initiative.\ \ The COVID-19 Host Genetics Initiative, a global initiative to elucidate the role of host genetic\ factors in susceptibility and severity of the SARS-CoV-2 virus pandemic.\ Eur J Hum Genet. 2020 Jun;28(6):715-718.\ PMID: 32404885; PMC: PMC7220587\
\ \\ Pairo-Castineira E, Clohisey S, Klaric L, Bretherick AD, Rawlik K, Pasko D, Walker S, Parkinson N,\ Fourman MH, Russell CD et al.\ \ Genetic mechanisms of critical illness in Covid-19.\ Nature. 2020 Dec 11;.\ PMID: 33307546\
\ \ \ phenDis 1 autoScale on\ bedNameLabel SNP\ chromosomes chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22\ compositeTrack on\ filter._effectSizeAbs 0\ filter.effectSize -1.6:2.2\ filter.pValueLog 3\ filter.sourceCount 1\ filterByRange.effectSize on\ filterLabel._effectSizeAbs Minimum effect size +-\ filterLabel.effectSize Effect size range\ filterLabel.sourceCount Minimum number of studies\ filterLimits.effectSize -1.6:2.2\ lollyField 13\ longLabel COVID risk variants from GWAS meta-analyses by the COVID-19 Host Genetics Initiative (Rel 4, Oct 2020)\ maxHeightPixels 48:75:128\ maxItems 500000\ mouseOver $name $ref/$alt effect $effectSize pVal $pValue studies $sourceCount\ noScoreFilter on\ priority 2\ shortLabel COVID GWAS v4\ superTrack covid pack\ track covidHgiGwasR4Pval\ type bigLolly 9 +\ viewLimits 0:10\ cq7Vcf CQ-7 Variants vcfTabix CQ-7 Variants 0 2 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/highRepro/CQ-7.sort.vcf.gz\ longLabel CQ-7 Variants\ parent highReproVcfs\ shortLabel CQ-7 Variants\ subGroups view=vcfs\ track cq7Vcf\ type vcfTabix\ crossTissueMapsFullDetails Cross Tissue Details bigBarChart Cross tissue nuclei full details 0 2 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$\
This track collection shows data from \
Single-nucleus cross-tissue molecular reference maps toward\
understanding disease gene function. The dataset covers ~200,000 single nuclei\
from a total of 16 human donors across 25 samples, using 4 different sample preparation\
protocols followed by droplet based single-cell RNA-seq. The samples were obtained from\
frozen tissue as part of the Genotype-Tissue Expression (GTEx) project.\
Samples were taken from the esophagus, skeletal muscle, heart, lung, prostate, breast,\
and skin. The dataset includes 43 broad cell classes, some specific to certain tissues\
and some shared across all tissue types.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ This track collection contains three bar chart tracks of RNA expression. The first track,\ Cross Tissue Nuclei, allows\ cells to be grouped together and faceted on up to 4 categories: tissue, cell class, cell subclass,\ and cell type. The second track,\ Cross Tissue Details, allows\ cells to be grouped together and faceted on up to 7 categories: tissue, cell class, cell subclass,\ cell type, granular cell type, sex, and donor. The third track,\ GTEx Immune Atlas,\ allows cells to be grouped together and faceted on up to 5 categories: tissue, cell type, cell\ class, sex, and donor.\
\ \\ Please see the\ GTEx portal\ for further interactive displays and additional data.
\ \\ Tissue-cell type combinations in the Full and Combined tracks are\ colored by which cell type they belong to in the below table:\
\
| Color | \Cell Type | \
|---|---|
| Endothelial | |
| Epithelial | |
| Glia | |
| Immune | |
| Neuron | |
| Stromal | |
| Other |
\ Tissue-cell type combinations in the Immune Atlas track are shaded according\ to the below table:\
| Color | \Cell Type | \
|---|---|
| Inflammatory Macrophage | |
| Lung Macrophage | |
| Monocyte/Macrophage FCGR3A High | |
| Monocyte/Macrophage FCGR3A Low | |
| Macrophage HLAII High | |
| Macrophage LYVE1 High | |
| Proliferating Macrophage | |
| Dendritic Cell 1 | |
| Dendritic Cell 2 | |
| Mature Dendritic Cell | |
| Langerhans | |
| CD14+ Monocyte | |
| CD16+ Monocyte | |
| LAM-like | |
| Other |
\ Using the previously collected tissue samples from the Genotype-Tissue Expression\ project, nuclei were isolated using four different protocols and sequenced\ using droplet based single cell RNA-seq. CellBender v2.1 and other standard quality\ control techniques were applied, resulting in 209,126 nuclei profiles across eight\ tissues, with a mean of 918 genes and 1519 transcripts per profile.\
\ \\ Data from all samples was integrated with a conditional variation autoencoder\ in order to correct for multiple sources of variation like sex, and protocol\ while preserving tissue and cell type specific effects.\
\ \\ For detailed methods, please refer to Eraslan et al, or the\ \ GTEx portal website.\
\ \\
The gene expression files were downloaded from the\
\
GTEx portal. The UCSC command line utilities matrixClusterColumns,\
matrixToBarChartBed, and bedToBigBed were used to transform\
these into a bar chart format bigBed file that can be visualized.\
The UCSC utilities can be found on\
our download server.\
\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions or our Data Access FAQ for more\ information.
\ \\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the GTEx Consortium for creating and analyzing these data.
\ \\ Eraslan G, Drokhlyansky E, Anand S, Fiskin E, Subramanian A, Slyper M, Wang J, Van Wittenberghe N,\ Rouhana JM, Waldman J et al.\ \ Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function.\ Science. 2022 May 13;376(6594):eabl4290.\ PMID: 35549429; PMC: PMC9383269\
\ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/crossTissueMaps/facet_detailed.categories\ barChartFacets tissue,cell_class,cell_subclass,cell_type,granular_cell_type,sex,donor\ barChartMerge on\ barChartMetric gene/genome\ barChartStatsUrl /gbdb/hg38/bbi/crossTissueMaps/facet_detailed.facets\ barChartStretchToItem on\ barChartUnit parts per million\ bigDataUrl /gbdb/hg38/bbi/crossTissueMaps/facet_detailed.bb\ defaultLabelFields name\ html crossTissueMaps\ labelFields name,name2\ longLabel Cross tissue nuclei full details\ maxWindowToDraw 10000000\ parent crossTissueMaps\ priority 2\ shortLabel Cross Tissue Details\ track crossTissueMapsFullDetails\ type bigBarChart\ url https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$NOTE:
\
While the DECIPHER database is \
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions.\
Because the UCSC Genes mappings for CNVs are based on associations from\ RefSeq and UniProt, they are dependent on any interpretations from those\ sources. Furthermore, because many DECIPHER records refer to multiple gene\ names, or syndromes not tightly mapped to individual genes, the associations\ in this track should be treated with skepticism and any conclusions\ based on them should be carefully scrutinized using independent\ resources.\
\Data Display Agreement Notice
\
The CNV/SNV data are only available for display in the Browser, and not for bulk\
download. Access to bulk data may be obtained directly from DECIPHER\
(https://www.deciphergenomics.org/about/data-sharing) and is subject to a\
Data Access Agreement, in which the user certifies that no attempt to\
identify individual patients will be undertaken. The same restrictions\
apply to the public data displayed at UCSC in the UCSC Genome Browser;\
no one is authorized to attempt to identify patients by any means.\
These data are made available as soon as possible and may be a\ pre-publication release. For information on the proper use of DECIPHER\ data, please see https://www.deciphergenomics.org/about/data-sharing.\
\The DECIPHER consortium provides these data in good faith as a research\ tool, but without verifying the accuracy, clinical validity, or utility of\ the data. The DECIPHER consortium makes no warranty, express or implied,\ nor assumes any legal liability or responsibility for any purpose for\ which the data are used.\
\\ The \ DECIPHER\ database of submicroscopic chromosomal imbalance \ collects clinical information about chromosomal \ microdeletions/duplications/insertions, translocations and inversions, \ and displays this information on the human genome map.\
\ The CNVs and SNVs tracks show genomic regions of reported cases and their \ associated phenotype information. All data have passed the strict\ consent requirements of the DECIPHER project and are approved for\ unrestricted public release. Clicking the Patient View ID link\ brings up a more detailed informational page on the patient at the \ DECIPHER web site.
\ \\ The Population CNVs track shows common copy-number variants (CNVs) and their\ population frequencies, lifted over from the hg19 assembly.
\ \\ The genomic locations of DECIPHER variants are labeled with the DECIPHER variant descriptions. \ Mouseover on items shows variant details, clinical interpretation, and associated conditions. \ Further information on each variant is displayed on the details page by a click onto any variant. \
\ \\ For the CNVs track, the entries are colored by the type of variant:\
\ A light-to-dark color gradient indicates the clinical significance of each variant, with \ the lightest shade being benign, to the darkest shade being pathogenic. Detailed information on the \ CNV color code is described here.\ Items can be filtered according to the size of the variant, variant type, and clinical significance \ using the track Configure options.\
\ \\ For the SNVs track, the entries are colored according to the estimated clinical significance \ of the variant:\
\ For the Population CNVs track, genomic variants are visually differentiated to facilitate quick and\ clear identification. Variants are colored according to their clinical significance and type:\
\\ The Population CNVs track's mouseover tooltip provides the following information\ about the data:\
\\ Data provided by the DECIPHER project group are imported and processed\ to create a simple BED track to annotate the genomic regions associated\ with individual patients.\
\ \ \\ For more information on DECIPHER, please contact\ \ contact@deciphergenomics.\ org\
\ \\ The DECIPHER data access and documentation can be found at\ DECIPHER Downloads.\
\ \\ Firth HV, Richards SM, Bevan AP, Clayton S, Corpas M, Rajan D, Van Vooren S, Moreau Y, Pettett RM,\ Carter NP.\ \ DECIPHER: Database of Chromosomal Imbalance and Phenotype in Humans Using Ensembl Resources.\ Am J Hum Genet. 2009 Apr;84(4):524-33.\ PMID: 19344873; PMC: PMC2667985\
\ phenDis 1 color 0,0,0\ group phenDis\ html decipherContainer\ longLabel DECIPHER: Chromosomal Imbalance and Phenotype in Humans (SNVs)\ nextExonText Right edge\ parent decipherContainer\ prevExonText Left edge\ priority 2\ shortLabel DECIPHER SNVs\ tableBrowser off decipherSnvsRaw\ track decipherSnvs\ type bed 4\ visibility pack\ dgvSupporting DGV Supp Var bigBed 9 + Database of Genomic Variants: Supporting Structural Var (CNV, Inversion, In/del) 0 2 0 0 0 127 127 127 0 0 0 http://dgv.tcag.ca/dgv/app/variant?id=$$&ref=$D varRep 1 bigDataUrl /gbdb/hg38/dgv/dgvSupporting.bb\ dataVersion 2020-02-25\ filter._size 1:9320633\ filterByRange._size on\ filterLabel._size Genomic size of variant\ filterValues.varType complex,deletion,duplication,gain,gain+loss,insertion,inversion,loss,mobile element insertion,novel sequence insertion,sequence alteration,tandem duplication\ longLabel Database of Genomic Variants: Supporting Structural Var (CNV, Inversion, In/del)\ mouseOver ID: $nameThis track displays genome-wide epigenomic signals and peaks from 3,201 individual\ ENCODE experiments, including DNase-seq and ATAC-seq for chromatin accessibility, and\ ChIP-seq for the histone modifications H3K4me3 and H3K27ac, as well as CTCF binding.
\ \The track includes two subtrack types:
\These datasets provide the underlying experimental data used to generate the\ corresponding layered summary tracks. Additional datasets are available at the\ ENCODE portal.
\ \Click a specific biosample type and organ/tissue combination to view available datasets.\ Subtracks can be further filtered by Assay (ATAC, DNase, CTCF, H3K27ac, and H3K4me3),\ Organ, Biosample Type, Data Type (Signal or Peak), and Life Stage.
\ \| Organ/Tissue | \DNase | \ATAC | \H3K4me3 | \H3K27ac | \CTCF | \
|---|---|---|---|---|---|
| adipose | ✓ | ✓ | ✓ | ✓ | ✓ |
| adrenal gland | ✓ | ✓ | ✓ | ✓ | ✓ |
| blood | ✓ | ✓ | ✓ | ✓ | ✓ |
| blood vessel | ✓ | ✓ | ✓ | ✓ | ✓ |
| bone | ✓ | – | ✓ | ✓ | ✓ |
| bone marrow | ✓ | ✓ | ✓ | ✓ | ✓ |
| brain | ✓ | ✓ | ✓ | ✓ | ✓ |
| breast | ✓ | ✓ | ✓ | ✓ | ✓ |
| connective tissue | ✓ | – | ✓ | ✓ | ✓ |
| embryo | ✓ | – | ✓ | ✓ | ✓ |
| epithelium | ✓ | – | ✓ | ✓ | ✓ |
| esophagus | ✓ | ✓ | ✓ | ✓ | ✓ |
| eye | ✓ | – | ✓ | ✓ | ✓ |
| gallbladder | ✓ | ✓ | – | – | – |
| heart | ✓ | ✓ | ✓ | ✓ | ✓ |
| kidney | ✓ | – | ✓ | ✓ | ✓ |
| large intestine | ✓ | ✓ | ✓ | ✓ | ✓ |
| limb | ✓ | – | – | – | – |
| liver | ✓ | ✓ | ✓ | ✓ | ✓ |
| lung | ✓ | ✓ | ✓ | ✓ | ✓ |
| lymphoid tissue | ✓ | – | – | – | – |
| mouth | ✓ | – | ✓ | ✓ | ✓ |
| muscle | ✓ | ✓ | ✓ | ✓ | ✓ |
| nerve | ✓ | ✓ | ✓ | ✓ | ✓ |
| nose | ✓ | – | – | – | – |
| ovary | ✓ | ✓ | ✓ | ✓ | ✓ |
| pancreas | ✓ | ✓ | ✓ | ✓ | ✓ |
| parathyroid gland | – | – | ✓ | ✓ | ✓ |
| penis | ✓ | – | ✓ | ✓ | ✓ |
| placenta | ✓ | – | ✓ | ✓ | ✓ |
| prostate | ✓ | ✓ | ✓ | ✓ | ✓ |
| skin | ✓ | ✓ | ✓ | ✓ | ✓ |
| small intestine | ✓ | ✓ | ✓ | ✓ | ✓ |
| spinal cord | ✓ | – | ✓ | ✓ | ✓ |
| spleen | ✓ | ✓ | ✓ | ✓ | ✓ |
| stomach | ✓ | ✓ | ✓ | ✓ | ✓ |
| testis | ✓ | ✓ | ✓ | ✓ | ✓ |
| thymus | ✓ | – | ✓ | ✓ | – |
| thyroid | ✓ | ✓ | ✓ | ✓ | ✓ |
| urinary bladder | ✓ | ✓ | ✓ | ✓ | – |
| uterus | ✓ | ✓ | ✓ | ✓ | ✓ |
| vagina | ✓ | – | ✓ | ✓ | ✓ |
\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.\ Clicking any accession in the track's configuration table links directly to the\ corresponding file details page on the ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://encode-public.s3.amazonaws.com/2021/02/25/f34812d4-08cd-4abb-956f-b722b516dcc6/ENCFF094EYJ.bigWig stdout
Data were generated by the ENCODE Consortium through the following production labs:\ Drs. Barbara Wold (Caltech), Bing Ren (UCSD), Bradley Bernstein (Broad), Gregory Crawford (Duke), John Stamatoyannopoulos (UW), Joseph Costello (UCSF), Michael Snyder (Stanford), Peggy Farnham (USC), Richard Myers (HAIB), Stephen Montgomery (Stanford), Vishwanath Iyer (UTA), Will Greenleaf (Stanford), and Yin Shen (UCSF).
\ \The data were further processed for visualization through a collaborative effort between\ the Weng lab and the\ Moore lab at UMass\ Chan Medical School (funded by NIH grant HG012343). Integration and visualization were\ developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\ \\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 1 colorSettingsUrl /gbdb/hg38/encode4/regulation/epi_colors.json\ compositeTrack faceted\ defaultSortField _Experiment\ html wgEncodeReg4Epigenetics.html\ longLabel Peaks and signal from individual DNase, ATAC, histone, and CTCF experiments from ENCODE 4\ maxCheckboxes 50\ metaDataUrl /gbdb/hg38/encode4/regulation/wgEncodeReg4Epigenetics_metadata.tsv\ noInherit on\ primaryKey Accession\ priority 2.0\ shortLabel DNase/ATAC/Histone/CTCF (Indiv.)\ subtrackUrls Accession=https://www.encodeproject.org/files/$$/ Experiment=https://www.encodeproject.org/experiments/$$/\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4Epigenetics\ type bed 3\ visibility hide\ ENCFF316SZE_ENCFF263CSV_ENCFF144JOJ_ENCFF035TJC ENCFF316SZE_ENCFF263CSV_ENCFF144JOJ_ENCFF035TJC bigBed 9 + 5 Adrenal gland, male adult (54 years): (1) cCREs 4 2 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF316SZE_ENCFF263CSV_ENCFF144JOJ_ENCFF035TJC.bb\ longLabel Adrenal gland, male adult (54 years): (1) cCREs\ mouseOver ID: ${name}\ The GENCODE Genes track (version 46, May 2024) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The v46 release was derived from the GTF file that contains annotations only on the main\ chromosomes. Statistics for this build and information on how they were generated can be found on\ the GENCODE site.
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is mostly based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \ \\ Within a gene using the pack display mode, transcripts below a specified rank will be\ condensed into a view similar to squish mode. The transcript ranking approach is\ preliminary and will change in future releases. The transcripts rankings are defined by the\ following criteria for protein-coding and non-coding genes:
\ Protein_coding genes\\
The GENCODE v46 track was built from the GENCODE downloads file \
gencode.v46.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources\
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney. This version of the track was\ generated by Jonathan Casper.
\ \\ Frankish A, Carbonell-Sala S, Diekhans M, Jungreis I, Loveland JE, Mudge JM, Sisu C, Wright JC,\ Arnan C, Barnes I et al.\ \ GENCODE: reference annotation for the human and mouse genomes in 2023.\ Nucleic Acids Res. 2023 Jan 6;51(D1):D942-D949.\ PMID: 36420896; PMC: PMC9825462\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV46.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV46\ group genes\ html knownGeneV46\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V46\ maxItems 50000\ parent knownGeneArchive\ priority 2\ searchIndex name\ shortLabel GENCODE V46\ squishyPackField rank\ squishyPackLabel Number of transcripts shown at full height (ranked by GENCODE transcript ranking)\ squishyPackPoint 1\ track knownGeneV46\ type bigGenePred knownGenePep knownGeneMrna\ visibility pack\ missenseByGene Gene Missense bigBed 12 + gnomAD Predicted Missense Constraint Metrics By Gene (Z-scores) v2.1.1 3 2 0 0 0 127 127 127 0 0 0 https://gnomad.broadinstitute.org/gene/$$?dataset=gnomad_r2_1 varRep 1 bigDataUrl /gbdb/hg38/gnomAD/pLI/missenseByGene.bb\ filter._zscore -19:11\ filterByRange._zscore on\ filterLabel._zscore Show only items between this Z-score range\ itemRgb on\ labelFields name,geneName\ longLabel gnomAD Predicted Missense Constraint Metrics By Gene (Z-scores) v2.1.1\ mouseOver Z: $_zscore\ GnomAD 4 used the whole-genome data from gnomAD 3 and added more exomes.\ The current v4.1 release includes a fix for the allele number\ issue.\ The v4.1 track shows variants from 807,162 individuals, including 730,947\ exomes and 76,215 genomes. This includes the 76,156 genomes from the gnomAD v3.1.2 release as well\ as new exome data from 416,555 UK Biobank individuals. For more detailed information on gnomAD\ v4.1, see the related blog post.\
\ \\ Following the conventions on the gnomAD browser, items are shaded according to their Annotation\ type:\
| pLoF | |
| Missense | |
| Synonymous | |
| Other |
\ Mouse hover on an item will display the following details about each variant:
\\ Clicking on an item will display additional details on the variant, including a population frequency\ table showing allele count in each sub-population.\
\ \\ To maintain consistency with the gnomAD website, variants are by default labeled according\ to their chromosomal start position followed by the reference and alternate alleles,\ for example "chr1-1234-T-CAG". dbSNP rsID's are also available as an additional\ label, if the variant is present in dbSnp.\
\ \\ Three filters are available for this track:\
\\ The gnomAD v4.1 data is unfiltered.
\ \\ For the full steps used to create the gnomAD tracks at UCSC, please see the\ hg38 gnomad makedoc.\
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below).
\ \\ The underlying bigBed only contains enough information necessary to use the track in the browser.\ The extra data like VEP annotations and CADD scores are available in the\ same directory\ as the bigBed but in the files details.tab.gz and details.tab.gz.gzi. The\ details.tab.gz contains the gzip compressed extra data in JSON format, and the .gzi file is\ available to speed searching of this data. Each variant has an associated md5sum in the name field\ of the bigBed which can be used along with the _dataOffset and _dataLen fields to get the\ associated external data. For example:
\ \\
# find an item of interest, the last two fields are _dataOffset and _dataLen:\
bigBedToBed genomes.bb stdout | head -4 | tail -1\
chr1 12416 12417 854246d79dc5d02dcdbd5f5438542b6e [..omitted..] 67293 902\
\
# use _dataOffset and _dataLen (add one to _dataLen for the newline character):\
bgzip -b 67293 -s 903 gnomad.v4.1.genomes.details.tab.gz\
854246d79dc5d02dcdbd5f5438542b6e {"DDX11L1": {"cons": ["non_coding_transcript_variant"...\
\
\
\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/v4.1/exomes/exomes.bb\ dataVersion Release v4.1 (April 19, 2024)\ defaultLabelFields _displayName\ detailsDynamicTable _jsonVep|Variant Effect Predictor,_jsonPopTable|Population Frequencies,_jsonHapTable|Haplotype Frequencies\ detailsTabUrls _dataOffset=/gbdb/hg38/gnomAD/v4.1/exomes/gnomad.v4.1.exomes.details.tab.gz\ filter.AF 0.0\ filterLabel.AF Minor Allele Frequency Filter\ filterType.FILTER multipleListAnd\ filterType.variation_type multipleListOr\ filterValues.FILTER PASS,InbreedingCoeff,RF,AC0,AS_VQSR,indel_stack (chrM only),npg (chrM only)\ filterValues.annot pLoF,missense,synonymous,other\ filterValues.variation_type 3_prime_UTR_variant,5_prime_UTR_variant,NMD_transcript_variant,coding_sequence_variant,frameshift_variant,incomplete_terminal_codon_variant,inframe_deletion,inframe_insertion,intron_variant,mature_miRNA_variant,missense_variant,non_coding_transcript_exon_variant,non_coding_transcript_variant,protein_altering_variant,splice_acceptor_variant,splice_donor_variant,splice_region_variant,start_lost,start_retained_variant,stop_gained,stop_lost,stop_retained_variant,synonymous_variant,transcript_ablation\ filterValuesDefault.FILTER PASS\ filterValuesDefault.annot pLoF,missense,synonymous\ html gnomadV4.1\ itemRgb on\ labelFields rsId,_displayName\ longLabel Genome Aggregation Database (gnomAD) Exomes Variants v4.1\ mouseOver Position: $chrom:${chromStart}-${chromEnd} ($ref/$alt)\ This container track helps call out sections of the genome that often cause problems or\ confusion when working with the genome. The hg19 genome has a track with the same name, but with\ more subtracks, as the GeT-RM and Genome-in-a-Bottle artifact variants do not exist \ for hg38.\ \
\ The Problematic Regions track contains the following subtracks:\
\ The Highly Reproducible Regions track highlights regions and variants\ from eight samples that can be used to assess variant detection pipelines. The\ "Highly Reproducible Regions" subtrack comprises the intersection of the reproducible\ regions across all eight samples, while the "Variants" subtracks contain the reproducible\ variants from each assayed sample. Both tracks contain data from the following samples:\
\The Genome in a Bottle (GIAB) Problematic Regions tracks provide stratifications of the\ genome to evaluate variant calls in complex regions. It is designed for use with Global Alliance\ for Genomic Health (GA4GH) benchmarking tools like\ hap.py\ and includes regions with low complexity, segmental duplications, functional regions,\ and difficult-to-sequence areas. Developed in collaboration with GA4GH, the\ Genome in a Bottle (GIAB) consortium, and the\ Telomere-to-Telomere Consortium (T2T), the dataset aims to standardize the\ analysis of genetic variation by offering pre-defined BED files for stratifying true and false\ positives in genomic studies, facilitating accurate assessments in complex areas of the genome.
\ \\ The creation of the GIAB Problematic Regions tracks involves using a pipeline and configuration to\ generate stratification BED files that categorize genomic regions based on specific challenges,\ such as low complexity or difficult mapping, to facilitate accurate benchmarking of variant calls.\ For more information on the pipeline and configuration used, please visit the following webpage:\ \ https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/genome-stratifications/v3.5/README.md.\ If you have questions or comments, please write to Justin Zook (jzook@nist.gov).
\ \\ The Panmask Easy 151b Regions subtrack contains a set of sample-agnostic easy regions where\ short-read variant calling reaches high accuracy. Easy regions are derived for variant filtration\ agnostic to individual samples. They are genomic intervals where general variant callers achieve\ high accuracy without sophisticated filtering.
\\ A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that\ could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples\ were inspected to exclude regions with excessively high or low coverage or those enriched with\ low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure\ are referred to as pm151:lenient, where "pm" stands for panmask. In addition, low\ complexity regions identified by SDUST were removed.
\The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and\ other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38,\ 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter\ variant calls for clinical or research human samples. Like the HighRepro track in this container\ (see above), it shows regions that are easy to sequence, not those that are problematic. The data\ was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.
\ \\ Each track contains a set of regions of varying length with no special configuration options. \ The UCSC Unusual Regions track has a mouse-over description, all other tracks have at most\ a name field, which can be shown in pack mode. The tracks are usually kept in dense mode.\
\ \\ The Hide empty subtracks control hides subtracks with no data in the browser window.\ Changing the browser window by zooming or scrolling may result in the display of a different\ selection of tracks.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator.\ \
\
For automated download and analysis, the genome annotation is stored in bigBed files that\
can be downloaded from\
our download server.\
Individual\
regions or the whole genome annotation can be obtained using our tool bigBedToBed\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g. \
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/problematic/comments.bb -chrom=chr21 -start=0 -end=100000000 stdout
\
\ Files were downloaded from the respective databases and converted to bigBed format.\ The procedure is documented in our\ hg38 makeDoc file.\
\ \\ Thanks to Anna Benet-Pagès, Max Haeussler, Angie Hinrichs, Daniel Schmelter, and Jairo\ Navarro at the UCSC Genome Browser for planning, building, and testing these tracks. The\ underlying data comes from the\ ENCODE Blacklist and some parts were copied manually from the HGNC and NCBI\ RefSeq tracks.\
\ \\ Amemiya HM, Kundaje A, Boyle AP.\ \ The ENCODE Blacklist: Identification of Problematic Regions of the Genome.\ Sci Rep. 2019 Jun 27;9(1):9354.\ PMID: 31249361; PMC: PMC6597582\
\ \\ Dwarshuis N, Kalra D, McDaniel J, Sanio P, Alvarez Jerez P, Jadhav B, Huang WE, Mondal R, Busby B,\ Olson ND et al.\ \ The GIAB genomic stratifications resource for human reference genomes.\ Nat Commun. 2024 Oct 19;15(1):9029.\ PMID: 39424793; PMC: PMC11489684\
\ \\ Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, Gonzalez-Porta M, Eberle MA,\ Tezak Z, Lababidi S et al.\ \ Best practices for benchmarking germline small-variant calls in human genomes.\ Nat Biotechnol. 2019 May;37(5):555-560.\ PMID: 30858580; PMC: PMC6699627\
\ \\ Li H.\ \ Finding easy regions for short-read variant calling from pangenome data.\ ArXiv. 2025 Aug 8;.\ PMID: 40799803; PMC: PMC12340882\
\ \\ Pan B, Ren L, Onuchic V, Guan M, Kusko R, Bruinsma S, Trigg L, Scherer A, Ning B, Zhang C et\ al.\ \ Assessing reproducibility of inherited variants detected with short-read whole genome\ sequencing.\ Genome Biol. 2022 Jan 3;23(1):2.\ PMID: 34980216; PMC: PMC8722114\
\ map 1 compositeTrack on\ html problematic\ longLabel Highly Reproducible genomic regions for sequencing\ parent problematicSuper\ priority 2\ shortLabel Highly Reproducible Regions\ subGroup1 view Views beds=Regions vcfs=Variants\ track highlyReproducible\ type bed 3\ visibility hide\ highReproBeds Highly Reproducible Regions bigBed 9 + Highly Reproducible Regions 1 2 0 0 0 127 127 127 0 0 0 map 1 longLabel Highly Reproducible Regions\ parent highlyReproducible\ shortLabel Highly Reproducible Regions\ track highReproBeds\ type bigBed 9 +\ view beds\ visibility dense\ highReproVcfs Highly Reproducible Variants vcfTabix Highly Reproducible Variants 0 2 0 0 0 127 127 127 0 0 0 map 1 hideEmptySubtracks on\ longLabel Highly Reproducible Variants\ parent highlyReproducible\ shortLabel Highly Reproducible Variants\ track highReproVcfs\ type vcfTabix\ view vcfs\ visibility hide\ hmc HMC bigWig HMC - Homologous Missense Constraint Score on PFAM domains 2 2 0 130 0 127 192 127 0 0 0\ The "Constraint scores" container track includes several subtracks showing the results of\ constraint prediction algorithms. These try to find regions of negative\ selection, where variations likely have functional impact. The algorithms do\ not use multi-species alignments to derive evolutionary constraint, but use\ primarily human variation, usually from variants collected by gnomAD (see the\ gnomAD V2 or V3 tracks on hg19 and hg38) or TOPMED (contained in our dbSNP\ tracks and available as a filter). One of the subtracks is based on UK Biobank\ variants, which are not available publicly, so we have no track with the raw data.\ The number of human genomes that are used as the input for these scores are\ 76k, 53k and 110k for gnomAD, TOPMED and UK Biobank, respectively.\
\ \Note that another important constraint score, gnomAD\ constraint, is not part of this container track but can be found in the hg38 gnomAD\ track.\
\ \ The algorithms included in this track are:\\ JARVIS scores are shown as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The scores were downloaded and converted to a single bigWig file.\ Move the mouse over the bars to display the exact values. A horizontal line is shown at the 0.733\ value which signifies the 90th percentile.
\ See hg19 makeDoc and\ hg38 makeDoc.\\ Interpretation: The authors offer a suggested guideline of > 0.9998 for identifying\ higher confidence calls and minimizing false positives. In addition to that strict threshold, the \ following two more relaxed cutoffs can be used to explore additional hits. Note that these\ thresholds are offered as guidelines and are not necessarily representative of pathogenicity.
\ \\
| Percentile | JARVIS score threshold |
|---|---|
| 99th | 0.9998 |
| 95th | 0.9826 |
| 90th | 0.7338 |
\ HMC scores are displayed as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The highly-constrained cutoff\ of 0.8 is indicated with a line.
\\ Interpretation: \ A protein residue with HMC score <1 indicates that missense variants affecting\ the homologous residues are significantly under negative selection (P-value <\ 0.05) and likely to be deleterious. A more stringent score threshold of HMC<0.8\ is recommended to prioritize predicted disease-associated variants.\
\ \\ Interpretation: The authors suggest the following guidelines for evaluating\ intolerance. By default, the MetaDome track displays a horizontal line at 0.7 which \ signifies the first intolerant bin. For more information see the MetaDome publication.
\ \\
| Classification | MetaDome Tolerance Score |
|---|---|
| Highly intolerant | ≤ 0.175 |
| Intolerant | ≤ 0.525 |
| Slightly intolerant | ≤ 0.7 |
\ MTR data can be found on two tracks, MTR All data and MTR Scores. In the\ MTR Scores track the data has been converted into 4 separate signal tracks\ representing each base pair mutation, with the lowest possible score shown when\ multiple transcripts overlap at a position. Overlaps can happen since this score\ is derived from transcripts and multiple transcripts can overlap. \ A horizontal line is drawn on the 0.8 score line\ to roughly represent the 25th percentile, meaning the items below may be of particular\ interest. It is recommended that the data be explored using\ this version of the track, as it condenses the information substantially while\ retaining the magnitude of the data.
\ \Any specific point mutations of interest can then be researched in the \ MTR All data track. This track contains all of the information from\ \ MTRV2 including more than 3 possible scores per base when transcripts overlap.\ A mouse-over on this track shows the ref and alt allele, as well as the MTR score\ and the MTR score percentile. Filters are available for MTR score, False Discovery Rate\ (FDR), MTR percentile, and variant consequence. By default, only items in the bottom\ 25 percentile are shown. Items in the track are colored according\ to their MTR percentile:
\\ Interpretation: Regions with low MTR scores were seen to be enriched with\ pathogenic variants. For example, ClinVar pathogenic variants were seen to\ have an average score of 0.77 whereas ClinVar benign variants had an average score\ of 0.92. Further validation using the FATHMM cancer-associated training dataset saw\ that scores less than 0.5 contained 8.6% of the pathogenic variants while only containing\ 0.9% of neutral variants. In summary, lower scores are more likely to represent\ pathogenic variants whereas higher scores could be pathogenic, but have a higher chance\ to be a false positive. For more information see the MTR-Viewer publication.
\ \\ Scores were downloaded and converted to a single bigWig file. See the\ hg19 makeDoc and the\ hg38 makeDoc for more info.\
\ \\ Scores were downloaded and converted to .bedGraph files with a custom Python \ script. The bedGraph files were then converted to bigWig files, as documented in our \ makeDoc hg19 build log.
\ \\
The authors provided a bed file containing codon coordinates along with the scores. \
This file was parsed with a python script to create the two tracks. For the first track\
the scores were aggregated for each coordinate, then the lowest score chosen for any\
overlaps and the result written out to bedGraph format. The file was then converted\
to bigWig with the bedGraphToBigWig utility. For the second track the file\
was reorganized into a bed 4+3 and conveted to bigBed with the bedToBigBed\
utility.
\ See the hg19 makeDoc for details including the build script.
\\ The raw MetaDome data can also be accessed via their Zenodo handle.
\ \\ V2\ file was downloaded and columns were reshuffled as well as itemRgb added for the\ MTR All data track. For the MTR Scores track the file was parsed with a python\ script to pull out the highest possible MTR score for each of the 3 possible mutations\ at each base pair and 4 tracks built out of these values representing each mutation.
\\ See the hg19 makeDoc entry on MTR for more info.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/hmc/hmc.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \ \\ Thanks to Jean-Madeleine Desainteagathe (APHP Paris, France) for suggesting the JARVIS, MTR, HMC tracks. Thanks to Xialei Zhang for providing the HMC data file and to Dimitrios Vitsios and Slave Petrovski for helping clean up the hg38 JARVIS files for providing guidance on interpretation. Additional\ thanks to Laurens van de Wiel for providing the MetaDome data as well as guidance on the track development and interpretation. \
\ \ \\ Vitsios D, Dhindsa RS, Middleton L, Gussow AB, Petrovski S.\ \ Prioritizing non-coding regions based on human genomic constraint and sequence context with deep\ learning.\ Nat Commun. 2021 Mar 8;12(1):1504.\ PMID: 33686085; PMC: PMC7940646\
\ \\ Xiaolei Zhang, Pantazis I. Theotokis, Nicholas Li, the SHaRe Investigators, Caroline F. Wright, Kaitlin E. Samocha, Nicola Whiffin, James S. Ware\ \ Genetic constraint at single amino acid resolution improves missense variant prioritisation and gene discovery.\ Medrxiv 2022.02.16.22271023\
\ \\ Wiel L, Baakman C, Gilissen D, Veltman JA, Vriend G, Gilissen C.\ \ MetaDome: Pathogenicity analysis of genetic variants through aggregation of homologous human protein\ domains.\ Hum Mutat. 2019 Aug;40(8):1030-1038.\ PMID: 31116477; PMC: PMC6772141\
\ \\ Silk M, Petrovski S, Ascher DB.\ \ MTR-Viewer: identifying regions within genes under purifying selection.\ Nucleic Acids Res. 2019 Jul 2;47(W1):W121-W126.\ PMID: 31170280; PMC: PMC6602522\
\ \\ Halldorsson BV, Eggertsson HP, Moore KHS, Hauswedell H, Eiriksson O, Ulfarsson MO, Palsson G,\ Hardarson MT, Oddsson A, Jensson BO et al.\ \ The sequences of 150,119 genomes in the UK Biobank.\ Nature. 2022 Jul;607(7920):732-740.\ PMID: 35859178; PMC: PMC9329122\
\ \ \\ Huang YF, Gulko B, Siepel A.\ \ Fast, scalable prediction of deleterious noncoding variants from functional and population genomic\ data.\ Nat Genet. 2017 Apr;49(4):618-624.\ PMID: 28288115; PMC: PMC5395419\
\ \ phenDis 0 bigDataUrl /gbdb/hg38/hmc/hmc.bw\ color 0,130,0\ html constraintSuper\ longLabel HMC - Homologous Missense Constraint Score on PFAM domains\ maxHeightPixels 128:40:8\ maxWindowToDraw 10000000\ mouseOverFunction noAverage\ parent constraintSuper\ priority 2\ shortLabel HMC\ track hmc\ type bigWig\ viewLimits 0:2\ viewLimitsMax 0:2\ visibility full\ yLineMark 0.8\ yLineOnOff on\ covidHgiGwasB2 Hosp COVID GWAS bigLolly 9 + Hospitalized COVID GWAS from the COVID-19 Host Genetics Initiative (3199 cases, 8 studies) 0 2 0 0 0 127 127 127 0 0 22 chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22, phenDis 1 bigDataUrl /gbdb/hg38/covidHgiGwas/covidHgiGwasB2.hg38.bb\ longLabel Hospitalized COVID GWAS from the COVID-19 Host Genetics Initiative (3199 cases, 8 studies)\ parent covidHgiGwas off\ shortLabel Hosp COVID GWAS\ track covidHgiGwasB2\ covidHgiGwasR4PvalB2 Hosp COVID vars bigLolly 9 + Hospitalized COVID risk variants from the COVID-19 HGI GWAS Analysis B2 (7885 cases, 21 studies, Rel 4: Oct 2020) 0 2 0 0 0 127 127 127 0 0 22 chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22, phenDis 1 bigDataUrl /gbdb/hg38/covidHgiGwas/covidHgiGwasR4.B2.hg38.bb\ longLabel Hospitalized COVID risk variants from the COVID-19 HGI GWAS Analysis B2 (7885 cases, 21 studies, Rel 4: Oct 2020)\ parent covidHgiGwasR4Pval on\ priority 2\ shortLabel Hosp COVID vars\ track covidHgiGwasR4PvalB2\ hgdp Human Genome Diversity Project, 1k WGS vcfTabix Phased Variants: Human Genome Diversity Project (HGDP) - 1043 samples, isolated populations 3 2 0 0 0 127 127 127 0 0 0\ This tracks contains variants of individual genotypes, usually phased, from the projects\ Human Diversity Genome Project, Simons Genome Diversity Project, gnomad's HGDP+1000 Genomes callset,\ and the Mexico Biobank.\ The original release of 1000 Genomes has its own, separate track.\ Projects where the released variants are not phased can be found in the container track "SNV Frequencies".\
\ \\ Available on hg19 and hg38:
\\ Available only on hg38:
\\ Full haplotype display:\ In "pack" mode, this track sorts the haplotypes. This can be\ useful for determining the similarity between the samples and inferring\ inheritance at a particular locus.\ Each sample's phased and/or homozygous genotypes are split into haplotypes,\ clustered by similarity around a central variant (in pink), and sorted for\ display by their position in the clustering tree. Click a variant to center on it.\ The tree (as space allows) is drawn in the label area next to the track image.\ Leaf clusters, in which all haplotypes are identical (at least for the variants\ used in clustering), are colored purple. \
\\ For a full description of how the display works, please see our \ Haplotype Display help page.\ \
\ MXB: Allele frequencies by geographical state and ancestry are available via\ the MexVar platform.\ Raw genotype data are available under controlled access at the\ EGA (Study: EGAS00001005797; Dataset: EGAD00010002361). For the VCFs, email\ andres.moreno@cinvestav.mx.\
\ \\ SGDP: The version used was\ https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/,\ merged with bcftools and lifted to hg38 with CrossMap. \
\ \\ MXB: We thank the Center for Research and Advanced Studies (Cinvestav) of Mexico for\ generating and providing the frequency data, the National Institute of Medical\ Sciences and Nutrition (INCMNSZ) for DNA extraction, and the Ministry of Health\ together with the National Institute of Public Health (INSP) for the design and\ implementation of the National Health Survey 2000 (ENSA 2000). We also thank\ the ENSA-Genomics Consortium for their contributions to sample collection and\ data processing that made possible the construction of the MXB genomic\ resource.\
\\ SGDP: This project was funded by the Simons Foundation. Thanks to David Reich and Swapan \ Mallick for help with importing the data.\
\ \\ Barberena-Jonas C, Medina-Muñoz SG, Cedillo-Castelán V, Sepúlveda-Morales T,\ Gonzaga-Jáuregui C, ENSA Genomics Consortium, García-García L, Ioannidis AG,\ Moreno-Estrada A.\ \ Clinical genetic variation across Hispanic populations in the Mexican Biobank.\ Nat Med. 2026 Jan 21;.\ DOI: 10.1038/s41591-025-04100-z; PMID: 41566040\
\ \\ Sohail M, Moreno-Estrada A.\ \ The Mexican Biobank Project promotes genetic discovery, inclusive science and local capacity\ building.\ Dis Model Mech. 2024 Jan 1;17(1).\ PMID: 38299665; PMC: PMC10855211\
\ \\ Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Corés CD, Barberena-Jonas C, Medina-Muñoz SG,\ Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L et al.\ \ Mexican Biobank advances population and medical genomics of diverse ancestries.\ Nature. 2023 Oct;622(7984):775-783.\ PMID: 37821706; PMC: PMC10600006\
\ \\ Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J\ et al.\ \ Insights into human genetic variation and population history from 929 diverse genomes.\ Science. 2020 Mar 20;367(6484).\ PMID: 32193295; PMC: PMC7115999\
\ \\ Koenig Z, Yohannes MT, Nkambule LL, Zhao X, Goodrich JK, Kim HA, Wilson MW, Tiao G, Hao SP, Sahakian\ N et al.\ \ A harmonized public resource of deeply sequenced diverse human genomes.\ Genome Res. 2024 Jun 25;34(5):796-809.\ PMID: 38749656; PMC: PMC11216312\
\ \\ Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S,\ Tandon A et al.\ \ The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.\ Nature. 2016 Oct 13;538(7624):201-206.\ PMID: 27654912; PMC: PMC5161557\
\ \ varRep 1 bigDataUrl /gbdb/hg38/phasedVars/hgdp/hgdp_wgs.20190516.full.vcf.gz\ dataVersion 2019-05-16\ html phasedVars.html\ longLabel Phased Variants: Human Genome Diversity Project (HGDP) - 1043 samples, isolated populations\ parent phasedVars on\ priority 2\ shortLabel Human Genome Diversity Project, 1k WGS\ track hgdp\ type vcfTabix\ visibility pack\ humanMethylationAtlasSignals Human Methylation Atlas Signals bigWig Human Methylation Atlas WGBS cell type signals 2 2 0 0 0 127 127 127 0 0 0\ The Human Methylation Atlas tracks display genome-wide DNA methylation profiles from \ deep whole-genome bisulfite sequencing (WGBS) of 39 primary human cell types \ sorted from 205 healthy tissue samples. This comprehensive resource enables fragment-level \ analysis across thousands of unique markers, providing a detailed reference for \ cell-type-specific methylation patterns.\
\ \\ The 205 samples from 39 cell type groups are organized into two data types:\
\\ DNA methylation patterns are highly reproducible across individuals of the same cell type \ (>99.5% identical), reflecting the robustness of cell identity programs.\
\ \\ Signal tracks display methylation beta values on a 0-1 scale, where 0 indicates fully \ unmethylated CpGs and 1 indicates fully methylated CpGs. A value of -1 indicates \ missing data. For optimal comparison across cell types, set the vertical viewing range \ to 0-1 with auto-scale off.\
\ \\ Tracks are colored by tissue/cell type category as follows:\
\ \| Color | Cell Type(s) |
|---|---|
| Neurons | |
| Oligodendrocytes | |
| Thyroid Epithelium | |
| Prostate Epithelium | |
| Bladder Epithelium | |
| Heart Cardiomyocytes | |
| Smooth Muscle | |
| Heart Fibroblasts | |
| Skeletal Muscle | |
| Erythrocyte Progenitors | |
| Blood Granulocytes | |
| Blood Monocytes/Macrophages | |
| Blood T Cells | |
| Blood B Cells | |
| Blood NK Cells | |
| Pancreas Beta Cells | |
| Pancreas Alpha Cells | |
| Pancreas Delta Cells | |
| Pancreas Duct Cells | |
| Pancreas Acinar Cells | |
| Colon Epithelium | |
| Colon Fibroblasts | |
| Small Intestine Epithelium | |
| Gastric Epithelium | |
| Gallbladder | |
| Liver Hepatocytes | |
| Lung Bronchus Epithelium | |
| Lung Alveolar Epithelium | |
| Kidney Epithelium | |
| Endothelial | |
| Breast Basal Epithelium | |
| Breast Luminal Epithelium | |
| Fallopian Epithelium | |
| Ovary Epithelium | |
| Adipocytes | |
| Epidermal Keratinocytes | |
| Dermal Fibroblasts | |
| Bone Osteoblasts | |
| Head Neck Epithelium |
\ Primary human cells were isolated from freshly dissociated adult healthy tissues using \ fluorescence-activated cell sorting (FACS), yielding high-purity preparations across major \ cell lineages. A total of 205 samples representing 77 primary cell types were collected from\ 137 consenting donors and merged into 39 cell type groups based on methylation similarity.\ Average sample purity exceeded 90% as determined by flow cytometry, gene expression, and\ DNA methylation analysis. Some cell types showed lower purity, including colon fibroblasts (78%),\ smooth muscle cells (82%), endothelial cells (86%), and adipocytes (87%).\
\ \\ Several cell types are absent from the atlas, typically due to limited availability of primary\ material. These include osteoblasts, cholangiocytes, cells of the adrenal gland, urethral\ epithelium, and haematopoietic stem cells. Subpopulations of interest, such as distinct neuronal or\ lymphocyte subtypes, were also not resolved separately.\
\ \\ Whole-genome bisulfite sequencing was performed using 150 bp paired-end reads at an average \ sequencing depth of 30× (minimum 6.62×). Libraries were prepared using the \ Accel-NGS Methyl-Seq DNA library preparation kit and sequenced on the Illumina NovaSeq 6000 \ platform.\
\ \\ Reads were mapped to the human genome (hg38) using bwa-meth, deduplicated with Sambamba, \ and processed into per-CpG methylation calls. The genome was segmented into 7.1 million \ non-overlapping methylation blocks using a multi-channel dynamic programming algorithm \ that identifies regions of homogeneous methylation across samples.\
\ \\ Cell-type-specific differentially methylated regions were identified using a one-versus-all \ comparison approach. Regions uniquely unmethylated in specific cell types were found to be \ enriched for transcriptional enhancers and tissue-specific transcription factor binding motifs.\
\ \\ Data processing was performed using \ wgbstools, an open-source \ computational suite for DNA methylation sequencing data representation, visualization, \ and analysis.\
\ \\ The raw data for these tracks can be explored interactively using the \ Table Browser or the \ Data Integrator. \ For automated analysis, the data may also be queried from our \ REST API.\
\ \\ The complete dataset, including all WGBS data files and processed methylation calls, \ is available from GEO accession \ GSE186458.\
\ \\ For questions regarding the data, please contact \ Prof. Tommy Kaplan at the Hebrew \ University of Jerusalem.\
\ \\ Data generation and analysis were performed at the Hebrew University of Jerusalem by the \ Dor, Kaplan, and Glaser laboratories and collaborators. Sample collection involved \ collaboration with Hadassah Medical Center, Oregon Health & Science University, \ Karolinska Institute, and University of Alberta.\
\ \\ Loyfer N, Magenheim J, Peretz A, Cann G, Bredno J, Klochendler A, Fox-Fisher I, \ Shabi-Porat S, Hecht M, Pelet T et al.\ \ A DNA methylation atlas of normal human cell types.\ Nature. 2023 Jan;613(7943):355-364.\ PMID: 36599988\
\ \\ Loyfer N, Rosenski J, Kaplan T.\ \ wgbstools: a computational suite for DNA methylation sequencing data analysis.\ Life Sci Alliance. 2026 Apr;9(4):e202503514.\ PMID: 41611450\
\ \ regulation 0 compositeTrack on\ dataVersion Data release version 2\ dimensions dimY=cellType dimX=dataType\ html methylationAtlasSignals.html\ longLabel Human Methylation Atlas WGBS cell type signals\ parent dnaMethylation\ priority 2\ shortLabel Human Methylation Atlas Signals\ subGroup1 cellType Cell_Type Neuron=Neuron Oligodend=Oligodend Heart-Cardio=Heart_Cardio Smooth-Musc=Smooth_Musc Heart-Fibro=Heart_Fibro Skeletal-Musc=Skeletal_Musc Adipocytes=Adipocytes Endothel=Endothel Blood-T=Blood_T Blood-B=Blood_B Blood-NK=Blood_NK Blood-Mono-Macro=Blood_Mono_Macro Blood-Granul=Blood_Granul Eryth-prog=Eryth_prog Head-Neck-Ep=Head_Neck_Ep Lung-Ep-Bron=Lung_Ep_Bron Lung-Ep-Alveo=Lung_Ep_Alveo Breast-Basal-Ep=Breast_Basal_Ep Breast-Luminal-Ep=Breast_Luminal_Ep Pancreas-Alpha=Pancreas_Alpha Pancreas-Beta=Pancreas_Beta Pancreas-Delta=Pancreas_Delta Pancreas-Acinar=Pancreas_Acinar Pancreas-Duct=Pancreas_Duct Liver-Hep=Liver_Hep Kidney-Ep=Kidney_Ep Gastric-Ep=Gastric_Ep Small-Int-Ep=Small_Int_Ep Colon-Ep=Colon_Ep Bladder-Ep=Bladder_Ep Prostate-Ep=Prostate_Ep Fallopian-Ep=Fallopian_Ep Ovary-Ep=Ovary_Ep Dermal-Fibro=Dermal_Fibro Epid-Kerat=Epid_Kerat Gallbladder=Gallbladder Colon-Fibro=Colon_Fibro Thyroid-Ep=Thyroid_Ep Bone-Osteob=Bone_Osteob\ subGroup2 dataType Data_Type Merged=Merged Replicate=Replicates\ track humanMethylationAtlasSignals\ type bigWig\ visibility full\ yLineOnOff on\ xGen_Research_Targets_V1 IDT xGen V1 T bigBed IDT - xGen Exome Research Panel V1 Target Regions 0 2 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/xgen-exome-research-panel-targets-hg38.bb\ color 100,143,255\ longLabel IDT - xGen Exome Research Panel V1 Target Regions\ parent exomeProbesets off\ shortLabel IDT xGen V1 T\ track xGen_Research_Targets_V1\ type bigBed\ nestedRepeats Interrupted Rpts bed 12 + Fragments of Interrupted Repeats Joined by RepeatMasker ID 0 2 0 0 0 127 127 127 1 0 0\ This track shows joined fragments of interrupted repeats extracted\ from the output of the \ RepeatMasker program which screens DNA sequences\ for interspersed repeats and low complexity DNA sequences using the\ \ Repbase Update library of repeats from the\ Genetic\ Information Research Institute (GIRI). Repbase Update is described in\ Jurka (2000) in the References section below.\
\ \\ The detailed annotations from RepeatMasker are in the RepeatMasker track. This\ track shows fragments of original repeat insertions which have been interrupted\ by insertions of younger repeats or through local rearrangements. The fragments\ are joined using the ID column of RepeatMasker output.\
\ \\ In pack or full mode, each interrupted repeat is displayed as boxes\ (fragments) joined by horizontal lines, labeled with the repeat name.\ If all fragments are on the same strand, arrows are added to the\ horizontal line to indicate the strand. In dense or squish mode, labels\ and arrows are omitted and in dense mode, all items are collapsed to\ fit on a single row.\
\ \\ Items are shaded according to the average identity score of their\ fragments. Usually, the shade of an item is similar to the shades of\ its fragments unless some fragments are much more diverged than\ others. The score displayed above is the average identity score,\ clipped to a range of 50% - 100% and then mapped to the range\ 0 - 1000 for shading in the browser.\
\ \\ UCSC has used the most current versions of the RepeatMasker software\ and repeat libraries available to generate these data. Note that these\ versions may be newer than those that are publicly available on the Internet.\
\ \\ Data are generated using the RepeatMasker -s flag. Additional flags\ may be used for certain organisms. See the\ FAQ for more information.\
\ \\ Thanks to Arian Smit, Robert Hubley and GIRI for providing the tools and\ repeat libraries used to generate this track.\
\ \\ Smit AFA, Hubley R, Green P.\ RepeatMasker Open-3.0.\ \ https://www.repeatmasker.org/. 1996-2010.\
\ \\ Repbase Update is described in:\
\ \\ Jurka J.\ \ Repbase Update: a database and an electronic journal of repetitive elements.\ Trends Genet. 2000 Sep;16(9):418-420.\ PMID: 10973072\
\ \\ For a discussion of repeats in mammalian genomes, see:\
\ \\ Smit AF.\ \ Interspersed repeats and other mementos of transposable elements in mammalian genomes.\ Curr Opin Genet Dev. 1999 Dec;9(6):657-63.\ PMID: 10607616\
\ \\ Smit AF.\ \ The origin of interspersed repeats in the human genome.\ Curr Opin Genet Dev. 1996 Dec;6(6):743-8.\ PMID: 8994846\
\ rep 1 exonNumbers off\ group rep\ longLabel Fragments of Interrupted Repeats Joined by RepeatMasker ID\ priority 2\ shortLabel Interrupted Rpts\ track nestedRepeats\ type bed 12 +\ useScore 1\ visibility hide\ jaspar2022 JASPAR 2022 TFBS bigBed 6 + JASPAR CORE 2022 - Predicted Transcription Factor Binding Sites 0 2 0 0 0 127 127 127 1 0 0 http://jaspar.genereg.net/search?q=$$&collection=all&tax_group=all&tax_id=all&type=all&class=all&family=all&version=all regulation 1 bigDataUrl /gbdb/hg38/jaspar/JASPAR2022.bb\ filterValues.TFName Ahr::Arnt,Alx1,ALX3,Alx4,Ar,ARGFX,Arid3a,Arid3b,Arid5a,Arnt,ARNT2,ARNT::HIF1A,Arntl,Arx,ASCL1,Ascl2,Atf1,ATF2,Atf3,ATF3,ATF4,ATF6,ATF7,Atoh1,ATOH7,BACH1,Bach1::Mafk,BACH2,BARHL1,BARHL2,BARX1,BARX2,BATF,BATF3,BATF::JUN,Bcl11B,BCL6,BCL6B,Bhlha15,BHLHA15,BHLHE22,BHLHE23,BHLHE40,BHLHE41,BNC2,BSX,CDX1,CDX2,CDX4,CEBPA,CEBPB,CEBPD,CEBPE,CEBPG,CLOCK,CREB1,CREB3,CREB3L1,Creb3l2,CREB3L4,Creb5,CREM,Crx,CTCF,CTCFL,CUX1,CUX2,DBP,Ddit3::Cebpa,DLX1,Dlx2,Dlx3,Dlx4,Dlx5,DLX6,Dmbx1,Dmrt1,DMRT3,DMRTA1,DMRTA2,DMRTC2,DPRX,DRGX,Dux,DUX4,DUXA,E2F1,E2F2,E2F3,E2F4,E2F6,E2F7,E2F8,EBF1,Ebf2,EBF3,EGR1,EGR2,EGR3,EGR4,EHF,ELF1,ELF2,ELF3,ELF4,Elf5,ELK1,ELK1::HOXA1,ELK1::HOXB13,ELK1::SREBF2,ELK3,ELK4,EMX1,EMX2,EN1,EN2,EOMES,ERF,ERF::FIGLA,ERF::FOXI1,ERF::FOXO1,ERF::HOXB13,ERF::NHLH1,ERF::SREBF2,Erg,ESR1,ESR2,ESRRA,ESRRB,Esrrg,ESX1,ETS1,ETS2,ETV1,ETV2,ETV2::DRGX,ETV2::FIGLA,ETV2::FOXI1,ETV2::HOXB13,ETV3,ETV4,ETV5,ETV5::DRGX,ETV5::FIGLA,ETV5::FOXI1,ETV5::FOXO1,ETV5::HOXA2,ETV6,ETV7,EVX1,EVX2,EWSR1-FLI1,FERD3L,FEV,FIGLA,FLI1,FLI1::DRGX,FLI1::FOXI1,FOS,FOSB::JUN,FOSB::JUNB,FOS::JUN,FOS::JUNB,FOS::JUND,FOSL1,FOSL1::JUN,FOSL1::JUNB,FOSL1::JUND,FOSL2,FOSL2::JUN,FOSL2::JUNB,FOSL2::JUND,FOXA1,FOXA2,FOXA3,FOXB1,FOXC1,FOXC2,FOXD1,FOXD2,FOXD3,FOXE1,Foxf1,FOXF2,FOXG1,FOXH1,FOXI1,Foxj2,FOXJ2::ELF1,Foxj3,FOXK1,FOXK2,FOXL1,Foxl2,Foxn1,FOXN3,Foxo1,FOXO1::ELF1,FOXO1::ELK1,FOXO1::ELK3,FOXO1::FLI1,Foxo3,FOXO4,FOXO6,FOXP1,FOXP2,FOXP3,Foxq1,GABPA,GATA1,GATA1::TAL1,GATA2,Gata3,GATA4,GATA5,GATA6,GBX1,GBX2,GCM1,GCM2,GFI1,Gfi1B,Gli1,Gli2,GLI3,GLIS1,GLIS2,GLIS3,Gmeb1,GMEB2,GRHL1,GRHL2,GSC,GSC2,GSX1,GSX2,Hand1::Tcf3,HAND2,HES1,HES2,HES5,HES6,HES7,HESX1,HEY1,HEY2,Hic1,HIC2,HIF1A,HINFP,HLF,HMBOX1,Hmx1,Hmx2,Hmx3,Hnf1A,HNF1A,HNF1B,HNF4A,HNF4G,HOXA1,HOXA10,Hoxa11,Hoxa13,HOXA2,HOXA4,HOXA5,HOXA6,HOXA7,HOXA9,HOXB13,HOXB2,HOXB2::ELK1,HOXB3,HOXB4,HOXB5,HOXB6,HOXB7,HOXB8,HOXB9,HOXC10,HOXC11,HOXC12,HOXC13,HOXC4,HOXC8,HOXC9,HOXD10,HOXD11,HOXD12,HOXD12::ELK1,Hoxd13,HOXD3,HOXD4,HOXD8,HOXD9,HSF1,HSF2,HSF4,IKZF1,Ikzf3,INSM1,Irf1,IRF2,IRF3,IRF4,IRF5,IRF6,IRF7,IRF8,IRF9,Isl1,ISL2,ISX,JDP2,Jun,JUN,JUNB,JUND,JUN::JUNB,KLF1,KLF10,KLF11,KLF12,KLF13,KLF14,KLF15,KLF16,KLF17,KLF2,KLF3,KLF4,KLF5,KLF6,KLF7,KLF9,LBX1,LBX2,Lef1,Lhx1,LHX2,Lhx3,Lhx4,LHX5,LHX6,Lhx8,LHX9,LIN54,LMX1A,LMX1B,MAF,MAFA,Mafb,MAFF,Mafg,MAFG::NFE2L1,MAFK,MAF::NFE2,MAX,MAX::MYC,MAZ,Mecom,MEF2A,MEF2B,MEF2C,MEF2D,MEIS1,MEIS2,MEIS3,MEOX1,MEOX2,MGA,MGA::EVX1,MITF,mix-a,MIXL1,MLX,Mlxip,MLXIPL,MNT,MNX1,MSANTD3,MSC,Msgn1,MSX1,MSX2,Msx3,MTF1,MXI1,MYB,MYBL1,MYBL2,MYC,MYCN,MYF5,MYF6,MYOD1,MYOG,MZF1,NEUROD1,Neurod2,NEUROG1,NEUROG2,Nfat5,Nfatc1,Nfatc2,NFATC3,NFATC4,NFE2,Nfe2l2,NFIA,NFIB,NFIC,NFIC::TLX1,NFIL3,NFIX,NFKB1,NFKB2,NFYA,NFYB,NFYC,NHLH1,NHLH2,Nkx2-1,NKX2-2,NKX2-3,NKX2-4,NKX2-5,NKX2-8,Nkx3-1,Nkx3-2,NKX6-1,NKX6-2,NKX6-3,Nobox,NOTO,Npas2,Npas4,NR1D1,NR1D2,Nr1H2,NR1H2::RXRA,Nr1h3::Rxra,Nr1H4,NR1H4::RXRA,NR1I2,NR1I3,NR2C1,NR2C2,Nr2e1,Nr2e3,NR2F1,NR2F2,Nr2f6,Nr2F6,NR2F6,NR3C1,NR3C2,NR4A1,NR4A2,NR4A2::RXRA,NR5A1,Nr5A2,NR6A1,Nrf1,NRL,OLIG1,Olig2,OLIG2,OLIG3,ONECUT1,ONECUT2,ONECUT3,OSR1,OSR2,OTX1,OTX2,OVOL1,OVOL2,PATZ1,PAX1,PAX2,PAX3,PAX4,PAX5,PAX6,Pax7,PAX9,PBX1,PBX2,PBX3,PDX1,PHOX2A,PHOX2B,PITX1,PITX2,PITX3,PKNOX1,PKNOX2,PLAG1,Plagl1,PLAGL2,POU1F1,POU2F1,POU2F1::SOX2,POU2F2,POU2F3,POU3F1,POU3F2,POU3F3,POU3F4,POU4F1,POU4F2,POU4F3,POU5F1,POU5F1B,Pou5f1::Sox2,POU6F1,POU6F2,PPARA::RXRA,PPARD,PPARG,Pparg::Rxra,PRDM1,Prdm14,Prdm15,Prdm4,Prdm5,PRDM9,PROP1,PROX1,PRRX1,PRRX2,Ptf1a,Ptf1A,RARA,RARA::RXRA,RARA::RXRG,Rarb,RARB,Rarg,RARG,RAX,RAX2,RBPJ,Rbpjl,REL,RELA,RELB,REST,RFX1,RFX2,RFX3,RFX4,RFX5,Rfx6,RFX7,Rhox11,RHOXF1,RORA,RORB,RORC,RREB1,Runx1,RUNX2,RUNX3,Rxra,RXRA::VDR,RXRB,RXRG,SATB1,SCRT1,SCRT2,Sf1,SHOX,Shox2,SIX1,SIX2,Six3,Six4,SMAD2,Smad2::Smad3,SMAD2::SMAD3::SMAD4,SMAD3,Smad4,SMAD5,SNAI1,SNAI2,SNAI3,SOHLH2,Sox1,SOX10,Sox11,SOX12,SOX13,SOX14,SOX15,Sox17,SOX18,SOX2,SOX21,Sox3,SOX4,Sox5,Sox6,SOX8,SOX9,SP1,SP2,SP3,SP4,SP5,SP8,SP9,SPDEF,Spi1,SPIB,SPIC,Spz1,SREBF1,SREBF2,SRF,SRY,STAT1,STAT1::STAT2,Stat2,STAT3,Stat4,Stat5a,Stat5a::Stat5b,Stat5b,Stat6,TAL1::TCF3,TBP,TBR1,TBX1,TBX15,TBX18,TBX19,TBX2,TBX20,TBX21,TBX3,TBX4,TBX5,Tbx6,TBXT,Tcf12,TCF12,Tcf21,TCF21,TCF3,TCF4,TCF7,TCF7L1,TCF7L2,TCFL5,TEAD1,TEAD2,TEAD3,TEAD4,TEF,TFAP2A,TFAP2B,TFAP2C,TFAP2E,TFAP4,TFAP4::ETV1,TFAP4::FLI1,TFCP2,Tfcp2l1,TFDP1,TFE3,TFEB,TFEC,TGIF1,TGIF2,TGIF2LX,TGIF2LY,THAP1,Thap11,THRA,THRB,TLX2,TP53,TP63,TP73,TRPS1,TWIST1,Twist2,UNCX,USF1,USF2,VAX1,VAX2,Vdr,VENTX,VEZF1,VSX1,VSX2,Wt1,XBP1,Yy1,YY2,ZBED1,ZBED2,ZBTB12,ZBTB14,ZBTB18,ZBTB26,ZBTB32,ZBTB33,ZBTB6,ZBTB7A,ZBTB7B,ZBTB7C,ZEB1,ZFP14,Zfp335,ZFP42,ZFP57,Zfx,ZIC1,Zic1::Zic2,Zic2,Zic3,ZIC4,ZIC5,ZIM3,ZKSCAN1,ZKSCAN3,ZKSCAN5,ZNF135,ZNF136,ZNF140,ZNF143,ZNF148,ZNF16,ZNF189,ZNF211,ZNF214,ZNF24,ZNF257,ZNF263,ZNF274,ZNF281,ZNF282,ZNF317,ZNF320,ZNF324,ZNF331,ZNF341,ZNF343,ZNF354A,ZNF354C,ZNF382,ZNF384,ZNF410,ZNF416,ZNF417,ZNF418,Znf423,ZNF449,ZNF454,ZNF460,ZNF528,ZNF530,ZNF549,ZNF574,ZNF582,ZNF610,ZNF652,ZNF667,ZNF669,ZNF675,ZNF680,ZNF682,ZNF684,ZNF692,ZNF701,ZNF707,ZNF708,ZNF740,ZNF75D,ZNF76,ZNF768,ZNF784,ZNF8,ZNF816,ZNF85,ZNF93,ZSCAN29,ZSCAN31,ZSCAN4\ labelFields TFName\ longLabel JASPAR CORE 2022 - Predicted Transcription Factor Binding Sites\ motifPwmTable hgFixed.jasparCore2022\ parent jaspar off\ priority 2\ shortLabel JASPAR 2022 TFBS\ track jaspar2022\ type bigBed 6 +\ visibility hide\ lovdLong LOVD Variants >= 50 bp bigBed 9 + LOVD: Leiden Open Variation Database Public Variants, long >= 50 bp variants 0 2 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/lovd/lovd.hg38.long.bb\ group phenDis\ longLabel LOVD: Leiden Open Variation Database Public Variants, long >= 50 bp variants\ mergeSpannedItems on\ noScoreFilter on\ parent lovdComp\ shortLabel LOVD Variants >= 50 bp\ track lovdLong\ type bigBed 9 +\ urls id="https://varcache.lovd.nl/redirect/$$"\ visibility hide\ alllowmapandsegdupregions LowMap+SegDup bigBed 3 Genome In a Bottle: lowMap+SegDup regions 1 2 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/GIAB/alllowmapandsegdupregions.bb\ longLabel Genome In a Bottle: lowMap+SegDup regions\ parent problematicGIAB on\ shortLabel LowMap+SegDup\ track alllowmapandsegdupregions\ type bigBed 3\ visibility dense\ tgpNA19675_m004_MXL m004 MXL Trio vcfPhasedTrio 1000 Genomes m004 Mexican Ancestry from Los Angeles Trio 2 2 0 0 0 127 127 127 0 0 23 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX, varRep 0 longLabel 1000 Genomes m004 Mexican Ancestry from Los Angeles Trio\ parent tgpTrios\ shortLabel m004 MXL Trio\ track tgpNA19675_m004_MXL\ type vcfPhasedTrio\ vcfChildSample NA19675|child\ vcfParentSamples NA19678|mother,NA19679|father\ visibility full\ mavedb_align_dna MaveDB DNA Align bigPsl Reference-Aligned DNA Sequences from MaveDB Experiments 3 2 0 0 0 127 127 127 0 0 0 expression 1 baseColorDefault diffBases\ baseColorUseSequence lfExtra\ bigDataUrl /gbdb/hg38/maveDB/mavedb_dna.bb\ longLabel Reference-Aligned DNA Sequences from MaveDB Experiments\ parent mavedb_align_composite\ shortLabel MaveDB DNA Align\ showDiffBasesAllScales .\ track mavedb_align_dna\ mavedb_maps MaveDB Heatmaps bigBed 12 + Variant Effect Maps from MaveDB 3 2 0 0 0 127 127 127 0 0 0\ This track provides heatmaps of multiplexed assays of variant effects (MAVE) from\ MaveDB. Each heatmap presents the results of an\ experiment where many small substitutions were tested within a gene to examine their\ functional consequences.\
\\ Heatmaps within the track display the consequence of substituting invididual amino acids within the\ genome with alternatives (alternatives are listed along the left edge of the heatmap). Score ranges\ vary among experiments, but each is presented with the highest scores in\ red, the lowest scores in\ blue, and scores at the midpoint between the two in\ silver. Higher scores correspond to a\ higher enrichment level for that variant compared to others in the experiment set.\
\ The column along the left edge of each heatmap provide single-letter amino acid codes do indicate what\ was substituted in for that piece of the experiment. = indicates a synonymous substitution, - indicates\ a deletion, and * indicates a stop codon.\
\ Cells where multiple scores were reported are marked with the score count (e.g. "2" if two scores were\ reported). Mousing over a cell in the heatmap will display the ID number of that particular substitution\ in the experiment, a MAVE-HGVS\ description of the substitution with three-letter amino acid codes, and the score (or scores, if more\ than one is present).\
\ When the display is zoomed out farther than a 200,000 base window, the display switches to a coverage plot\ of where the MaveDB heatmaps can be found.\
\\ The track controls include a filter for the URN ID of the experiment (e.g. 00000103-a-1). The display supports\ full, pack, and squish modes. In dense mode, the track switches to a dense BED-like display where items mark\ the extent of individual heatmaps and exons indicate where the heatmap includes score values.\
\\ Methods for the various experiments are described briefly on the individual details pages for each heatmap\ in the track, and in greater detail on the MaveDB page for that experiment (link available on our own\ details pages).\
\ JSON files containing data from these experiments were processed into UCSC's heatmap extension of the\ standard BED format for display here.\
\\ Direct access to the data files for these experiments can be obtained from\ MaveDB.\
\\ Rubin AF, Stone J, Bianchi AH, Capodanno BJ, Da EY, Dias M, Esposito D, Frazer J, Fu Y, Grindstaff\ SB et al.\ \ MaveDB 2024: a curated community database with over seven million variant effects from multiplexed\ functional assays.\ Genome Biol. 2025 Jan 21;26(1):13.\ PMID: 39838450; PMC: PMC11753097\
\ expression 1 bigDataUrl /gbdb/hg38/maveDB/all_mave.bb\ filterLabel.name Experiment ID\ filterText.name .*\ filterType.name regexp\ longLabel Variant Effect Maps from MaveDB\ maxWindowCoverage 200000\ parent mavedb\ priority 2\ shortLabel MaveDB Heatmaps\ style heatmap\ track mavedb_maps\ type bigBed 12 +\ visibility pack\ MaxCounts_Rev Max counts of CAGE reads (rev) bigWig Max counts of CAGE reads reverse 2 2 0 0 255 127 127 255 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/fantom5/ctssMaxCounts.rev.bw\ color 0,0,255\ dataVersion FANTOM5 reprocessed7\ longLabel Max counts of CAGE reads reverse\ parent Max_counts_multiwig\ shortLabel Max counts of CAGE reads (rev)\ subGroups category=max strand=reverse\ track MaxCounts_Rev\ type bigWig\ mitoMapDiseaseMuts MITOMAP Disease Muts bigBed 9 + 14 MITOMAP Disease Mutations 0 2 0 0 0 127 127 127 0 0 2 chrM,chrMT, https://www.mitomap.org/foswiki/bin/view/MITOMAP/$<_mutsCodingOrRNA> phenDis 1 bigDataUrl /gbdb/hg38/bbi/mitoMapDiseaseMuts.bb\ exonNumbers off\ group phenDis\ longLabel MITOMAP Disease Mutations\ mouseOverField _mouseOver\ parent mitoMap on\ priority 2\ shortLabel MITOMAP Disease Muts\ track mitoMapDiseaseMuts\ type bigBed 9 + 14\ url https://www.mitomap.org/foswiki/bin/view/MITOMAP/$<_mutsCodingOrRNA>\ urlLabel MITOMAP link\ mpraVarDb MPRAVarDB bigBed 9 + 13 MPRAs: MPRAVarDB - MPRA-tested Regulatory Variant Effects 1 2 0 0 0 127 127 127 0 0 0\ The MPRAVarDB track shows 239,028 variants successfully mapped to hg38\ (from 242,818 total) across 18 MPRA studies compiled in the MPRAVarDB database\ (Jin et al., 2024).\ Each variant was experimentally tested in an MPRA experiment to evaluate whether it\ affects regulatory activity. The database covers over 30 cell lines and 30 human\ diseases and traits, including neurodegenerative diseases, immune disorders,\ melanoma, multiple myeloma, and autoimmune diseases.\
\\ Note on cell lines: The cell line shown for each variant is the reporter\ cell line in which the human regulatory element was assayed. Several studies\ used mouse cell lines (e.g. Neuro-2a, N2A, NIH/3T3, MIN6) as reporter systems\ for human sequences; these variants retain human (hg38) coordinates.\
\ \\ Note on study type: Not all studies measure transcriptional regulation\ in the same sense. Two of the larger contributors,\ Griesemer\ et al., 2021 (72,546 variants) and\ Schuster\ et al., 2023 (26,546 variants), test 3'UTR variants placed downstream\ of the reporter, where the log2 fold change between alleles reflects changes\ in mRNA stability, decay, RBP or miRNA binding, or translation efficiency\ rather than transcriptional activation. The remaining studies test 5'\ regulatory elements (promoters and enhancers) where log2FC reflects changes\ in transcription. Together, the 3'UTR studies account for 99,092 of the\ 239,028 variants in the track (~41%).\
\ \\ Items are colored by statistical significance:\
\ Each item shows the variant name (rsID when available, otherwise chr:pos:ref>alt),\ the reference and alternate alleles, the associated disease or trait, cell line,\ log2 fold change, p-value, and FDR.\
\ \\ Cell-type specificity: MPRA results are typically cell-type-specific,\ and significance in one cell line does not imply activity in another. For\ example, Tewhey\ et al., 2016 found only modest correlation (R ≈ 0.63)\ between LCL and HepG2 measurements of the same eQTL variants, and\ McAfee\ et al., 2023 reported that only 205 of 1,004 HEK293-positive\ variants overlapped HNP-positive variants. The cell line filter can be used\ to narrow results to a relevant context.\
\ \\ Note on Kircher et al., 2019:\ This study\ contributes 44,647 variants (~19% of the track) using a saturation mutagenesis\ design that tests nearly every possible nucleotide substitution at each\ position of 20 disease-associated regulatory elements at single-base-pair\ resolution: 10 promoters (TERT, LDLR, HBB, HBG1, HNF4A, MSMB, PKLR, F9,\ FOXE1, GP1BB) and 10 enhancers (SORT1, ZRS, BCL11A, IRF4, IRF6, MYC tested\ with two distinct enhancers, RET, TCF7L2, and the UC88 ultraconserved\ enhancer). Regions over those elements show many densely-packed Kircher\ variants that may dominate visualization at those loci.\
\ \\ The log2 fold change is computed as\ log2(alt RNA/DNA) − log2(ref RNA/DNA).\ A positive value means the alternate allele drove more reporter activity than\ the reference allele in this assay; a negative value means the reverse. The\ linear allelic ratio is approximately 2log2FC: log2FC = 0.5\ corresponds to roughly 1.41× allelic difference, log2FC = 1.0\ to 2×, and log2FC = 2.0 to 4×. As noted in the\ Description section, log2FC reflects transcriptional activation for\ 5'-regulatory studies and steady-state mRNA abundance, decay, or translation\ efficiency for 3'UTR studies (Griesemer et al., 2021; Schuster\ et al., 2023).\
\ \\ The following table lists the 18 MPRA studies included in MPRAVarDB, with the number of\ tested variants, diseases/traits, cell lines, and a brief description of the variant selection.\
\ \| Study | \Variants | \Disease/Trait | \Cell Line(s) | \Description | \
|---|---|---|---|---|
| Griesemer et al., 2021 | \72,546 | \NHGRI-EBI GWAS catalog | \GM12878, HEK293FT, HMEC, HepG2, K562, SKNSH | \3'UTR SNPs and indels in LD with GWAS catalog variants, variants under positive selection, and rare outlier expression variants from GTEx | \
| Kircher et al., 2019 | \44,647 | \Various (18 diseases including diabetes, cancer, blood disorders, limb malformations) | \HEK293T, HEL92.1.7, HaCaT, HeLa, HepG2, K562, LNCaP, MIN6, NIH/3T3, Neuro-2a, SK-MEL-28, SF7996 | \Saturation mutagenesis of 20 disease-associated regulatory elements at single base-pair resolution | \
| Abell et al., 2022 | \29,564 | \eQTL (no specific disease) | \GM12878 | \30,893 variants in LD with independent, common, top-ranked eQTL across 744 eGenes in the CEU cohort | \
| Tewhey et al., 2016 | \23,430 | \eQTL (no specific disease) | \GM12878 | \32,373 variants associated with eQTLs in lymphoblastoid cell lines | \
| Schuster et al., 2023 | \26,546 | \Prostate cancer | \PC3 | \14,497 single-nucleotide mutations enriched in oncogenic pathways and 3'UTR regulatory elements | \
| Mouri et al., 2022 | \14,549 | \Autoimmune diseases (Crohn's, IBD, psoriasis, MS, RA, T1D, ulcerative colitis) | \Jurkat | \GWAS variants from autoimmune disease loci tested for regulatory element activity in T cells | \
| McAfee et al., 2023 | \10,302 | \Schizophrenia | \HEK293s, HNPS | \5,173 fine-mapped schizophrenia GWAS variants | \
| Cooper et al., 2022 | \5,330 | \Alzheimer's disease, Progressive supranuclear palsy | \HEK293T | \5,706 noncoding SNVs from 25 AD and 9 PSP genome-wide significant loci | \
| Long et al., 2022 | \3,980 | \Melanoma | \C283T, UACC903 | \1,992 risk-associated variants in tight LD (r2>0.8) from 54 melanoma risk loci | \
| Myint et al., 2020 | \2,158 | \Schizophrenia, Alzheimer's disease | \K562, SH-SY5Y | \1,049 SZ and 30 AD variants in 64 SZ loci and 9 AD loci | \
| Choi et al., 2020 | \1,664 | \Melanoma | \HEK293FT, UACC903 | \GWAS melanoma risk variants | \
| Ajore et al., 2022 | \1,582 | \Multiple myeloma | \L363, MOLP8 | \1,039 variants in high LD (r2>0.8) at 23 MM risk loci | \
| Klein et al., 2019 | \1,119 | \Osteoarthritis | \Saos-2 | \1,605 SNPs in high LD (r2>0.8) at 35 lead SNPs associated with OA via GWAS | \
| Lu et al., 2021 | \1,036 | \Systemic lupus erythematosus | \GM12878, Jurkat | \18,312 variants in tight LD (r2>0.8) with 578 GWAS index variants at 531 loci | \
| Mulvey & Dougherty, 2021 | \275 | \Major depressive disorder | \N2A | \Over 1,000 SNPs from 39 neuropsychiatric GWAS loci, selected by overlap with eQTL and histone marks | \
| Ferraro et al., 2020 | \150 | \Rare variant expression (no specific disease) | \GM12878 | \Rare variants contributing to extreme expression, allelic expression, and splicing across 49 GTEx tissues | \
| Rao et al., 2021 | \88 | \Alcohol use disorder | \BLA, CE, NAC, SFC | \SNPs in 3'UTR of 88 genes from allele-specific expression analysis (30 AUD subjects vs 30 controls) | \
| Ulirsch et al., 2016 | \62 | \Red blood cell traits | \K562, K562+GATA1 | \2,756 variants in strong LD with 75 sentinel variants associated with RBC traits | \
\ Variant counts above are from the source publications (pre-liftOver totals).\ Of 242,818 total source variants, 239,028 lifted successfully to hg38; see Methods.\
\ \\
Data was downloaded from the\
MPRAVarDB web server.\
Variants originally mapped to hg19 (213,689 of 242,818) were lifted to hg38\
using liftOver. 114 variants could not be mapped and were excluded.\
The remaining variants were merged with the 29,129 natively hg38-mapped variants\
to produce a total of 239,028 hg38 records.\
\ Significance thresholds across studies: The source studies in MPRAVarDB\ do not all use the same significance framework. Most studies apply a\ Benjamini-Hochberg FDR threshold (commonly 0.05 or 0.10), but some report only\ nominal regression p-values. For example,\ Tewhey\ et al., 2016 uses BH FDR < 0.05 to call "emVars",\ Griesemer\ et al., 2021 and\ McAfee\ et al., 2023 use BH FDR < 0.10, and\ Kircher\ et al., 2019 reports raw regression p-values rather than FDR. The\ track applies a uniform FDR < 0.05 / nominal\ p < 0.05 color cutoff for visual consistency, which is the\ more conservative of the FDR thresholds reported by the source studies. For\ any variant of interest, consult the source publication for the original\ significance call.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator\ and exported from there to spreadsheet or tab-sep tables.\ From scripts, the data can be accessed through our\ API, track=mpraVarDb.\
\\ For automated download and analysis, the genome annotation is stored in a bigBed\ file that can be downloaded from\ our download server.\ The file for this track is called mpravardb.bb. Individual\ regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a\ precompiled binary for your system. Instructions for downloading source code and\ binaries can be found\ here.\ The tool can also be used to obtain features within a given range, e.g.\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/mpra/mpravardb/mpravardb.bb -chrom=chr21 -start=0 -end=100000000 stdout\
\\ The original annotation source data can be downloaded from the\ MPRAVarDB web server.\
\ \\ Thanks to Weijia Jin and colleagues at the University of Florida for creating\ and maintaining the MPRAVarDB database.\
\ \\ Abell NS, DeGorter MK, Gloudemans MJ, Greenwald E, Smith KS, He Z, Montgomery SB.\ \ Multiple causal variants underlie genetic associations in humans.\ Science. 2022 Mar 18;375(6586):1247-1254.\ PMID: 35298243; PMC: PMC9725108\
\ \\ Ajore R, Niroula A, Pertesi M, Cafaro C, Thodberg M, Went M, Bao EL, Duran-Lozano L, Lopez de\ Lapuente Portilla A, Olafsdottir T et al.\ \ Functional dissection of inherited non-coding variation influencing multiple myeloma risk.\ Nat Commun. 2022 Jan 10;13(1):151.\ PMID: 35013207; PMC: PMC8748989\
\ \\ Choi J, Zhang T, Vu A, Ablain J, Makowski MM, Colli LM, Xu M, Hennessey RC, Yin J, Rothschild H\ et al.\ \ Massively parallel reporter assays of melanoma risk variants identify MX2 as a gene promoting\ melanoma.\ Nat Commun. 2020 Jun 1;11(1):2718.\ PMID: 32483191; PMC: PMC7264232\
\ \\ Cooper YA, Teyssier N, Dräger NM, Guo Q, Davis JE, Sattler SM, Yang Z, Patel A, Wu S, Kosuri S\ et al.\ \ Functional regulatory variants implicate distinct transcriptional networks in dementia.\ Science. 2022 Aug 19;377(6608):eabi8654.\ PMID: 35981026\
\ \\ Ferraro NM, Strober BJ, Einson J, Abell NS, Aguet F, Barbeira AN, Brandt M, Bucan M, Castel SE,\ Davis JR et al.\ \ Transcriptomic signatures across human tissues identify functional rare genetic variation.\ Science. 2020 Sep 11;369(6509).\ PMID: 32913073; PMC: PMC7646251\
\ \\ Griesemer D, Xue JR, Reilly SK, Ulirsch JC, Kukreja K, Davis JR, Kanai M, Yang DK, Butts JC, Guney\ MH et al.\ \ Genome-wide functional screen of 3'UTR variants uncovers causal variants for human disease and\ evolution.\ Cell. 2021 Sep 30;184(20):5247-5260.e19.\ PMID: 34534445; PMC: PMC8487971\
\ \\ Jin W, Xia Y, Nizomov J, Liu Y, Li Z, Lu Q, Chen L.\ \ MPRAVarDB: an online database and web server for exploring regulatory effects of genetic variants.\ Bioinformatics. 2024 Oct 1;40(10).\ PMID: 39325859; PMC: PMC11464417\
\ \\ Kircher M, Xiong C, Martin B, Schubach M, Inoue F, Bell RJA, Costello JF, Shendure J, Ahituv N.\ \ Saturation mutagenesis of twenty disease-associated regulatory elements at single base-pair\ resolution.\ Nat Commun. 2019 Aug 8;10(1):3583.\ PMID: 31395865; PMC: PMC6687891\
\ \\ Klein JC, Keith A, Rice SJ, Shepherd C, Agarwal V, Loughlin J, Shendure J.\ \ Functional testing of thousands of osteoarthritis-associated variants for regulatory activity.\ Nat Commun. 2019 Jun 4;10(1):2434.\ PMID: 31164647; PMC: PMC6547687\
\ \\ Long E, Yin J, Funderburk KM, Xu M, Feng J, Kane A, Zhang T, Myers T, Golden A, Thakur R et\ al.\ \ Massively parallel reporter assays and variant scoring identified functional variants and target\ genes for melanoma loci and highlighted cell-type specificity.\ Am J Hum Genet. 2022 Dec 1;109(12):2210-2229.\ PMID: 36423637; PMC: PMC9748337\
\ \\ Lu X, Chen X, Forney C, Donmez O, Miller D, Parameswaran S, Hong T, Huang Y, Pujato M, Cazares T\ et al.\ \ Global discovery of lupus genetic risk variant allelic enhancer activity.\ Nat Commun. 2021 Mar 12;12(1):1611.\ PMID: 33712590; PMC: PMC7955039\
\ \\ McAfee JC, Lee S, Lee J, Bell JL, Krupa O, Davis J, Insigne K, Bond ML, Zhao N, Boyle AP et\ al.\ \ Systematic investigation of allelic regulatory activity of schizophrenia-associated common\ variants.\ Cell Genom. 2023 Oct 11;3(10):100404.\ PMID: 37868037; PMC: PMC10589626\
\ \\ Mouri K, Guo MH, de Boer CG, Lissner MM, Harten IA, Newby GA, DeBerg HA, Platt WF, Gentili M, Liu DR\ et al.\ \ Prioritization of autoimmune disease-associated genetic variants that perturb regulatory element\ activity in T cells.\ Nat Genet. 2022 May;54(5):603-612.\ PMID: 35513721; PMC: PMC9793778\
\ \\ Mulvey B, Dougherty JD.\ \ Transcriptional-regulatory convergence across functional MDD risk variants identified by massively\ parallel reporter assays.\ Transl Psychiatry. 2021 Jul 22;11(1):403.\ PMID: 34294677; PMC: PMC8298436\
\ \\ Myint L, Wang R, Boukas L, Hansen KD, Goff LA, Avramopoulos D.\ \ A screen of 1,049 schizophrenia and 30 Alzheimer's-associated variants for regulatory\ potential.\ Am J Med Genet B Neuropsychiatr Genet. 2020 Jan;183(1):61-73.\ PMID: 31503409; PMC: PMC7233147\
\ \\ Rao X, Thapa KS, Chen AB, Lin H, Gao H, Reiter JL, Hargreaves KA, Ipe J, Lai D, Xuei X et\ al.\ \ Allele-specific expression and high-throughput reporter assay reveal functional genetic variants\ associated with alcohol use disorders.\ Mol Psychiatry. 2021 Apr;26(4):1142-1151.\ PMID: 31477794; PMC: PMC7050407\
\ \\ Schuster SL, Arora S, Wladyka CL, Itagi P, Corey L, Young D, Stackhouse BL, Kollath L, Wu QV, Corey\ E et al.\ \ Multi-level functional genomics reveals molecular and cellular oncogenicity of patient-based\ 3'-untranslated region mutations.\ Cell Rep. 2023 Aug 29;42(8):112840.\ PMID: 37516102; PMC: PMC10540565\
\ \\ Tewhey R, Kotliar D, Park DS, Liu B, Winnicki S, Reilly SK, Andersen KG, Mikkelsen TS, Lander ES,\ Schaffner SF et al.\ \ Direct Identification of Hundreds of Expression-Modulating Variants using a Multiplexed Reporter\ Assay.\ Cell. 2016 Jun 2;165(6):1519-1529.\ PMID: 27259153; PMC: PMC4957403\
\ \\ Ulirsch JC, Nandakumar SK, Wang L, Giani FC, Zhang X, Rogov P, Melnikov A, McDonel P, Do R,\ Mikkelsen TS et al.\ \ Systematic Functional Dissection of Common Genetic Variation Affecting Red Blood Cell Traits.\ Cell. 2016 Jun 2;165(6):1530-1545.\ PMID: 27259154; PMC: PMC4893171\
\ regulation 1 bigDataUrl /gbdb/hg38/mpra/mpravardb/mpravardb.bb\ dataVersion MPRAVarDB snapshot 2026-03-10\ defaultLabelFields name\ filter.fdr 0:1\ filter.log2FC -5:5\ filterByRange.fdr on\ filterByRange.log2FC on\ filterLabel.fdr Filter by false discovery rate\ filterLabel.log2FC Filter by log2 fold change (alt vs ref)\ filterLimits.fdr 0:1\ filterLimits.log2FC -5:5\ filterValues.cellLine GM12878,PC3 cell,HepG2,K562,Jurkat,HEK293FT,HEK293T,SKNSH,HMEC,HNPS,HEK293s,HEL92.1.7,Neuro-2a,MIN6,NIH/3T3,HEK293T,,SF7996,N2A,SH-SY5Y,HaCaT,HeLa,LNCaP,SK-MEL-28,Saos-2,MOLP8,L363,C283T,UACC903,K562+GATA1,BLA,CE,NAC,SFC\ itemRgb on\ labelFields name,cellLine,disease\ longLabel MPRAs: MPRAVarDB - MPRA-tested Regulatory Variant Effects\ maxWindowToDraw 10000000\ mouseOver Variant: $name\ The NCBI RefSeq Genes composite track shows human protein-coding and non-protein-coding\ genes taken from the NCBI RNA reference sequences collection (RefSeq). All subtracks use\ coordinates provided by RefSeq, except for the UCSC RefSeq track, which UCSC produces by\ realigning the RefSeq RNAs to the genome. This realignment may result in occasional differences\ between the annotation coordinates provided by UCSC and NCBI. For RNA-seq analysis, we advise\ using NCBI aligned tables like RefSeq All or RefSeq Curated. See the \ Methods section for more details about how the different tracks were \ created.
\\ Please visit NCBI's Feedback for Gene and Reference Sequences (RefSeq) page to make suggestions, \ submit additions and corrections, or ask for help concerning RefSeq records.
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ This track is a composite track that contains differing data sets.\ To show only a selected set of subtracks, uncheck the boxes next to the tracks that you wish to \ hide. Note: Not all subtracks are available on all assemblies.
\ \ The possible subtracks include:\\ The RefSeq All, RefSeq Curated, RefSeq Predicted, RefSeq HGMD,\ RefSeq Select/MANE and UCSC RefSeq tracks follow the display conventions for\ gene prediction tracks.\ The color shading indicates the level of review the RefSeq record has undergone:\ predicted (light), provisional (medium), or reviewed (dark), as defined by RefSeq.
\ \\
| Color | \Level of review | \
|---|---|
| \ | Reviewed: the RefSeq record has been reviewed by NCBI staff or by a collaborator. The NCBI review process includes assessing available sequence data and the literature. Some RefSeq records may incorporate expanded sequence and annotation information. | \
| \ | Provisional: the RefSeq record has not yet been subject to individual review. The initial sequence-to-gene association has been established by outside collaborators or NCBI staff. | \
| \ | Predicted: the RefSeq record has not yet been subject to individual review, and some aspect of the RefSeq record is predicted. | \
\ The item labels and codon display properties for features within this track can be configured \ through the check-box controls at the top of the track description page. To adjust the settings \ for an individual subtrack, click the wrench icon next to the track name in the subtrack list.
\The RefSeq Diffs track contains five different types of inconsistency between the\ reference genome sequence and the RefSeq transcript sequences. The five types of differences are\ as follows:\
\ When reporting HGVS with RefSeq sequences, to make sure that results from\ research articles can be mapped to the genome unambiguously, \ please specify the RefSeq annotation release displayed on the transcript's\ Genome Browser details page and also the RefSeq transcript ID with version\ (e.g. NM_012309.4 not NM_012309). \
\ \ \ \\ Tracks contained in the RefSeq annotation and RefSeq RNA alignment tracks were created at UCSC using \ data from the NCBI RefSeq project. Data files were downloaded from RefSeq in GFF file format and \ converted to the genePred and PSL table formats for display in the Genome Browser. Information about\ the NCBI annotation pipeline can be found \ here.
\ \The RefSeq Diffs track is generated by UCSC using NCBI's RefSeq RNA alignments.
\\ The UCSC RefSeq Genes track is constructed using the same methods as previous RefSeq Genes tracks.\ RefSeq RNAs were aligned against the human genome using BLAT. Those with an alignment of\ less than 15% were discarded. When a single RNA aligned in multiple places, the alignment\ having the highest base identity was identified. Only alignments having a base identity\ level within 0.1% of the best and at least 96% base identity with the genomic sequence were\ kept.
\\ The NCBI Orthologs track was generated using the latest\ NCBI files (gene2accession and\ gene_orthologs). NCBI chromosome identifiers were mapped to UCSC-compatible IDs using\ species-specific chromosome alias files, and genes were filtered to include only those located on\ valid NCBI chromosomes. A custom Python script processed the ortholog relationships and created bed files for\ each species. The bed files were then converted to BigBed format, with indexing for search\ functionality. The procedure is documented in the makeDoc from our GitHub repository.
\ \\ The raw data for these tracks can be accessed in multiple ways. It can be explored interactively \ using the REST API,\ Table Browser or\ Data Integrator. The tables can also be accessed programmatically through our\ public MySQL server or downloaded from our\ downloads server for local processing. The previous track versions are available\ in the archives of our downloads server. You can also access any RefSeq table\ entries in JSON format through our \ JSON API.
\\ The data in the RefSeq Other, RefSeq Diffs, and NCBI Orthologs tracks are organized in\ bigBed file format; more\ information about accessing the information in this bigBed file can be found\ below. The other subtracks are associated with database tables as follows:
\\ The first column of each of these tables is "bin". This column is designed\ to speed up access for display in the Genome Browser, but can be safely ignored in downstream\ analysis. You can read more about the bin indexing system\ here.
\\ The annotations in the RefSeqOther, RefSeqDiffs, and NCBI Orthologs tracks are stored in bigBed\ files, which can be obtained from our downloads server here,\ ncbiRefSeqOther.bb,\ ncbiRefSeqDiffs.bb, and\ ncbiOrtho.bb.\ Individual regions or the whole set of genome-wide annotations can be obtained using our tool\ bigBedToBed which can be compiled from the source code or downloaded as a precompiled\ binary for your system from the utilities directory linked below. For example, to extract only\ annotations in a given region, you could use the following command:
\\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/ncbiRefSeq/ncbiRefSeqOther.bb\ -chrom=chr16 -start=34990190 -end=36727467 stdout
\\ You can download a GTF format version of the RefSeq All table from the \ GTF downloads directory.\ The genePred format tracks can also be converted to GTF format using the\ genePredToGtf utility, available from the\ utilities directory on the UCSC downloads \ server. The utility can be run from the command line like so:
\ genePredToGtf hg38 ncbiRefSeqPredicted ncbiRefSeqPredicted.gtf\\ Note that using genePredToGtf in this manner accesses our public MySQL server, and you therefore \ must set up your hg.conf as described on the MySQL page linked near the beginning of the Data Access\ section.
\\ A file containing the RNA sequences in FASTA format for all items in the RefSeq All, RefSeq Curated, \ and RefSeq Predicted tracks can be found on our downloads server\ here.
\\ Please refer to our mailing list archives for questions.
\ \\ Previous versions of the ncbiRefSeq set of tracks can be found on our archive download server.\
\ \\ This track was produced at UCSC from data generated by scientists worldwide and curated by the\ NCBI RefSeq project.
\ \\ Kent WJ.\ BLAT - the BLAST-like \ alignment tool. Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518
\\ Pruitt KD, Brown GR, Hiatt SM, Thibaud-Nissen F, Astashyn A, Ermolaeva O, Farrell CM, Hart J,\ Landrum MJ, McGarvey KM et al.\ RefSeq: an update on mammalian reference sequences.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D756-63.\ PMID: 24259432; PMC: \ PMC3965018
\\ Pruitt KD, Tatusova T, Maglott DR.\ \ NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts \ and proteins.\ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4.\ PMID: 15608248; PMC: PMC539979
\ genes 1 allButtonPair on\ compositeTrack on\ dataVersion /gbdb/$D/ncbiRefSeq/ncbiRefSeqVersion.txt\ dbPrefixLabels hg="HGNC" dm="FlyBase" ce="WormBase" rn="RGD" sacCer="SGD" danRer="ZFIN" mm="MGI" xenTro="XenBase"\ dbPrefixUrls hg="http://www.genenames.org/cgi-bin/gene_symbol_report?hgnc_id=$$" dm="https://flybase.org/reports/$$" ce="http://www.wormbase.org/db/gene/gene?name=$$" rn="https://rgd.mcw.edu/rgdweb/search/search.html?term=$$" sacCer="https://www.yeastgenome.org/locus/$$" danRer="https://zfin.org/$$" mm="https://www.informatics.jax.org//marker/$$" xenTro="https://www.xenbase.org/gene/showgene.do?method=display&geneId=$$"\ dragAndDrop subTracks\ group genes\ longLabel RefSeq genes from NCBI\ noInherit on\ priority 2\ shortLabel NCBI RefSeq\ track refSeqComposite\ type genePred\ visibility dense\ chainHg19ReMap NCBI ReMap hg19 chain hg19 NCBI ReMap alignments to hg19/GRCh37 0 2 0 0 0 127 127 127 0 0 0 map 1 chainLinearGap medium\ chainMinScore 3000\ longLabel NCBI ReMap alignments to hg19/GRCh37\ matrix 16 91,-114,-31,-123,-114,100,-125,-31,-31,-125,100,-114,-123,-31,-114,91\ matrixHeader A, C, G, T\ otherDb hg19\ parent liftHg19\ priority 2\ shortLabel NCBI ReMap hg19\ track chainHg19ReMap\ type chain hg19\ nmdDetectiveA NMDetective-A bigWig NMDetective-A: Random forest prediction of NMD efficiency (Lindeboom 2016) 0 2 0 128 255 127 191 255 0 0 0\ The NMDetective tracks display genome-wide predictions of nonsense-mediated mRNA\ decay (NMD) efficiency from\ Lindeboom et al. 2016.\ NMDetective scores predict whether a premature termination codon (PTC) at a given position\ will trigger NMD and mRNA degradation, or whether the transcript will escape NMD and\ potentially produce a truncated protein.\
\ \\ Scores range from approximately −1 to +1. Positive values indicate that a PTC at\ that position is predicted to trigger NMD (the mRNA is degraded). Negative values indicate\ that the PTC is predicted to escape NMD (the truncated mRNA may be translated into an\ aberrant protein). Values near zero indicate intermediate or uncertain NMD efficiency.\
\ \| Track | Description |
|---|---|
| NMDetective-A | \Random forest model predicting NMD efficiency for all possible PTCs introduced\ by single-nucleotide variants. Explains ~71% of systematic variance in NMD\ efficiency. |
| NMDetective-B | \Simplified decision tree model for all possible PTCs. Slightly lower accuracy\ (~68% variance explained) but more interpretable, making it suitable for\ clinical applications. |
| NMDetective-A PTC | \Random forest model predicting NMD efficiency specifically for the first\ out-of-frame PTC introduced by frameshifting indel mutations. |
| NMDetective-B PTC | \Decision tree model for the first out-of-frame PTC from frameshifting\ indels. |
\ Each subtrack is displayed as a signal (bigWig) track. By default, the vertical axis\ ranges from −1 to +1. Regions with positive values (predicted NMD-triggering) are\ shown above the baseline; regions with negative values (predicted NMD escape) are shown\ below.\
\\ The NMDetective models were trained on somatic nonsense mutation data from 9,769 cancer\ patients and validated with frameshift mutations and germline variants\ (Lindeboom et al. 2019).\ The models incorporate the following features to predict NMD efficiency:\
\\ NMDetective-A (random forest regression) captures non-linear interactions among\ these features and achieves the highest predictive accuracy.\ NMDetective-B (decision tree) applies a simpler rule-based classification that\ is more transparent, with a modest reduction in accuracy.\
\ \\ The predictions were generated for every possible PTC-introducing single-nucleotide\ variant and for the first out-of-frame PTC from every possible single-nucleotide\ frameshifting indel across all human protein-coding transcripts. The original bedGraph\ custom track files were downloaded from the\ NMDetective Figshare page\ resource and converted to bigWig format at UCSC.\
\ \\ The data underlying these tracks can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Rik Lindeboom for providing custom tracks and the original NMDetective data\ on Figshare.\
\ \\ Lindeboom RG, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Lindeboom RGH, Vermeulen M, Lehner B, Supek F.\ \ The impact of nonsense-mediated mRNA decay on genetic disease, gene editing and cancer\ immunotherapy.\ Nat Genet. 2019 Nov;51(11):1645-1651.\ PMID: 31659324; PMC: PMC6858879\
\ \ genes 0 autoScale off\ bigDataUrl /gbdb/hg38/nmd/NMDetectiveA.bw\ color 0,128,255\ html nmdDetective\ longLabel NMDetective-A: Random forest prediction of NMD efficiency (Lindeboom 2016)\ maxHeightPixels 128:32:8\ parent nmd off\ priority 2\ shortLabel NMDetective-A\ track nmdDetectiveA\ type bigWig\ viewLimits -1:1\ visibility hide\ omimGene2 OMIM Genes bed 4 OMIM Gene Phenotypes - Dark Green Can Be Disease-causing 1 2 0 80 0 127 167 127 0 0 0 http://www.omim.org/entry/NOTE:
\
OMIM is intended for use primarily by physicians and other\
professionals concerned with genetic disorders, by genetics researchers, and\
by advanced students in science and medicine. While the OMIM database is\
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions. Further, please be\
sure to click through to omim.org for the very latest, as they are continually \
updating data.
NOTE ABOUT DOWNLOADS:
\
OMIM is the property \
of Johns Hopkins University and is not available for download or mirroring \
by any third party without their permission. Please see \
OMIM\
for downloads.
OMIM is a compendium of human genes and genetic phenotypes. The full-text,\ referenced overviews in OMIM contain information on all known Mendelian\ disorders and over 12,000 genes. OMIM is authored and edited at the\ McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University\ School of Medicine, under the direction of Dr. Ada Hamosh. This database\ was initiated in the early 1960s by Dr. Victor A. McKusick as a catalog\ of Mendelian traits and disorders, entitled Mendelian Inheritance\ in Man (MIM).\
\ \\ The OMIM data are separated into three separate tracks:\
\ \OMIM Alleles \
Variants in the OMIM database that have associated \
dbSNP identifiers. This track is currently unavailable on the hg38 assembly,\
as it depends on dbSNP data that has not been released yet.\
\
OMIM Genes\
The genomic positions of gene entries in the OMIM \
database. The coloring indicates the associated OMIM phenotype map key.\
OMIM Phenotypes - Gene Unknown\
Regions known to be associated with a phenotype, \
but for which no specific gene is known to be causative. This track \
also includes known multi-gene syndromes.\
\ This track shows the genomic positions of all gene entries in the Online Mendelian\ Inheritance in Man (OMIM) database.\
\ \Genomic locations of OMIM gene entries are displayed as solid blocks. The entries are colored\ according to the associated OMIM phenotype map key (if any):\
Gene symbol and disease information, when available, are displayed on the details page for an\ item, and links to related RefSeq Genes and UCSC Genes are given.\
\The descriptions of the OMIM entries are shown on the main browser display when Full display\ mode is chosen. In Pack mode, the descriptions are shown when mousing over each entry. Items\ displayed can be filtered according to phenotype map key on the track controls page. \
\ \\ The mappings displayed in this track are based on OMIM gene entries, their Entrez Gene IDs, and\ the corresponding RefSeq Gene locations:\
\ Because OMIM has only allowed Data queries within individual chromosomes, no download files are\ available from the Genome Browser. Full genome datasets can be downloaded directly from the\ OMIM Downloads page.\ All genome-wide downloads are freely available from OMIM after registration.
\\ If you need the OMIM data in exactly the format of the UCSC Genome Browser,\ for example if you are running a UCSC Genome Browser local installation (a partial "mirror"),\ please create a user account on omim.org and contact OMIM via\ https://omim.org/contact. Send them your OMIM\ account name and request access to the UCSC Genome Browser "entitlement". They will\ then grant you access to a MySQL/MariaDB data dump that contains all UCSC\ Genome Browser OMIM tables.
\\ UCSC offers queries within chromosomes from\ Table Browser that include a variety\ of filtering options and cross-referencing other datasets using our\ Data Integrator tool.\ UCSC also has an API\ that can be used to retrieve data in JSON format from a particular chromosome range.
\\ Please refer to our searchable\ mailing list archives\ for more questions and example queries, or our\ Data Access FAQ\ for more information.
\ \chr1\ 11106534\ 11262551\ MTOR\ 601231,\ Smith-Kingsmore syndrome,Focal cortical dysplasia, type II, somatic,\ 3,\ Autosomal dominant
For a quick link to pre-fill these options, click \ \ this session link.\ \
\ \\ Thanks to OMIM and NCBI for the use of their data. This track was\ constructed by Fan Hsu, Robert Kuhn, and Brooke Rhead of the UCSC Genome Bioinformatics Group.
\ \Amberger J, Bocchini CA, Scott AF, Hamosh A. \ McKusick's Online Mendelian Inheritance in Man (OMIM®). \ Nucleic Acids Res. 2009 Jan;37(Database issue):D793-6. Epub 2008 Oct 8.\
\\ Hamosh A, Scott AF, Amberger JS, Bocchini CA, McKusick VA. \ Online Mendelian Inheritance in Man (OMIM), a knowledgebase of \ human genes and genetic disorders. \ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D514-7.\
\ phenDis 1 color 0, 80, 0\ hgsid on\ longLabel OMIM Gene Phenotypes - Dark Green Can Be Disease-causing\ noGenomeReason Distribution restrictions by OMIM. See the track documentation for details. You can download the complete OMIM dataset for free from omim.org\ parent omimContainer\ priority 2\ shortLabel OMIM Genes\ tableBrowser noGenome omimGeneMap omimGeneMap2 omimPhenotype omimGeneSymbol omim2gene\ track omimGene2\ type bed 4\ url http://www.omim.org/entry/\ visibility dense\ panelAppCNVs PanelApp GE CNVs bigBed 9 + Genomics England PanelApp CNV Regions 3 2 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/panelApp/cnv.bb\ filter.versionCreated 1\ filterLabel.versionCreated Minimum panel version to display\ filterValues.confidenceLevel 3,2,1,0\ itemRgb on\ labelFields entityName\ longLabel Genomics England PanelApp CNV Regions\ mouseOver Gene: $entityName\ These tracks contain pseudogene predictions and their parents as identified by PseudoPipe.\ PseudoPipe is a homology-based\ computational pipeline that can search a mammalian genome and identify pseudogene sequences\ comprehensively and consistently.\
\\ Pseudogenes are genomic sequences that bear similarity to specific protein-coding genes, but are\ unable to produce functional proteins due to the existence of frameshifts, premature stop codons, or\ other deleterious mutations. They arise from gene duplication or retrotransposition events and are\ important resources in understanding the evolutionary history of genes and genomes.
\ \This composite track consists of two subtracks: the Pseudogenes track and the Pseudogene\ Parents track.
\\ The Pseudogene Parents track displays parent genes and pseudogenes\ labeled with their HUGO\ IDs, which were derived from Ensembl gene IDs provided by the Gerstein lab after dataset creation. It includes indicators for pseudogenes. \ These indicators do not show pseudogene locations directly but instead indicate how many pseudogenes\ are associated with each gene and link to their genomic regions in the Pseudogenes track.
\\ The Pseudogenes track shows pseudogenes labeled with their parent HUGO ID and colored\ according to pseudogene type. The authors assigned PGOHUMG IDs to genes and PGOHUMT IDs to\ transcripts. Note: Not all PseudoPipe IDs could be mapped back to their original Ensembl\ IDs. In these cases, the gene ID is listed as NA.
\ \ Pseudogene types:\Each parent gene is shown with associated pseudogenes represented as grey blocks. These blocks\ do not reflect actual pseudogene locations but rather indicate the count of pseudogenes linked to\ the gene.\
\\ If a parent gene has four grey blocks beneath it, this indicates the presence of four pseudogenes\ elsewhere in the genome. Hovering over an item displays the gene type, ID (Ensembl transcript ID\ or PseudoPipe transcript ID), and the genome position of the gene or pseudogene, with a link to\ that genomic region.\
\ \Pseudogenes are colored by type.
\\ Hovering over a pseudogene item shows the pseudogene type, parent HUGO gene symbol, and the Ensembl\ parent transcript ID, which links to the genome position of the parent gene.
\ \\ The PseudoPipe pipeline identifies pseudogenes through a series of steps. It first uses BLAST to\ rapidly cross-reference potential parent proteins against the intergenic regions of the genome. The\ resulting raw hits are then processed by removing redundancies, clustering neighboring sequences,\ and aligning each cluster with a unique parent gene. Finally, pseudogenes are classified based on a\ combination of criteria, including homology, intron-exon structure, and the presence of stop codons\ or frameshifts. This method is designed to detect pseudogenes that are unable to be translated into\ proteins.
\\ These tracks were generated using a Bash script that processes a GTF file with pseudogene\ annotations by removing duplicates, correcting overlapping exons, and converting the data to BED\ format with pseudoPipeToBed.py. This script extracts gene and transcript IDs, merges overlapping\ exons, assigns colors based on pseudogene type, and outputs a BED file with gene and parent\ annotations. PseudoPipeParents.py then links pseudogenes to their functional genes by determining\ parent gene coordinates, updating pseudogene entries with interactive browser links and generating a\ parent BED file. The final data are formatted into pseudoPipePgenes.bb and pseudoPipeParents.bb BigBed\ files. The detailed documentation (makeDoc) and \ Python scripts are available in our GitHub repository.\
\ \The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ The data may also be explored interactively using our\ REST API.
\For automated download and analysis, the genome annotation is stored at UCSC in bigBed files\ that can be downloaded from the\ download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed which can be compiled from the source code or downloaded as a precompiled\ binary for your system.
\\ Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/hg38/pseudogenes/pseudoPipePgenes.bb -chrom=chr21 -start=0 -end=10000000 stdout\ \ \Thanks to the Gerstein lab at Yale University for making this data available, and to Cristina\ Sisu for providing data in GTF format with parent annotations.
\ \\ Zhang Z, Carriero N, Zheng D, Karro J, Harrison PM, Gerstein M.\ \ PseudoPipe: an automated pseudogene identification pipeline.\ Bioinformatics. 2006 Jun 15;22(12):1437-9.\ PMID: 16574694\
\ genes 1 bigDataUrl /gbdb/hg38/pseudogenes/pseudoPipePgenes.bb\ defaultLabelFields pgenehugo\ html pseudogenes.html\ itemRgb on\ labelFields pgenehugo\ labelSeparator " "\ longLabel Yale Pseudogenes\ mouseOver Pseudogene type: ${pgeneType}\ The recombination rate track represents calculated rates of recombination based\ on the genetic maps from deCODE (Halldorsson et al., 2019) and 1000 Genomes\ (2013 Phase 3 release, lifted from hg19). The deCODE map is more recent, has a higher \ resolution and was natively created on hg38 and therefore recommended. \ For the Recomb. deCODE average track, the recombination rates for chrX represent the female rate.\
\ \This track also includes a subtrack with all the\ individual deCODE recombination events and another subtrack with several thousand\ de-novo mutations found in the deCODE sequencing data. These two tracks are hidden by\ default and have to be switched on explicitly on the configuration page.\
\ \\ This is a super track that contains different subtracks, three with the deCODE\ recombination rates (paternal, maternal and average) and one with the 1000\ Genomes recombination rate (average). These tracks are in \ signal graph\ (wiggle) format. By default, to show most recombination hotspots, their maximum\ value is set to 100 cM, even though many regions have values higher than 100.\ The maximum value can be changed on the configuration pages of the tracks.\
\ \\ There are two more tracks that show additional details provided by deCODE: one\ subtrack with the raw data of all cross-overs tagged with their proband ID and\ another one with around 8000 human de-novo mutation variants that are linked to\ cross-over changes.\
\ \\ The deCODE genetic map was created at \ deCODE Genetics. It is based \ on microarrays assaying 626,828 SNP markers that allowed to identify 1,476,140 crossovers in\ 56,321 paternal meioses and 3,055,395 crossovers in 70,086 maternal meioses.\ In total, the data is based on 4,531,535 crossovers in 126,427 meioses. By\ using WGS data with 9,305,070 SNPs, the boundaries for 761,981 crossovers were\ refined: 247,942 crossovers in 9423 paternal meioses and 514,039 crossovers in\ 11,750 maternal meioses. The average resolution of the genetic map is 682 base\ pairs (bp): 655 and 708 bp for the paternal and maternal maps, respectively.\
\ \The 1000 Genomes genetic map is based on the IMPUTE genetic map based on 1000 Genomes Phase 3, on hg19 coordinates. It\ was converted to hg38 by Po-Ru Loh at the Broad Institute. After a run of \ liftOver, he post-processed the data to deal with situations in which\ consecutive map locations became much closer/farther after lifting. The\ heuristic used is sufficient for statistical phasing but may not be optimal for\ other analyses. For this reason, and because of its higher resolution, the DeCODE\ map is therefore recommended for hg38.\
\ \As with all other tracks, the data conversion commands and pointers to the\ original data files are documented in the \ makeDoc file of this track.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr17 -start=45941345 -end=45942345 http://hgdownload.soe.ucsc.edu/gbdb/hg38/recombRate/recombAvg.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ This track was produced at UCSC using data that are freely available for\ the deCODE\ and 1000 Genomes genetic maps. Thanks to Po-Ru Loh at the\ Broad Institute for providing the code to lift the hg19 1000 Genomes map data to hg38.\
\ \\ 1000 Genomes Project Consortium., Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA,\ Hurles ME, McVean GA.\ \ A map of human genome variation from population-scale sequencing.\ Nature. 2010 Oct 28;467(7319):1061-73.\ PMID: 20981092; PMC: PMC3042601\
\ \\ Halldorsson BV, Palsson G, Stefansson OA, Jonsson H, Hardarson MT, Eggertsson HP, Gunnarsson B,\ Oddsson A, Halldorsson GH, Zink F et al.\ \ Characterizing mutagenic effects of recombination through a sequence-level genetic map.\ Science. 2019 Jan 25;363(6425).\ PMID: 30679340\
\ map 0 bigDataUrl /gbdb/hg38/recombRate/recombPat.bw\ html recombRate2.html\ longLabel Recombination rate: deCODE Genetics, paternal\ maxHeightPixels 128:60:8\ parent recombRate2\ priority 2\ shortLabel Recomb. deCODE Pat\ track recombPat\ type bigWig\ viewLimits 0.0:100\ viewLimitsMax 0:150000\ visibility full\ ncbiRefSeqCurated RefSeq Curated genePred NCBI RefSeq genes, curated subset (NM_*, NR_*, NP_* or YP_*) 1 2 12 12 120 133 133 187 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 12,12,120\ idXref ncbiRefSeqLink mrnaAcc name\ longLabel NCBI RefSeq genes, curated subset (NM_*, NR_*, NP_* or YP_*)\ parent refSeqComposite on\ priority 2\ shortLabel RefSeq Curated\ track ncbiRefSeqCurated\ ReMapTFs ReMap ChIP-seq bigBed 9 + ReMap Atlas of Regulatory Regions 4 2 0 0 0 127 127 127 0 0 0\ This track represents the ReMap Atlas of regulatory regions, which consists of a\ large-scale integrative analysis of all Public ChIP-seq data for transcriptional\ regulators from GEO, ArrayExpress, and ENCODE. \
\ \\ Below is a schematic diagram of the types of regulatory regions: \
\
\
\ This 4th release of ReMap (2022) presents the analysis of a total of 8,103 \ quality controlled ChIP-seq (n=7,895) and ChIP-exo (n=208) data sets from public\ sources (GEO, ArrayExpress, ENCODE). The ChIP-seq/exo data sets have been mapped\ to the GRCh38/hg38 human assembly. The data set is defined as a ChIP-seq \ experiment in a given series (e.g. GSE46237), for a given TF (e.g. NR2C2), in a\ particular biological condition (i.e. cell line, tissue type, disease state, or\ experimental conditions; e.g. HELA). Data sets were labeled by concatenating\ these three pieces of information, such as GSE46237.NR2C2.HELA. \ \
\Those merged analyses cover a total of 1,211 DNA-binding proteins\ (transcriptional regulators) such as a variety of transcription factors (TFs),\ transcription co-activators (TCFs), and chromatin-remodeling factors (CRFs) for\ 182 million peaks. \
\ \
\
\
\
Public ChIP-seq data sets were extracted from Gene Expression Omnibus (GEO) and\
ArrayExpress (AE) databases. For GEO, the query\
\
'('chip seq' OR 'chipseq' OR\
'chip sequencing') AND 'Genome binding/occupancy profiling by high throughput\
sequencing' AND 'homo sapiens'[organism] AND NOT 'ENCODE'[project]'\
\
was used to return a list of all potential data sets to analyze, which were then manually \
assessed for further analyses. Data sets involving polymerases (i.e. Pol2 and\
Pol3), and some mutated or fused TFs (e.g. KAP1 N/C terminal mutation, GSE27929)\
were excluded.\
\ Available ENCODE ChIP-seq data sets for transcriptional regulators from the\ ENCODE portal were processed with the\ standardized ReMap pipeline. The list of ENCODE data was retrieved as FASTQ files from the\ ENCODE portal\ using the following filters:\
\ Both Public and ENCODE data were processed similarly. Bowtie 2 (PMC3322381) (version 2.2.9) with options -end-to-end -sensitive was used to align all\ reads on the genome. Biological and technical\ replicates for each unique combination of GSE/TF/Cell type or Biological condition\ were used for peak calling. TFBS were identified using MACS2 peak-calling tool\ (PMC3120977) (version 2.1.1.2) in order to follow ENCODE ChIP-seq guidelines,\ with stringent thresholds (MACS2 default thresholds, p-value: 1e-5). An input data\ set was used when available.\
\ \ \\ To assess the quality of public data sets, a score was computed based on the\ cross-correlation and the FRiP (fraction of reads in peaks) metrics developed by\ the ENCODE Consortium (https://genome.ucsc.edu/ENCODE/qualityMetrics.html). Two\ thresholds were defined for each of the two cross-correlation ratios (NSC,\ normalized strand coefficient: 1.05 and 1.10; RSC, relative strand coefficient:\ 0.8 and 1.0). Detailed descriptions of the ENCODE quality coefficients can be\ found at https://genome.ucsc.edu/ENCODE/qualityMetrics.html. The\ phantompeak tools suite was used\ (https://code.google.com/p/phantompeakqualtools/) to compute\ RSC and NSC.\
\\ Please refer to the ReMap 2022, 2020, and 2018 publications for more details\ (citation below).\
\ \ \ \\ ReMap Atlas of regulatory regions data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ Individual BED files for specific TFs, cells/biotypes, or data sets can be\ found and downloaded on the ReMap website.\
\ \\ Chèneby J, Gheorghe M, Artufel M, Mathelier A, Ballester B.\ \ ReMap 2018: an updated atlas of regulatory regions from an integrative analysis of DNA-binding ChIP-\ seq experiments.\ Nucleic Acids Res. 2018 Jan 4;46(D1):D267-D275.\ PMID: 29126285; PMC: PMC5753247\
\\ Chèneby J, Ménétrier Z, Mestdagh M, Rosnet T, Douida A, Rhalloussi W, Bergon A, Lopez\ F, Ballester B.\ \ ReMap 2020: a database of regulatory regions from an integrative analysis of Human and Arabidopsis\ DNA-binding sequencing experiments.\ Nucleic Acids Res. 2020 Jan 8;48(D1):D180-D188.\ PMID: 31665499; PMC: PMC7145625\
\\ Griffon A, Barbier Q, Dalino J, van Helden J, Spicuglia S, Ballester B.\ \ Integrative analysis of public ChIP-seq experiments reveals a complex multi-cell regulatory\ landscape.\ Nucleic Acids Res. 2015 Feb 27;43(4):e27.\ PMID: 25477382; PMC: PMC4344487\
\\ Hammal F, de Langen P, Bergon A, Lopez F, Ballester B.\ \ ReMap 2022: a database of Human, Mouse, Drosophila and Arabidopsis regulatory regions from an\ integrative analysis of DNA-binding sequencing experiments.\ Nucleic Acids Res. 2022 Jan 7;50(D1):D316-D325.\ PMID: 34751401; PMC: PMC8728178\
\ \ regulation 1 bigDataUrl /gbdb/hg38/reMap/reMap2022.bb\ denseCoverage 100\ filterLabel.Biotypes Biotypes (cell lines, tissues...)\ filterLabel.TF Transcriptional regulators\ filterText.Biotypes *\ filterText.TF *\ filterType.Biotypes multipleListOnlyOr\ filterType.TF multipleListOnlyOr\ filterValues.Biotypes 12Z,143B,226LDM,22Rv1,402-91,501-mel,697,786-M1A,786-O,81-3,A-137,A139,A-1847,A1A3,A2780,A2780cis,A-375,A-498,A-549,A-673,A-673-clone-Asp114,AB32,AB-LCL,AC16,adipocyte,adrenal-gland,adult-duodenal-cell,AF22,aggregated-lymphoid-nodules,AGS,ALL,ALL-SIL,AMIPS6,AMIPS8,AML,AMLPZ12,anterior-temporal-cortex,aorta,aortic-endothelial-cell,aortic-smooth-muscle-cell,arterial-endothelial-cells,artery,ASC,ascending-aorta,Aska-SS,AsPC-1,astrocyte,BA10,BA40,BC-3,BCBL-1,B-cell,BCP-ALL,BCR-ABL1,BDMC,BE2C,BEAS-2B,BG01V,BG03,BH-LCLs,BICR,BIN-67,BJ,BJ1-hTERT,BJAB,BL41,BLM,blood,BLUE1,bonchial,BPE,BPLER,brain-prefrontal-cortex,breast,breast-cancer,breast-organoid,BT-16,BT-20,BT-474,BT-549,BxPC-3,CA46,Caco-2,CAL-1,Calu-1,Calu-3,cardiac,cardiac-muscle,cardiomyocyte,cartilage,CaSki,CC-LP-1,CCLP1,ccRCC,CCRF-CEM,CD14,CD34,CD34-pos,CD4,CD4-pos,CD8,CFPAC-1,CHL-1,chondrosarcoma,choroid-plexus,CHP-134,CHRF28811,CLB-Ga,CLL,COG-N-415,COLO-205,COLO-320,COLO-741,COLO-800,COLO-829,colon,colorectal-cancer,coronary-artery,cortical-interneuron,CRL-7250,CTV-1,CUTLL1,D283-Med,D341-Med,D54,DAOY,DC,delta-47,dendrite,dermal,dermal-fibroblast,Detroit-562,DKO,DLBCL,DLD-1,DND41,DOHH2,dopaminergic-neuron,DU145,DU528,DUCAP,EDOMIPS2,EM-3,embryonic-kidney,EndoC-betaH2,endoderm,endometrial-epithelial-cells,endometrial-stromal-cell,endometrioid-adenocarcinoma,endometrium,endothelial,EP156T,epididymis,epithelial,erythroblast,erythroid,erythroid-progenitor,ESF,ESO-26,esophagus,esophagus-muscularis-mucosa,esophagus-squamous-epithelium,FaDu,fetal,fibroblast,FLP143HA,FLP76,foregut,foreskin,FT282,G296S,G-401,G523NS,gastric-epithelial-cell,gastrocnemius-medialis,gastroesophageal-sphincter,GBM1A,GEN2-2,GIC,GIST,GIST48,GIST882,GIST-T1,glioblastoma,glioma,GM00011,GM01310,GM04025,GM04604,GM04648,GM06077,GM06170,GM06990,GM08714,GM09236,GM09237,GM10248,GM10266,GM10847,GM12801,GM12864,GM12865,GM12866,GM12867,GM12868,GM12869,GM12870,GM12871,GM12872,GM12873,GM12874,GM12875,GM12878,GM12891,GM12892,GM13976,GM13977,GM15510,GM15850,GM17942,GM18505,GM18526,GM18951,GM19099,GM19193,GM20000,GM23248,GM23338,GP5D,GRANT-A519,GSC,GSC23,GSC8-11,H-1,H69,H9,HaCaT,HAEC,HAP1,HASMC,HBE,HBTEC,HCAEC,HCASMC,HCC1143,HCC1187,HCC1395,HCC1428,HCC1599,HCC1806,HCC1937,HCC1954,HCC2157,HCC2814,HCC70,HCC95,HCCLM3,HCT-116,HCT-15,HDF,heart,HEC-1-A,HEC-1-B,HEE,HEK,HEK293,HEK293-FT,HEK293T,HEL,HeLa,HeLa-B2,HeLa-Kyoto,HeLa-S3,HeLa-Tet-On,HEP,HEP10-01008-LCLs,HEP14-00120-LCLs,HEP14-0079-LCLs,HEP14-0080-LCLs,Hep-3B2-1-7,HepaRG,hepatocellular-carcinoma-cell,hepatocyte,Hep-G2,hESC,hESC-1,HEY-A8,HFF,HFOB,HGrC1,HIES,hiF-T,hippocampus,hiPSC,HKC,HL-60,hMADS,HMEC-1,HMELBRAF,HMLE,HMLER,HMLE-Twist-ER,HMS001,hMSC,hMSC-TERT,hMSC-TERT4,HNPC,HNSC,HPBALL,Hs-352-Sk,HS578T,HSPC,HSPC-CD34,HSPC-CD34pos,HS-SY-2,HT-1080,HT29,hTERT-HME1,HUCCT1,HUDEP-2,HUES-64,HUES-8,HUG1N,Huh-7,HUVEC-C,ID00014,ID00015,ID00016,IM95,IMEC,IMR-5,IMR-90,IMS-M2,induced-endothelial-cell,intestinal-cell,Ishikawa,islet,JD-LCLs,JHU-029,J-Lat,JL-LCLs,JMSU-1,Jurkat,K-562,Karpas-299,Karpas-422,KARPAS422,Karpas-45,Kasumi-1,KATO-III,KB,Kelly,keratinocyte,KerCT,KG-1,KGN,kidney,kidney-cortex,KK-1,KMS-11,KNS-62,KOPN-8,KOPT-K1,KYSE-150,KYSE-70,L1236,L826,LA-N-5,LA-N-6,LAPC-4,LBCL,LCL,LCLGM10861,leiomyoma,leukemia,LHSAR,liver,LK2,LNAR,LNCaP,LNCaP-95,LNCaP-abl,LNCaP-C4-2,LNCaP-C4-2B,LNCaP-clone-FGC,LNCaP-FGC,Loucy,LoVo,LOX-IMVI,LP1,LPS141,LREX,LS174T,LS180,LTAD,Lu-130,lung,LX2,lymphoblast,lymphocyte,macrophage,MALME-3,mammary-epithelial-cell,MCF-10A,MCF10A-Er-Src,MCF-10AT1,MCF-10CA1a,MCF-7,MCF-7L,MCF-7-Luc,MCF-7-Luc-Y537S,MCF-7-TAMR-1,MCF7-Tet-On,MCF-7-WS8,MDA-BoM-1833,MDA-MB-134-VI,MDA-MB-157,MDA-MB-231,MDA-MB-361,MDA-MB-435,MDA-MB-436,MDA-MB-453,MDA-MB-468,MDA-Pca-2b,MDM,ME-1,medulloblastoma,Mel270,melanocyte,mesenchymal,metastatic-neuroblastoma,MG-63-3,MIA-PaCa-2,MKN28,MKN74,ML-2,MM1-S,MNNG-HOS,MO91,MOLM-13,MOLM-14,MOLT-3,MOLT-4,monocyte,MPNST,MRC-5,MSTO,Mutu-1,MUTUL,MV4-11,MV4-11-B,MYCN-3,myoblast,myofibroblast,myometrium,myotube,NALM-6,Namalwa,NB-1643,NB4,NB69,NCCIT,NCI-H1048,NCI-H128,NCI-H1299,NCI-H1703,NCI-H1819,NCI-H1963,NCI-H1975,NCI-H2087,NCI-H2107,NCI-H2171,NCI-H23,NCI-H295R,NCI-H3122,NCI-H3396,NCI-H358,NCI-H441,NCI-H460,NCI-H520,NCI-H524,NCI-H526,NCI-H82,NCI-H838,NCI-H889,NCI-H929,nerve,neural,neural-progenitor,neuroblastoma,neuroepithelilal-cells,neuron,neuron-progenitor,neutrophil,NGP,NHEK,NMC24335,NOMO1,NPC,NSC,NT2-D1,NTERA2,NUT,NY15,OACP4-C,OCI-AML-3,OCI-AML3,OCI-Ly1,OCI-Ly10,OCI-Ly19,OCI-Ly3,OCI-Ly7,ocular-melanoma-cell,OE33,omental-fat-pad,OSK,OSKM,osteoblast,OSvK,OSvKM,ovary,OVCA429,OVCAR-3,OVCAR-5,OVCAR-8,OVSAHO,P12,P493-6,PANC-1,pancreas,pancreatic-progenitor,PATU8988,PAVE,PBMC,PC-3,PC-9,PDAC,PEO1,PER-117,peripheral-blood-mononuclear-cell,peripheral-blood-neutrophil,Peyers-patch,PF-382,Pfeiffer,PFSK1,PK-LCLs,placenta,plasmablast,pleural-effusion,pre-B-cell,PrEC,PRIMA2,PRIMA5,primary-B-cell,primary-breast-cancer,primary-bronchial-epithelial,primary-chondrocyte,primary-dermal-fibroblasts,primary-endometrial-stromal-cell,primary-endometrium-cancer,primary-epidermal-keratinocyte,primary-glioblastoma,primary-keratinocyte,primary-lung-fibroblast,primary-monocyte,primary-neutrophil,primary-prostate-cancer,primary-prostate-epithelial-cell,primordial-germ-cell-like-cell,ProEs,proliferating-human-fibroblast,prostate,prostate-cancer,pulmonary-artery,Raji,Ramos,RCC10,RCC4,RCH-ACV,RD,REC-1,Reh,RENVM,retina,Rh18,RH3,RH30,RH4,Rh41,RH5,rhabdomyosarcoma,RKO,RL,RMG-I,RPE,RPMI8402,RS4-11,RWPE-1,RWPE-2,SaOS-2,SCC,SCC-25,SCC-9,SCCOHT-1,SCLC,SCMC,SEM,SET-2,SF8628,SGBS,SH-EP,SHEP-21N,SHI-1,SH-SY5Y,sigmoid-colon,SiHa,SJSA-1,SK-BR-3,SKH1,skin,SKM-1,SK-MEL-147,SK-MEL-239,SK-MEL-28,SK-MEL-5,SK-N-AS,SK-N-BE2,SK-N-BE2-C,SK-N-MC,SKNO-1,SK-N-SH,SK-UT-1,SLK,SMMC-7721,smooth-muscle-cell,SMS-CTR,SMS-KCN,SMS-KCNR,SNU-216,SNU-398,SP-49,spleen,ST-1,stomach,subcutaneous-adipose-tissue,SU-DHL-10,SU-DHL-2,SU-DHL-4,SU-DHL-5,SU-DHL-6,SUIT-2,SUM1315,SUM149,SUM149PT,SUM159,SUM159PT,SUM185,SUM229PE,SUM44PE,SUP-B15,SVOG-3e,SW1353,SW1783,SW1990,SW480,SW620,SYO-1,T-47D,T-47D-A,T47D-A1-2,T-47D-B,T-47D-MTVL,T778,T98G,TALL-1,TC-32,TC-71,T-cell,TE-5,testis,TF1,Th1,Th17,T-HESCs,thoracic-aorta,THP-1,THP-6,thymocyte,thymus,thyroid-cancer,thyroid-gland,tibial-artery,tibial-nerve,TMD8,tonsil,TOV-21G,T-REx-293,TSU-1621MT,TT,TTC-1240,TTC-549,U266,U266B1,U2932,U2OS,U-87MG,U-937,UACC-257,UACC-62,UAE,UCLA1-hESCs,UCSD-AML1,UM-RC-6,UO-31,UPCI-SCC-090,UTEIPS11,UTEIPS4,UTEIPS6,UTEIPS7,uterus,vagina,VCaP,VCaP-LTAD,VU-SCC-147,WA01,WA09,WERI-Rb-1,WHIM12,WI-38,WI-38VA13,WIBR3,WN8532,WPMY-1,WSU-DLCL2,YCC-3,ZR-75-1,ZR751\ filterValues.TF AATF,ADNP,AEBP2,AFF1,AFF4,AGO1,AHR,AHRR,APC,AR,ARHGAP35,ARID1A,ARID1B,ARID2,ARID3A,ARID3B,ARID4A,ARID4B,ARID5B,ARNT,ARNTL,ARRB1,ASCL1,ASH1L,ASH2L,ASXL1,ASXL3,ATF1,ATF2,ATF3,ATF4,ATF7,ATM,ATOH8,ATRX,ATXN7L3,BACH1,BACH2,BAF155,BAHD1,BAP1,BATF,BATF3,BCL11A,BCL11B,BCL3,BCL6,BCL6B,BCLAF1,BCOR,BDP1,BHLHE22,BHLHE40,BICRA,BMI1,BMPR1A,BNC2,BPTF,BRCA1,BRD1,BRD2,BRD3,BRD4,BRD7,BRD9,BRF1,BRF2,C17orf49,CARM1,CASZ1,CBFA2T2,CBFA2T3,CBFB,CBX1,CBX2,CBX3,CBX4,CBX5,CBX7,CBX8,CC2D1A,CCAR2,CCNT2,CD74,CDC5L,CDK2,CDK6,CDK7,CDK8,CDK9,CDK9-HEXIM1,CDKN1B,CDX2,CEBPA,CEBPB,CEBPD,CEBPG,CEBPZ,CERS6,CHAF1B,CHAMP1,CHD1,CHD2,CHD4,CHD7,CHD8,CIITA,CLOCK,COBLL1,CREB1,CREB3,CREB3L1,CREB5,CREBBP,CREM,CRX,CRY1,CRY2,CSDC2,CSNK2A1,CTBP1,CTBP2,CTCF,CTCFL,CTNNB1,CUX1,CXXC4,CXXC5,DACH1,DAXX,DDX20,DDX21,DDX5,DEAF1,DEK,DIDO1,DLX4,DLX6,DMAP1,DNMT1,DNMT3B,DPF1,DPF2,DR1,DRAP1,DUX4,E2F1,E2F3,E2F4,E2F5,E2F6,E2F7,E2F8,E4F1,EBF1,EBF3,EED,EGR1,EHF,EHMT2,ELF1,ELF2,ELF3,ELF4,ELF5,ELK1,ELK4,ELL,ELL2,EOMES,EP300,EP400,EPAS1,ERF,ERG,ESR1,ESR2,ESRRA,ESRRB,ESRRG,ETS1,ETS2,ETV1,ETV2,ETV4,ETV6,EVI1,EWSR1,EZH1,EZH2,FANCD2,FANCL,FEZF1,FIP1L1,FLI1,FOS,FOSB,FOSL1,FOSL2,FOXA1,FOXA2,FOXF1,FOXF2,FOXJ2,FOXJ3,FOXK1,FOXK2,FOXL2,FOXM1,FOXO1,FOXO1-PAX3,FOXO3,FOXP1,FOXP2,FOXP4,FOXS1,FUS,GABPA,GABPB1,GATA1,GATA2,GATA3,GATA4,GATA6,GATAD1,GATAD2A,GATAD2B,GFI1,GFI1B,GLI1,GLI2,GLI4,GLIS1,GLIS2,GLIS3,GLYR1,GMEB1,GMEB2,GPS2,GR,GRHL1,GRHL2,GSPT2,GTF2A2,GTF2B,GTF2F1,GTF3A,GTF3C2,GTF3C5,HAND2,HBP1,HCFC1,HCFC1R1,HDAC1,HDAC2,HDAC3,HDAC6,HDAC8,HDGF,HES1,HEXIM1,HEXIM1-CDK9,HEY1,HEY2,HHEX,HIC1,HIF1A,HIF3A,HINFP,HIVEP1,HKR1,HLF,HMBOX1,HMGA1,HMGB1,HMGB2,HMGN3,HMGXB4,HNF1A,HNF1B,HNF4A,HNF4G,HNRNPC,HNRNPH1,HNRNPK,HNRNPL,HNRNPLL,HNRNPUL1,HOMEZ,HOXA3,HOXA7,HOXA9,HOXB13,HOXB5,HOXB7,HOXB8,HOXC5,HOXC6,HSF1,HSF2,ICE1,ICE2,ID3,IFNA1,IKZF1,IKZF2,IKZF3,ILF3,ILK,INO80,INSM2,INTS11,INTS13,IRF1,IRF2,IRF2BP2,IRF3,IRF4,IRF5,IRF8,IRF9,ISL1,ISL2,JARID2,JDP2,JMJD1C,JMJD6,JUN,JUNB,JUND,KAT2A,KAT2B,KAT7,KAT8,KDM1A,KDM3A,KDM4A,KDM4B,KDM4C,KDM5A,KDM5B,KDM6B,KLF1,KLF10,KLF12,KLF13,KLF14,KLF15,KLF16,KLF17,KLF3,KLF4,KLF5,KLF6,KLF7,KLF8,KLF9,KMT2A,KMT2B,KMT2C,KMT2D,L3MBTL2,L3MBTL4,LCORL,LDB1,LEF1,LHX2,LIN54,LIN9,LMO1,LMO2,LYL1,MAF,MAF1,MAFB,MAFF,MAFG,MAFK,MAML1,MAML3,MAX,MAZ,MBD1,MBD2,MBD3,MBD4,MCM2,MCM3,MCM5,MCM7,MCRS1,MECOM,MECP2,MED,MED1,MED12,MED25,MED26,MEF2A,MEF2B,MEF2C,MEF2D,MEIS1,MEIS2,MEN1,MGA,MIER1,MITF,MLL4,MLLT1,MLLT3,MLX,MLXIP,MNT,MNX1,MORC2,MPHOSPH8,MRTFA,MRTFB,MSX2,MTA1,MTA2,MTA3,MTF2,MXD4,MXI1,MYB,MYBL2,MYC,MYC-DAXX,MYCN,MYF5,MYNN,MYOCD,MYOD1,MYOG,MZF1,NAB2,NANOG,NBN,NCAPH2,NCBP1,NCOA1,NCOA2,NCOA3,NCOA4,NCOA6,NCOR1,NCOR2,NELFA,NELFCD,NELFE,NEUROD1,NEUROG2,NFAT5,NFATC1,NFATC2,NFATC3,NFE2,NFE2L1,NFE2L2,NFIA,NFIB,NFIC,NFIL3,NFIX,NFKB1,NFKB2,NFKBIA,NFKBIZ,NFRKB,NFXL1,NFYA,NFYB,NFYC,NIPBL,NKX2-1,NKX2-5,NKX3-1,NME2,NONO,NOTCH1,NOTCH3,NR0B1,NR1H2,NR1H3,NR2C1,NR2C2,NR2F1,NR2F2,NR2F6,NR3C1,NR4A1,NR5A1,NR5A2,NRF1,NRIP1,NRL,NSD2,NUFIP1,NUP98-HOXA9,NUTM1,OGG1,OGT,OLIG2,ONECUT1,ONECUT2,OSR2,OTX2,OVOL1,OVOL3,PAF1,PALB2,PARP1,PATZ1,PAX3-FOXO1,PAX5,PAX6,PAX7,PAX8,PAXIP1,PBX1,PBX1-2-3,PBX2,PBX3,PCBP1,PCBP2,PCGF1,PCGF2,PDX1,PGR,PHB2,PHC1,PHF19,PHF20,PHF21A,PHF5A,PHF8,PHIP,PHOX2B,PITX3,PKNOX1,PLAG1,PLRG1,PML,POU2AF1,POU2F1,POU2F2,POU2F3,POU3F1,POU3F2,POU4F2,POU5F1,PPARA,PPARG,PPARGC1A,PRDM1,PRDM10,PRDM12,PRDM14,PRDM15,PRDM2,PRDM4,PRDM6,PREB,PRKDC,PRMT5,PROX1,PRPF4,PSIP1,PTBP1,PTRF,PTTG1,PYGO2,RAD21,RAD51,RARA,RB1,RBAK,RBBP4,RBBP5,RBFOX2,RBM14,RBM15,RBM22,RBM25,RBM34,RBM39,RBP2,RBPJ,RCOR1,REL,RELA,RELB,REPIN1,REST,RFX1,RFX2,RFX3,RFX5,RFXAP,RING1,RLF,RNF2,RORB,RORC,RPA2,RREB1,RUNX1,RUNX1-3,RUNX1-RUNX1T1,RUNX1T1,RUNX2,RUVBL1,RUVBL2,RXR,RXRA,RYBP,SAFB,SAFB2,SALL1,SALL2,SALL3,SALL4,SAP30,SATB1,SCRT1,SETDB1,SETX,SFMBT1,SFPQ,SGF29,SHOX2,SIN3A,SIN3B,SIRT3,SIRT6,SIX1,SIX2,SIX4,SIX5,SKI,SKIL,SMAD1,SMAD1-5,SMAD1-5-8,SMAD2,SMAD2-3,SMAD3,SMAD3-EPAS1,SMAD3-HIF1A,SMAD4,SMAD5,SMARCA2,SMARCA4,SMARCA5,SMARCB1,SMARCC1,SMARCC2,SMARCD3,SMARCE1,SMC1,SMC1A,SMC1A-B,SMC3,SMC4,SNAI1,SNAI2,SNAPC1,SNAPC4,SND1,SNIP1,SNRNP70,SOX10,SOX11,SOX13,SOX2,SOX21,SOX3,SOX4,SOX6,SOX8,SOX9,SP1,SP140L,SP2,SP3,SP4,SP5,SP7,SPDEF,SPI1,SPIB,SPIN1,SRC,SREBF1,SREBF2,SREBP2,SRF,SRSF1,SRSF3,SRSF4,SRSF7,SRSF9,SS18,SS18-SSX,SSRP1,STAG1,STAG2,STAT1,STAT2,STAT3,STAT5A,STAT5B,SUPT16H,SUPT5H,SUPT6H,SUZ12,SVIL,T,TAF1,TAF15,TAF2,TAF3,TAF7,TAF9B,TAL1,TARDBP,TASOR,TBL1X,TBL1XR1,TBP,TBX18,TBX2,TBX21,TBX3,TBX5,TCF12,TCF21,TCF25,TCF3,TCF3-PBX1,TCF4,TCF7,TCF7L2,TCFL5,TCOF1,TEAD1,TEAD2,TEAD4,TERF1,TERF2,TERT,TET2,TFAP2A,TFAP2C,TFAP4,TFCP2,TFDP1,TFDP2,TFE3,TFEB,TFIIIC,TGIF2,THAP1,THAP11,THRA,THRAP3,THRB,TLE3,TOP1,TOP2A,TOX2,TP53,TP63,TP73,TRIM22,TRIM24,TRIM25,TRIM28,TRIP13,TRPS1,TRRAP,TSC22D4,TSHZ1,TSHZ2,TWIST1,U2AF1,U2AF2,UBN1,UBTF,USF1,USF2,USP7,UTX,VDR,VEZF1,WDHD1,WDR5,WRNIP1,WT1,XBP1,XRCC3,XRCC5,XRN2,YAP1,YBX1,YBX3,YY1,YY1AP1,YY2,ZBED1,ZBED2,ZBED4,ZBTB1,ZBTB10,ZBTB11,ZBTB12,ZBTB14,ZBTB16,ZBTB18,ZBTB2,ZBTB20,ZBTB21,ZBTB24,ZBTB26,ZBTB33,ZBTB40,ZBTB42,ZBTB44,ZBTB48,ZBTB49,ZBTB5,ZBTB6,ZBTB7A,ZBTB7B,ZBTB8A,ZC3H11A,ZC3H8,ZEB1,ZEB2,ZFP14,ZFP28,ZFP3,ZFP36,ZFP37,ZFP41,ZFP42,ZFP57,ZFP64,ZFP69,ZFP69B,ZFP82,ZFP90,ZFP91,ZFX,ZFY,ZGPAT,ZHX1,ZHX2,ZIC2,ZIC5,ZIK1,ZIM3,ZKSCAN1,ZKSCAN2,ZKSCAN3,ZKSCAN5,ZKSCAN8,ZMIZ1,ZMYM2,ZMYM3,ZMYND11,ZMYND8,ZNF10,ZNF101,ZNF112,ZNF114,ZNF12,ZNF121,ZNF124,ZNF132,ZNF133,ZNF134,ZNF135,ZNF136,ZNF138,ZNF140,ZNF141,ZNF142,ZNF143,ZNF146,ZNF148,ZNF154,ZNF155,ZNF157,ZNF16,ZNF165,ZNF169,ZNF17,ZNF174,ZNF175,ZNF18,ZNF180,ZNF182,ZNF184,ZNF189,ZNF19,ZNF195,ZNF197,ZNF2,ZNF202,ZNF205,ZNF207,ZNF211,ZNF212,ZNF213,ZNF214,ZNF215,ZNF217,ZNF22,ZNF221,ZNF222,ZNF223,ZNF224,ZNF225,ZNF23,ZNF232,ZNF239,ZNF24,ZNF248,ZNF25,ZNF250,ZNF253,ZNF256,ZNF257,ZNF26,ZNF260,ZNF263,ZNF264,ZNF266,ZNF267,ZNF273,ZNF274,ZNF276,ZNF28,ZNF280A,ZNF280C,ZNF280D,ZNF281,ZNF282,ZNF283,ZNF284,ZNF285,ZNF287,ZNF292,ZNF3,ZNF30,ZNF300,ZNF302,ZNF304,ZNF311,ZNF316,ZNF317,ZNF318,ZNF319,ZNF320,ZNF322,ZNF324,ZNF329,ZNF331,ZNF333,ZNF335,ZNF337,ZNF33A,ZNF33B,ZNF34,ZNF341,ZNF343,ZNF35,ZNF350,ZNF354A,ZNF354B,ZNF354C,ZNF362,ZNF366,ZNF37A,ZNF383,ZNF384,ZNF391,ZNF394,ZNF395,ZNF397,ZNF398,ZNF404,ZNF407,ZNF408,ZNF41,ZNF410,ZNF416,ZNF417,ZNF418,ZNF423,ZNF425,ZNF426,ZNF429,ZNF430,ZNF431,ZNF432,ZNF433,ZNF436,ZNF44,ZNF440,ZNF441,ZNF444,ZNF445,ZNF449,ZNF454,ZNF460,ZNF462,ZNF467,ZNF468,ZNF473,ZNF479,ZNF48,ZNF480,ZNF483,ZNF484,ZNF485,ZNF487,ZNF488,ZNF490,ZNF491,ZNF492,ZNF493,ZNF496,ZNF501,ZNF502,ZNF503,ZNF506,ZNF507,ZNF510,ZNF512,ZNF512B,ZNF513,ZNF514,ZNF518A,ZNF519,ZNF521,ZNF524,ZNF527,ZNF528,ZNF529,ZNF530,ZNF532,ZNF534,ZNF540,ZNF543,ZNF544,ZNF547,ZNF548,ZNF549,ZNF550,ZNF554,ZNF555,ZNF557,ZNF558,ZNF560,ZNF561,ZNF563,ZNF565,ZNF566,ZNF567,ZNF57,ZNF570,ZNF571,ZNF572,ZNF573,ZNF574,ZNF577,ZNF579,ZNF580,ZNF582,ZNF583,ZNF584,ZNF585A,ZNF585B,ZNF586,ZNF587,ZNF589,ZNF592,ZNF595,ZNF596,ZNF597,ZNF598,ZNF605,ZNF609,ZNF610,ZNF611,ZNF613,ZNF614,ZNF616,ZNF621,ZNF622,ZNF623,ZNF624,ZNF626,ZNF627,ZNF629,ZNF639,ZNF641,ZNF644,ZNF645,ZNF649,ZNF652,ZNF654,ZNF658,ZNF660,ZNF662,ZNF664,ZNF667,ZNF669,ZNF670,ZNF671,ZNF674,ZNF675,ZNF677,ZNF680,ZNF681,ZNF684,ZNF687,ZNF692,ZNF695,ZNF696,ZNF697,ZNF7,ZNF700,ZNF701,ZNF704,ZNF707,ZNF708,ZNF711,ZNF714,ZNF716,ZNF730,ZNF736,ZNF737,ZNF740,ZNF747,ZNF749,ZNF750,ZNF75A,ZNF76,ZNF764,ZNF765,ZNF766,ZNF768,ZNF77,ZNF770,ZNF774,ZNF776,ZNF777,ZNF778,ZNF780A,ZNF781,ZNF783,ZNF784,ZNF785,ZNF786,ZNF789,ZNF79,ZNF791,ZNF792,ZNF799,ZNF8,ZNF800,ZNF808,ZNF81,ZNF816,ZNF823,ZNF83,ZNF830,ZNF837,ZNF84,ZNF843,ZNF846,ZNF85,ZNF860,ZNF879,ZNF880,ZNF883,ZNF891,ZNF90,ZNF92,ZNF93,ZSCAN16,ZSCAN18,ZSCAN2,ZSCAN21,ZSCAN22,ZSCAN23,ZSCAN26,ZSCAN29,ZSCAN30,ZSCAN31,ZSCAN4,ZSCAN5A,ZSCAN5C,ZXDB,ZXDC,ZZZ3\ html ../reMap\ itemRgb on\ labelFields name, TF, Biotypes\ longLabel ReMap Atlas of Regulatory Regions\ maxItems 10000\ maxWindowCoverage 20000\ parent ReMap on\ priority 2\ shortLabel ReMap ChIP-seq\ showCfg on\ track ReMapTFs\ type bigBed 9 +\ urls TF="http://remap.univ-amu.fr/target_page/$$:9606" Biotypes="http://remap.univ-amu.fr/biotype_page/$$:9606"\ visibility squish\ rmskJoinedCurrent RepeatMasker Viz. bed 3 + RepeatMasker v4.0.7 Dfam_2.0 : Current Dataset 0 2 0 0 0 127 127 127 1 0 0 rep 0 longLabel RepeatMasker v4.0.7 Dfam_2.0 : Current Dataset\ parent joinedRmsk on\ priority 2\ shortLabel RepeatMasker Viz.\ track rmskJoinedCurrent\ gnomad35XPercentage Sample % > 5X bigWig gnomAD Percentage of Genome Samples with at least 5X Coverage v3.0.1 2 2 225 0 30 240 127 142 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_5.bw\ color 225,0,30\ longLabel gnomAD Percentage of Genome Samples with at least 5X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 2\ shortLabel Sample % > 5X\ track gnomad35XPercentage\ viewLimits 0:1\ gnomad4Exome5XPercentage Sample % > 5X bigWig gnomAD Percentage of Exome Samples with at least 5X Coverage v4.0 2 2 225 0 30 240 127 142 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_5.bw\ color 225,0,30\ longLabel gnomAD Percentage of Exome Samples with at least 5X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 2\ shortLabel Sample % > 5X\ track gnomad4Exome5XPercentage\ viewLimits 0:1\ miRnaAtlasSample2BarChart Sample 2 bigBarChart miRNA Tissue Atlas microRna Expression 2 2 0 0 0 127 127 127 0 0 0\ The Human miRNA Tissue Atlas is a\ catalog of tissue-specific microRNA (miRNA) expression across 62 tissues. This track contains\ quantile normalized miRNA expression data sampled from two individuals and mapped to\ miRBase v21 coordinates. The track contains two subtracks, one\ for each individual sampled.
\ \\ The Tissue Specificity Index (TSI) is analogous to the "tau" value for mRNA expression,\ and is calculated as described in the\ \ associated publication. Values closer to 0 indicate miRNAs expressed in many or all tissues,\ while values closer to 1 indicate miRNAs expressed only in a specific tissue or tissues. To\ browse miRNAs by TSI value, please see the\ miRNA Tissue Atlas.
\ \\ This track is formatted as a barChart track,\ similar to the GTEx or the\ TCGA Cancer Expression tracks, where the\ heights of each bar indicate the expression value for the miRNA in a specific tissue. The tissues\ sampled are described in the table below:\
\| Bar Color | Sample 1 | Sample 2 |
| Adipocyte | Adipocyte | |
| Artery | Artery | |
| Colon | Colon | |
| Dura mater | Dura mater | |
| Kidney | Kidney | |
| Liver | Liver | |
| Lung | Lung | |
| Muscle | Muscle | |
| Myocardium | Myocardium | |
| Skin | Skin | |
| Spleen | Spleen | |
| Stomach | Stomach | |
| Testis | Testis | |
| Thyroid | Thyroid | |
| Small intestine | ||
| Bone | ||
| Gallbladder | ||
| Fascia | ||
| Bladder | ||
| Epididymis | ||
| Tunica albuginea | ||
| Nervus intercostalis | ||
| Arachnoid mater | ||
| Brain | ||
| Small intestine duodenum | ||
| Small intestine jejunum | ||
| Pancreas | ||
| Kidney glandula suprarenalis | ||
| Kidney cortex renalis | ||
| Esophagus | ||
| Prostate | ||
| Bone marrow | ||
| Vein | ||
| Lymph node | ||
| Nerve not specified | ||
| Pleura | ||
| Pituitary gland | ||
| Spinal cord | ||
| Thalamus | ||
| Brain white matter | ||
| Nucleus caudatus | ||
| Kidney medulla renalis | ||
| Brain gray_matter | ||
| Cerebral cortex temporal | ||
| Cerebral cortex frontal | ||
| Cerebral cortex occipital | ||
| Cerebellum |
\ The 14 shared tissues sampled across both individuals are presented in the same order for easier comparison.\
\ \\ The underlying expression matrix and TSI values can be obtained from the\ miRNA tissue atlas website, in the\ data_matrix_quantile.txt and tsi_quantile.csv files.\
\ \\ Ludwig N, Leidinger P, Becker K, Backes C, Fehlmann T, Pallasch C, Rheinheimer S, Meder B,\ Stähler C, Meese E et al.\ \ Distribution of miRNA expression across human tissues.\ Nucleic Acids Res. 2016 May 5;44(8):3865-77.\ PMID: 26921406; PMC: PMC4856985\
\ expression 1 barChartBars adipocyte artery colon dura_mater kidney liver lung muscle myocardium skin spleen stomach testis thyroid small_intestine_duodenum small_intestine_jejunum pancreas kidney_glandula_suprarenalis kidney_cortex_renalis kidney_medulla_renalis esophagus prostate bone_marrow vein lymph_node nerve_not_specified pleura brain_pituitary_gland spinal_cord brain_thalamus brain_white_matter brain_nucleus_caudatus brain_gray_matter brain_cerebral_cortex_temporal brain_cerebral_cortex_frontal brain_cerebral_cortex_occipital brain_cerebellum\ barChartColors #F7A028 #F73528 #DEBE98 #86BF80 #CDB79E #CDB79E #9ACD32 #7A67AE #9745AC #1E90FF \\#CDB79E #FFD39B #A6A6A6 #008B45 #CDB79E #CDB79E #CD9B1D \\#CDB79E #CDB79E #CDB79E #AC8F69 #D9D9D9 #BD3487 \\#FF00FF #EE82EE #F7E300 #73A585 #B4EEB4 #EEEE00 \\#EEEE00 #EEEE00 #EEEE00 #EEEE00 \\#EEEE00 #EEEE00 \\#EEEE00 #EEEE00\ barChartLabel Tissue\ barChartMatrixUrl /gbdb/hgFixed/human/expMatrix/miRnaAtlasSample2Matrix.txt\ barChartSampleUrl /gbdb/hgFixed/human/expMatrix/miRnaAtlasSample2.txt\ barChartUnit Quantile_Norm_Expr\ bigDataUrl /gbdb/hg38/bbi/miRnaAtlasSample2.bb\ configurable on\ group expression\ html miRnaAtlas\ longLabel miRNA Tissue Atlas microRna Expression\ maxLimit 52000\ parent miRnaAtlasSample2\ searchIndex name\ shortLabel Sample 2\ subGroups view=b_B\ track miRnaAtlasSample2BarChart\ url2 http://www.mirbase.org/cgi-bin/query.pl?terms=$$\ url2Label miRBase v21 Precursor Accession:\ visibility full\ snpediaText SNPedia with text bed 4 SNPedia pages with manually typed text 0 2 50 0 100 152 127 177 0 0 0 https://www.snpedia.com/index.php/$$ phenDis 1 color 50,0,100\ exonNumbers off\ itemDetailsHtmlTable snpediaTextHtml\ longLabel SNPedia pages with manually typed text\ parent snpedia\ shortLabel SNPedia with text\ track snpediaText\ type bed 4\ url https://www.snpedia.com/index.php/$$\ urlLabel Link to SNPedia page:\ spliceAiAccMinus SpliceAI Acceptor Minus bigWig 0 1 SpliceAI Splice Acceptor Sites, Minus Strand 2 2 0 0 0 127 127 127 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/bbi/spliceAi/wildtype/spliceAiAcceptorMinus.bw\ longLabel SpliceAI Splice Acceptor Sites, Minus Strand\ parent spliceAIWt on\ priority 2\ shortLabel SpliceAI Acceptor Minus\ track spliceAiAccMinus\ type bigWig 0 1\ spliceAIWt SpliceAI Wildtype bigWig SpliceAI Wildtype: Splicing of the reference genome sequence 2 2 0 0 0 127 127 127 0 0 0\ The "Splicing Impact" container track contains tracks showing the predicted or validated effect of variants\ close to splice sites.\
\ \AbSplice is a method that predicts aberrant splicing across human tissues, as described in Wagner,\ Çelik et al., 2023. This track displays precomputed AbSplice scores for all possible\ single-nucleotide variants genome-wide. The scores represent the probability that a given variant\ causes aberrant splicing in a given tissue.\ AbSplice scores\ can be computed from VCF files and are based on quantitative tissue-specific splice site annotations\ (SpliceMaps).\ While SpliceMaps can be generated for any tissue of interest from a cohort of RNA-seq samples, this\ track includes 49 tissues available from the\ Genotype-Tissue\ Expression (GTEx) dataset.\
\ \SpliceAI is an open-source deep\ learning splicing prediction algorithm that can predict splicing alterations caused by DNA variations.\ To score variants, the spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added to the sequence and the new sequence is\ scored. Variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a\ lookup tool\ is provided by the Broad institute. \
\ \\ This SpliceAI "Wildtype" container track shows the scores for the genome sequence itself,\ without variants, from predicted splice donor (5' intron boundaries) and splice acceptor\ (3' intron boundaries) sites. Predictions are strand-specific, with separate subtracks for the\ plus and minus strands. These tracks are useful in combination with the variants track for\ evaluating new transcript models. They can be used to assess potential exon boundaries or\ possible splice acceptor sites.
\ \ Why are some variants not scored by SpliceAI?\\ SpliceAI only annotates variants within genes defined by the gene\ annotation file. Additionally, SpliceAI does not annotate variants if they are close to chromosome\ ends (5kb on either side), deletions of length greater than twice the input parameter -D, or\ inconsistent with the reference fasta file.\
\ \ What are the differences between masked and unmasked tracks?\\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \SpliceVarDB is an online database consolidating over 50,000 variants assayed\ for their effects on splicing in over 8,000 human genes. The authors evaluated\ over 500 published data sources and established a spliceogenicity scale to\ standardize, harmonize, and consolidate variant validation data generated by a\ range of experimental protocols. Genes and variant locations were obtained using\ GENCODE v44. Splice regions were calculated as specific distances from the closest\ canonical exon, including 5' and 3' untranslated regions (UTRs). The\ database is available at\ splicevardb.org.
\ \The AbSplice score is a probability estimate of how likely aberrant splicing of some sort takes\ place in a given tissue. The authors suggest three cutoffs which are represented by color in the track.\
\ \\ Mouseover on items shows the gene name, maximum score, and tissues that had this score. Clicking on\ any item brings up a table with scores for all 49 GTEX tissues.\
\ \\ Variants are colored according to Walker et al. 2023 splicing impact:\
\\ The scores range from 0 to 1 and can be interpreted as the\ probability of the variant being splice-altering. In the paper, a detailed characterization is\ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\ These tracks are in bigWig format. The signal height represents the SpliceAI probability score.\ This track may be configured in a variety of ways to highlight different aspects of the displayed\ information. Click the "Graph configuration help" link for an explanation of configuration\ options.
\ \According to the strength of their supporting\ evidence, variants were classified as "splice-altering" (~25%), "not\ splice-altering" (~25%), and "low-frequency splice-altering" (~50%), which\ correspond to weak or indeterminate evidence of spliceogenicity. 55% of the\ splice-altering variants in SpliceVarDB are outside the canonical splice sites\ (5.6% are deep intronic). The data is shown as lollipop plots that can be clicked, \ the details page then shows a link to SpliceVarDB with full details.\
\ \The classification thresholds primarily follow those established by the original study.\ However, most studies only defined criteria for splice-altering variants and did not define\ criteria for variants that resulted in normal splicing. The authors implemented stringent\ thresholds to define the normal category and ensure a high-quality set of control variants.\ Variants that did not meet these criteria were classified as low-frequency splice-altering\ variants with a wide range of sub-optimal scores. Variants that fell between the normal and\ splice-altering classifications were placed into a low-frequency splice-altering category.\ In situations where a variant was validated multiple times, if at least one validation\ returned splice-altering and another returned normal, the "conflicting" category\ was applied.\
\ \\ The lollipop plots are color-coded based on the score value, which corresponds\ to the following classifications:\
Data was converted from the files (AbSplice_DNA_ hg38 _snvs_high_scores.zip) provided by the authors\ at zenodo.org. Files in the\ score_cutoff=0.01 directory were concatenated. To convert the data to bigBed format, scores and\ their tissues were selected from the AbSplice_DNA fields and maximum scores, and then calculated\ using a custom Python script, which can be found in the\ \ makeDoc from our GitHub repository.
\ \\
The data were downloaded from Illumina.\
The spliceAI scores are represented in the VCF INFO field as\
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain\
\ Since most of the values are 0 or almost 0, we selected only those variants\ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \Data was provided by the Michael Hiller lab. SpliceAI was run on the entire genome reference\ chromosomes. Since the algorithm does not know where transcripts start or end, the scores\ can differ from those on other websites, especially for splice sites before the last exon or\ around the first exon.
\ \ \The data was converted by Patricia Sullivan from SpliceVarDB to\ bigLolly format, and the UCSC\ Browser staff downloaded it for display.\
\ \Precomputed AbSplice-DNA scores in all 49 GTEx tissues are available at\ \ Zenodo.
\ \ License\\ The SpliceAI data is not available for download from the Genome Browser.\ The raw data can be found directly on\ Illumina.\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are\ made available by Illumina only for academic or not-for-profit research only.\ By accessing the SpliceAI data, you acknowledge and agree that you may only\ use this data for your own personal academic or not-for-profit research only,\ and not for any other purposes. You may not use this data for any for-profit,\ clinical, or other commercial purpose without obtaining a commercial license\ from Illumina, Inc.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. For automated analysis, the data may\ be queried from our REST API.
\ \\
For automated download and analysis, the genome annotation is stored in a bigBed or a bigWig file\
that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools, e.g.\
\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg19/splicevardb/SVADB.bb\
\ -chrom=chr21 -start=0 -end=100000000 stdout\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500\
\ http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/spliceAi/wildtype/spliceAiAcceptorMinus.bw\
\ stdout\
\
\
These tools can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.
Thanks to Illumina for making SpliceAI available, both the model and the precomputed data files.
\ \Thanks to Francois Lecoquierre from the University of Oxford, Jean-Madeleine de Sainte Agathe\ from Institut Pasteur Paris, and Michael Hiller from the Senckenberg Museum Frankfurt for\ suggesting and then creating the SpliceAI Wildtype annotations.
\ \Thanks to Nils Wagner for helpful comments and suggestions for the AbSplice track.
\ \Thanks to the SpliceVarDB team for converting the data into our data formats.
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Sullivan PJ, Quinn JMW, Wu W, Pinese M, Cowley MJ.\ \ SpliceVarDB: A comprehensive database of experimentally validated human splicing variants.\ Am J Hum Genet. 2024 Oct 3;111(10):2164-2175.\ PMID: 39226898; PMC: PMC11480807\
\ \\ Wagner N, Çelik MH, Hölzlwimmer FR, Mertes C, Prokisch H, Yépez VA, Gagneur J.\ \ Aberrant splicing prediction across human tissues.\ Nat Genet. 2023 May;55(5):861-870.\ PMID: 37142848\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ \ phenDis 0 compositeTrack on\ group phenDis\ html spliceImpactSuper\ longLabel SpliceAI Wildtype: Splicing of the reference genome sequence\ parent spliceImpactSuper on\ priority 2\ shortLabel SpliceAI Wildtype\ track spliceAIWt\ type bigWig\ visibility full\ recount3_tcga TCGA bigBed 9 + recount3 TCGA introns 0 2 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/recount3/tcgav2.bb\ filter.readcount 10000:2000000000\ filter.size 30:100000\ filterByRange.readcount on\ filterByRange.size on\ filterLabel.readcount Filter by supporting split reads\ filterLabel.size Filter by intron size\ filterLabel.sjPair splice junctions (format GT/AG)\ filterLabel.strand Strand\ filterLimits.readcount 0:2000000000\ filterText.sjPair *\ filterType.sjPair wildcard\ filterType.strand multiple\ filterValues.strand +,-,.\ iframeOptions height='300' width='1000' scrolling='yes'\ iframeUrl https://snaptron.cs.jhu.edu/snaptron-studies/jxn2studies?compilation=tcgav2&jid=$$&coords=$S:${-$}\ itemRgb on\ labelFields none\ longLabel recount3 TCGA introns\ mouseOver Split read count: $readcount\ The FANTOM5 track shows mapped transcription start sites (TSS) and their usage in primary cells,\ cell lines, and tissues to produce a comprehensive overview of gene expression across the human\ body by using single molecule sequencing.\
\ \Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \Individual biological states are profiled by HeliScopeCAGE, which is a variation of the CAGE\ (Cap Analysis Gene Expression) protocol based on a single molecule sequencer. The standard protocol\ requiring 5 µg of total RNA as a starting material is referred to as hCAGE, and an\ optimized version for a lower quantity (~ 100 ng) is referred to as LQhCAGE (Kanamori-Katyama\ et al. 2011).\
Transcription start sites (TSSs) were mapped and their usage in human and mouse primary cells,\ cell lines, and tissues was to produce a comprehensive overview of mammalian gene expression across the\ human body. 5′-end of the mapped CAGE reads are counted at a single base pair resolution\ (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the\ sample. Individual samples shown in "TSS activity" tracks are grouped as below.\
TSS (CAGE) peaks across the panel of the biological states (samples) are identified by DPI\ (decomposition based peak identification, Forrest et al. 2014), where each of the peaks consists of\ neighboring and related TSSs. The peaks are used as anchors to define promoters and units of\ promoter-level expression analysis. Two subsets of the peaks are defined based on evidence of read\ counts, depending on scopes of subsequent analyses, and the first subset (referred as a\ robust set of the peaks, thresholded for expression analysis is shown as TSS peaks. They are\ named "p#@GENE_SYMBOL" if associated with 5'-end of known genes, or "p@CHROM:START..END,STRAND"\ otherwise. The summary tracks consist of the TSS (CAGE) peaks and summary profiles of TSS\ activities (total and maximum values). The summary track consists of the following tracks.\
\ 5′-end of the mapped CAGE reads are counted at a single base pair resolution (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the sample. The read counts tracks indicate raw counts of CAGE reads, and the TPM tracks indicate normalized counts as TPM (tags per million).\
\ \\ FANTOM5 data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ The FANTOM5 reprocessed data can be found and downloaded on the FANTOM website.
\ \\ Thanks to the FANTOM5 consortium,\ the Large Scale Data Managing Unit and Preventive Medicine and\ Applied Genomics Unit, the Center for Integrative Medical Sciences (IMS), and\ RIKEN for providing this data\ and its analysis.
\ \\ FANTOM Consortium and the RIKEN PMI and CLST (DGT), Forrest AR, Kawaji H, Rehli M, Baillie JK, de\ Hoon MJ, Haberle V, Lassmann T, Kulakovskiy IV, Lizio M et al.\ \ A promoter-level mammalian expression atlas.\ Nature. 2014 Mar 27;507(7493):462-70.\ PMID: 24670764; PMC: PMC4529748\
\ \\ Kanamori-Katayama M, Itoh M, Kawaji H, Lassmann T, Katayama S, Kojima M, Bertin N, Kaiho A, Ninomiya\ N, Daub CO et al.\ \ Unamplified cap analysis of gene expression on a single-molecule sequencer.\ Genome Res. 2011 Jul;21(7):1150-9.\ PMID: 21596820; PMC: PMC3129257\
\ \\ Lizio M, Harshbarger J, Shimoji H, Severin J, Kasukawa T, Sahin S, Abugessaisa I, Fukuda S, Hori F,\ Ishikawa-Kato S et al.\ \ Gateways to the FANTOM5 promoter level mammalian expression atlas.\ Genome Biol. 2015 Jan 5;16(1):22.\ PMID: 25723102; PMC: PMC4310165\
\ regulation 0 boxedCfg on\ compositeTrack on\ dataVersion FANTOM5 reprocessed7\ dimensions dimX=sequenceTech dimY=category dimA=strand\ html fantom5.html\ longLabel FANTOM5: TSS activity per sample read counts\ priority 2\ shortLabel TSS activity - read counts\ showSubtrackColorOnUi off\ sortOrder category=+ sequenceTech=+\ subGroup1 sequenceTech Sequence_Tech hCAGE=hCAGE LQhCAGE=LQhCAGE\ subGroup2 category Category cellLine=cellLine fractionation=fractionation primaryCell=primaryCell tissue=tissue AoSMC_response_to_FGF2=AoSMC_response_to_FGF2_timecourse AoSMC_response_to_IL1b=AoSMC_response_to_IL1b_timecourse ES_to_cardiomyocyte=ES_to_cardiomyocyte_timecourse Embryoid_body_to_melanocyte=Embryoid_body_to_melanocyte_timecourse Epithelial_to_mesenchymal=Epithelial_to_mesenchymal_timecourse Human_iPS_to_neuron_Downs_syndrome_1=Human_iPS_to_neuron_Downs_syndrome_1_timecourse Human_iPS_to_neuron_Downs_syndrome_2=Human_iPS_to_neuron_Downs_syndrome_2_timecourse Human_iPS_to_neuron_wt_1=Human_iPS_to_neuron_wt_1_timecourse Human_iPS_to_neuron_wt_2=Human_iPS_to_neuron_wt_2_timecourse Lymphatic_EC_response_to_VEGFC=Lymphatic_EC_response_to_VEGFC_timecourse MCF7_response_to_EGF=MCF7_response_to_EGF_timecourse MCF7_response_to_HRG=MCF7_response_to_HRG_timecourse MSC_to_adipocyte_human=MSC_to_adipocyte_human_timecourse Macrophage_influenza_infection=Macrophage_influenza_infection_timecourse Macrophage_response_to_LPS=Macrophage_response_to_LPS_timecourse Myoblast_to_myotube_wt_and_DMD=Myoblast_to_myotube_wt_and_DMD_timecourse Preadipocyte_to_adipocyte=Preadipocyte_to_adipocyte_timecourse Rinderpest_infection_series=Rinderpest_infection_series_timecourse Saos_calcification=Saos_calcification_timecourse timecourse=other_samples_in_timecourse\ subGroup3 strand Strand forward=forward reverse=reverse\ superTrack fantom5\ track TSS_activity_read_counts\ type bigWig\ visibility hide\ umap36 Umap S36 bigBed 6 Single-read mappability with 36-mers 0 2 80 70 240 167 162 247 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k36.Unique.Mappability.bb\ color 80,70,240\ longLabel Single-read mappability with 36-mers\ parent umapBigBed off\ priority 2\ shortLabel Umap S36\ subGroups view=SR\ track umap36\ visibility hide\ cpgIslandExtUnmasked Unmasked CpG bed 4 + CpG Islands on All Sequence (Islands < 300 Bases are Light Green) 0 2 0 100 0 128 228 128 0 0 0CpG islands are associated with genes, particularly housekeeping\ genes, in vertebrates. CpG islands are typically common near\ transcription start sites and may be associated with promoter\ regions. Normally a C (cytosine) base followed immediately by a \ G (guanine) base (a CpG) is rare in\ vertebrate DNA because the Cs in such an arrangement tend to be\ methylated. This methylation helps distinguish the newly synthesized\ DNA strand from the parent strand, which aids in the final stages of\ DNA proofreading after duplication. However, over evolutionary time,\ methylated Cs tend to turn into Ts because of spontaneous\ deamination. The result is that CpGs are relatively rare unless\ there is selective pressure to keep them or a region is not methylated\ for some other reason, perhaps having to do with the regulation of gene\ expression. CpG islands are regions where CpGs are present at\ significantly higher levels than is typical for the genome as a whole.
\ \\ The unmasked version of the track displays potential CpG islands\ that exist in repeat regions and would otherwise not be visible\ in the repeat masked version.\
\ \\ By default, only the masked version of the track is displayed. To view the\ unmasked version, change the visibility settings in the track controls at\ the top of this page.\
\ \CpG islands were predicted by searching the sequence one base at a\ time, scoring each dinucleotide (+17 for CG and -1 for others) and\ identifying maximally scoring segments. Each segment was then\ evaluated for the following criteria:\ \
\ The entire genome sequence, masking areas included, was\ used for the construction of the track Unmasked CpG.\ The track CpG Islands is constructed on the sequence after\ all masked sequence is removed.\
\ \The CpG count is the number of CG dinucleotides in the island. \ The Percentage CpG is the ratio of CpG nucleotide bases\ (twice the CpG count) to the length. The ratio of observed to expected \ CpG is calculated according to the formula (cited in \ Gardiner-Garden et al. (1987)):\ \
Obs/Exp CpG = Number of CpG * N / (Number of C * Number of G)\ \ where N = length of sequence.\
\ The calculation of the track data is performed by the following command sequence:\
\
twoBitToFa assembly.2bit stdout | maskOutFa stdin hard stdout \\\
| cpg_lh /dev/stdin 2> cpg_lh.err \\\
| awk '{$2 = $2 - 1; width = $3 - $2; printf("%s\\t%d\\t%s\\t%s %s\\t%s\\t%s\\t%0.0f\\t%0.1f\\t%s\\t%s\\n", $1, $2, $3, $5, $6, width, $6, width*$7*0.01, 100.0*2*$6/width, $7, $9);}' \\\
| sort -k1,1 -k2,2n > cpgIsland.bed\
\
The unmasked track data is constructed from\
twoBitToFa -noMask output for the twoBitToFa command.\
\
\
\ CpG islands and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator.\ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\\ The source for the cpg_lh program can be obtained from\ src/utils/cpgIslandExt/.\ The cpg_lh program binary can be obtained from: http://hgdownload.soe.ucsc.edu/admin/exe/linux.x86_64/cpg_lh (choose "save file")\
\ \This track was generated using a modification of a program developed by G. Micklem and L. Hillier \ (unpublished).
\ \\ Gardiner-Garden M, Frommer M.\ \ CpG islands in vertebrate genomes.\ J Mol Biol. 1987 Jul 20;196(2):261-82.\ PMID: 3656447\
\ regulation 1 html cpgIslandSuper\ longLabel CpG Islands on All Sequence (Islands < 300 Bases are Light Green)\ parent cpgIslandSuper hide\ priority 2\ shortLabel Unmasked CpG\ track cpgIslandExtUnmasked\ cons241way Zoonomia 241 Placent bed 4 Zoonomia Alignment - 241 Placental Mammal Genomes aligned by the Zoonomia Project with Cactus 0 2 0 0 0 127 127 127 0 0 0\ Downloads for data in this track are available:\
Warning: Unlike other alignment tracks on the genome browser, this one does not show\ insertions in the query genomes. Also, all other alignment tracks show one query\ genome sequence for each target genome sequence, but in this track, each\ target genome sequence can be aligned to multiple query genome sequences.\ Only the first sequence is shown on the genome browser itself, the others are shown on the details page,\ when one clicks on the alignment. If you are interested in this track and want\ these shortcomings to be fixed, please contact us.\
\ \\ This track shows multiple alignments of 241 vertebrate\ species and measurements of evolutionary conservation\ from the Zoonomia Project.\
\ \\ The multiple alignments were generated using the\ Cactus comparative genomics alignment system.\ Cactus produces reference-free, whole-genome multiple alignments.\
\ \ \
\ The base-wise conservation scores are computed using phyloP from the\ PHAST package, for all species.\ This version was prepared by Michael Dong (Uppsala U) with an improved neutral\ model incorporating better versions of ancestral repeats.\
\ \\ For genome assemblies not available in the genome browser, there are\ alternative assembly hub genome browsers. Missing sequence in any assembly is\ highlighted in the track display by regions of yellow when zoomed out and by\ Ns when displayed at base level (see Gap Annotation, below).
\\
\\ \\
\ count \common \
nameCLADE \group \scientific \
namesequencing \
sourceNCBI \
assemblyspecies \
status\ 1 \Cape golden mole \AFROSORICIDA \Chrysochloridae \Chrysochloris asiatica \1. Zoonomia \GCA_004027935.1 \LC \\ 2 \Small madagascar hedgehog \AFROSORICIDA \Tenrecidae \Echinops telfairi \2. Existing assembly \GCF_000313985.1 \LC \\ 3 \Talazac's shrew tenrec \AFROSORICIDA \Tenrecidae \Microgale talazaci \1. Zoonomia \GCA_004026705.1 \LC \\ 4 \Cheetah \CARNIVORA \Felidae \Acinonyx jubatus \2. Existing assembly \GCF_001443585.1 \CR \\ 5 \Giant panda \CARNIVORA \Ursidae \Ailuropoda melanoleuca \2. Existing assembly \GCA_002007445.1 \VU \\ 6 \Lesser panda \CARNIVORA \Ailuridae \Ailurus fulgens \2. Existing assembly \GCA_002007465.1 \EN \\ 7 \Domestic dog \CARNIVORA \Canidae \Canis lupus familiaris \2. Existing assembly \GCF_000002285.3 \LC \\ 8 \Domestic dog (village dog) \CARNIVORA \Canidae \Canis lupus familiaris \1. Zoonomia \GCA_004027395.1 \LC \\ 9 \Fossa \CARNIVORA \Eupleridae \Cryptoprocta ferox \1. Zoonomia \GCA_004023885.1 \VU \\ 10 \Sea otter \CARNIVORA \Mustelidae \Enhydra lutris \2. Existing assembly \GCF_002288905.1 \EN \\ 11 \Domestic cat \CARNIVORA \Felidae \Felis catus \2. Existing assembly \GCF_000181335.2 \LC \\ 12 \Black-footed cat \CARNIVORA \Felidae \Felis nigripes \1. Zoonomia \GCA_004023925.1 \VU \\ 13 \Dwarf mongoose \CARNIVORA \Herpestidae \Helogale parvula \1. Zoonomia \GCA_004023845.1 \LC \\ 14 \Striped hyena \CARNIVORA \Hyaenidae \Hyaena hyaena \1. Zoonomia \GCA_004023945.1 \NT \\ 15 \Weddell seal \CARNIVORA \Phocidae \Leptonychotes weddellii \2. Existing assembly \GCF_000349705.1 \LC \\ 16 \African hunting dog \CARNIVORA \Canidae \Lycaon pictus \2. Existing assembly \GCA_001887905.1 \EN \\ 17 \Honey badger \CARNIVORA \Mustelidae \Mellivora capensis \1. Zoonomia \GCA_004024625.1 \LC \\ 18 \Northern elephant seal \CARNIVORA \Phocidae \Mirounga angustirostris \1. Zoonomia \GCA_004023865.1 \LC \\ 19 \South African banded mongoose \CARNIVORA \Herpestidae \Mungos mungo \1. Zoonomia \GCA_004023785.1 \LC \\ 20 \Domestic ferret \CARNIVORA \Mustelidae \Mustela putorius \2. Existing assembly \GCF_000239315.1 \LC \\ 21 \Hawaiian monk seal \CARNIVORA \Phocidae \Neomonachus schauinslandi \2. Existing assembly \GCA_002201575.1 \EN \\ 22 \Pacific walrus \CARNIVORA \Odobenidae \Odobenus rosmarus \2. Existing assembly \GCF_000321225.1 \DD \\ 23 \Jaguar \CARNIVORA \Felidae \Panthera onca \1. Zoonomia \GCA_004023805.1 \NT \\ 24 \Leopard \CARNIVORA \Felidae \Panthera pardus \2. Existing assembly \GCA_001857705.1 \VU \\ 25 \Amur tiger \CARNIVORA \Felidae \Panthera tigris \2. Existing assembly \GCF_000464555.1 \EN \\ 26 \Asian palm civet \CARNIVORA \Viverridae \Paradoxurus hermaphroditus \1. Zoonomia \GCA_004024585.1 \LC \\ 27 \Giant otter \CARNIVORA \Mustelidae \Pteronura brasiliensis \1. Zoonomia \GCA_004024605.1 \EN \\ 28 \Puma \CARNIVORA \Felidae \Puma concolor \2. Existing assembly \GCF_003327715.1 \LC \\ 29 \Western spotted skunk \CARNIVORA \Mephitidae \Spilogale gracilis \1. Zoonomia \GCA_004023965.1 \LC \\ 30 \Meerkat \CARNIVORA \Herpestidae \Suricata suricatta \1. Zoonomia \GCA_004023905.1 \LC \\ 31 \Polar bear \CARNIVORA \Ursidae \Ursus maritimus \2. Existing assembly \GCF_000687225.1 \VU \\ 32 \Arctic fox \CARNIVORA \Canidae \Vulpes lagopus \1. Zoonomia \GCA_004023825.1 \LC \\ 33 \California sea lion \CARNIVORA \Otariidae \Zalophus californianus \1. Zoonomia \GCA_004024565.1 \LC \\ 34 \Aoudad \CETARTIODACTYLA \Bovidae \Ammotragus lervia \2. Existing assembly \GCA_002201775.1 \VU \\ 35 \Pronghorn \CETARTIODACTYLA \Antilocapridae \Antilocapra americana \1. Zoonomia \GCA_004027515.1 \LC \\ 36 \Minke whale \CETARTIODACTYLA \Balaenopteridae \Balaenoptera acutorostrata \2. Existing assembly \GCF_000493695.1 \LC \\ 37 \Antarctic minke whale \CETARTIODACTYLA \Balaenopteridae \Balaenoptera bonaerensis \2. Existing assembly \GCA_000978805.1 \DD \\ 38 \Hirola \CETARTIODACTYLA \Bovidae \Beatragus hunteri \1. Zoonomia \GCA_004027495.1 \CR \\ 39 \American bison \CETARTIODACTYLA \Bovidae \Bison bison \2. Existing assembly \GCF_000754665.1 \NT \\ 40 \Zebu cattle \CETARTIODACTYLA \Bovidae \Bos indicus \2. Existing assembly \GCA_000247795.2 \LC \\ 41 \Wild yak \CETARTIODACTYLA \Bovidae \Bos mutus \2. Existing assembly \GCF_000298355.1 \VU \\ 42 \Cattle \CETARTIODACTYLA \Bovidae \Bos taurus \2. Existing assembly \GCF_000003205.7 \LC \\ 43 \Water buffalo \CETARTIODACTYLA \Bovidae \Bubalus bubalis \2. Existing assembly \GCF_000471725.1 \LC \\ 44 \Bactrian camel \CETARTIODACTYLA \Camelidae \Camelus bactrianus \2. Existing assembly \GCF_000767855.1 \LC \\ 45 \Arabian camel \CETARTIODACTYLA \Camelidae \Camelus dromedarius \2. Existing assembly \GCF_000767585.1 \LC \\ 46 \Wild bactrian camel \CETARTIODACTYLA \Camelidae \Camelus ferus \2. Existing assembly \GCF_000311805.1 \CR \\ 47 \Wild goat \CETARTIODACTYLA \Bovidae \Capra aegagrus \2. Existing assembly \GCA_000978405.1 \VU \\ 48 \Goat \CETARTIODACTYLA \Bovidae \Capra hircus \2. Existing assembly \GCF_001704415.1 \LC \\ 49 \Chacoan peccary \CETARTIODACTYLA \Tayassuidae \Catagonus wagneri \1. Zoonomia \GCA_004024745.1 \EN \\ 50 \Beluga whale \CETARTIODACTYLA \Monodontidae \Delphinapterus leucas \2. Existing assembly \GCF_002288925.1 \LC \\ 51 \Pere david's deer \CETARTIODACTYLA \Cervidae \Elaphurus davidianus \2. Existing assembly \GCA_002443075.1 \CR \\ 52 \Grey whale \CETARTIODACTYLA \Eschrichtiidae \Eschrichtius robustus \1. Zoonomia \GCA_004363415.1 \LC \\ 53 \North Pacific right whale \CETARTIODACTYLA \Balaenidae \Eubalaena japonica \1. Zoonomia \GCA_004363455.1 \EN \\ 54 \Giraffe \CETARTIODACTYLA \Giraffidae \Giraffa tippelskirchi \2. Existing assembly \GCA_001651235.1 \VU \\ 55 \Nilgiri tahr \CETARTIODACTYLA \Bovidae \Hemitragus hylocrius \1. Zoonomia \GCA_004026825.1 \EN \\ 56 \Hippopotamus \CETARTIODACTYLA \Hippopotamidae \Hippopotamus amphibius \1. Zoonomia \GCA_004027065.1 \VU \\ 57 \Amazon river dolphin \CETARTIODACTYLA \Iniidae \Inia geoffrensis \1. Zoonomia \GCA_004363515.1 \DD \\ 58 \Pygmy sperm whale \CETARTIODACTYLA \Physeteridae \Kogia breviceps \1. Zoonomia \GCA_004363705.1 \DD \\ 59 \Yangtze river dolphin \CETARTIODACTYLA \Iniidae \Lipotes vexillifer \2. Existing assembly \GCF_000442215.1 \CR \\ 60 \Sowerby's beaked whale \CETARTIODACTYLA \Ziphiidae \Mesoplodon bidens \1. Zoonomia \GCA_004027085.1 \DD \\ 61 \Narwhal \CETARTIODACTYLA \Monodontidae \Monodon monoceros \1. Zoonomia \GCA_004026685.1 \LC \\ 62 \Siberian musk deer \CETARTIODACTYLA \Moschidae \Moschus moschiferus \1. Zoonomia \GCA_004024705.1 \VU \\ 63 \Yangtze finless porpoise \CETARTIODACTYLA \Phocoenidae \Neophocaena asiaeorientalis \2. Existing assembly \GCA_003031525.1 \EN \\ 64 \White-tailed deer \CETARTIODACTYLA \Cervidae \Odocoileus virginianus \2. Existing assembly \GCA_002102435.1 \LC \\ 65 \Okapi \CETARTIODACTYLA \Giraffidae \Okapia johnstoni \2. Existing assembly \GCA_001660835.1 \EN \\ 66 \Killer whale \CETARTIODACTYLA \Delphinidae \Orcinus orca \2. Existing assembly \GCF_000331955.2 \DD \\ 67 \Sheep \CETARTIODACTYLA \Bovidae \Ovis aries \2. Existing assembly \GCF_000298735.2 \LC \\ 68 \Peninsular bighorn sheep \CETARTIODACTYLA \Bovidae \Ovis canadensis cremnobates \1. Zoonomia \GCA_004026945.1 \EN \\ 69 \Chiru \CETARTIODACTYLA \Bovidae \Pantholops hodgsonii \2. Existing assembly \GCF_000400835.1 \NT \\ 70 \Harbor porpoise \CETARTIODACTYLA \Phocoenidae \Phocoena phocoena \1. Zoonomia \GCA_004363495.1 \LC \\ 71 \Indus river dolphin \CETARTIODACTYLA \Platanistidae \Platanista gangetica minor \1. Zoonomia \GCA_004363435.1 \EN \\ 72 \Siberian reindeer \CETARTIODACTYLA \Cervidae \Rangifer tarandus \1. Zoonomia \GCA_004026565.1 \VU \\ 73 \Russian saiga \CETARTIODACTYLA \Bovidae \Saiga tatarica tatarica \1. Zoonomia \GCA_004024985.1 \CR \\ 74 \Pig \CETARTIODACTYLA \Suidae \Sus scrofa \2. Existing assembly \GCF_000003025.5 \LC \\ 75 \Java lesser chevrotain \CETARTIODACTYLA \Tragulidae \Tragulus javanicus \1. Zoonomia \GCA_004024965.1 \DD \\ 76 \Bottlenose dolphin \CETARTIODACTYLA \Delphinidae \Tursiops truncatus \2. Existing assembly \GCA_001922835.1 \LC \\ 77 \Alpaca \CETARTIODACTYLA \Camelidae \Vicugna pacos \2. Existing assembly \GCA_000767525.1 \LC \\ 78 \Cuvier's beaked whale \CETARTIODACTYLA \Ziphiidae \Ziphius cavirostris \1. Zoonomia \GCA_004364475.1 \LC \\ 79 \Tailed tailless bat \CHIROPTERA \Phyllostomidae \Anoura caudifer \1. Zoonomia \GCA_004027475.1 \LC \\ 80 \Jamacian fruit-eating bat \CHIROPTERA \Phyllostomidae \Artibeus jamaicensis \1. Zoonomia \GCA_004027435.1 \LC \\ 81 \Seba's short-tailed bat \CHIROPTERA \Phyllostomidae \Carollia perspicillata \1. Zoonomia \GCA_004027735.1 \LC \\ 82 \Bumblebee bat \CHIROPTERA \Craseonycteridae \Craseonycteris thonglongyai \1. Zoonomia \GCA_004027555.1 \VU \\ 83 \Common vampire bat \CHIROPTERA \Phyllostomidae \Desmodus rotundus \2. Existing assembly \GCA_002940915.2 \LC \\ 84 \Straw-colored fruit bat \CHIROPTERA \Pteropodidae \Eidolon helvum \2. Existing assembly \GCA_000465285.1 \NT \\ 85 \Big brown bat \CHIROPTERA \Vespertilionidae \Eptesicus fuscus \2. Existing assembly \GCF_000308155.1 \LC \\ 86 \Great roundleaf bat \CHIROPTERA \Hipposideridae \Hipposideros armiger \2. Existing assembly \GCA_001890085.1 \LC \\ 87 \Cantor's leaf-nosed bat \CHIROPTERA \Hipposideridae \Hipposideros galeritus \1. Zoonomia \GCA_004027415.1 \LC \\ 88 \Eastern red bat \CHIROPTERA \Vespertilionidae \Lasiurus borealis \1. Zoonomia \GCA_004026805.1 \LC \\ 89 \Long-tongued fruit bat \CHIROPTERA \Pteropodidae \Macroglossus sobrinus \1. Zoonomia \GCA_004027375.1 \LC \\ 90 \Greater false vampire bat \CHIROPTERA \Megadermatidae \Megaderma lyra \1. Zoonomia \GCA_004026885.1 \LC \\ 91 \Hairy big-eared bat \CHIROPTERA \Phyllostomidae \Micronycteris hirsuta \1. Zoonomia \GCA_004026765.1 \LC \\ 92 \Natal long-fingered bat \CHIROPTERA \Vespertilionidae \Miniopterus natalensis \2. Existing assembly \GCF_001595765.1 \LC \\ 93 \Common bent-wing bat \CHIROPTERA \Vespertilionidae \Miniopterus schreibersii \1. Zoonomia \GCA_004026525.1 \NT \\ 94 \Ghost-faced bat \CHIROPTERA \Mormoopidae \Mormoops blainvillei \1. Zoonomia \GCA_004026545.1 \LC \\ 95 \Ashy-gray tube-nosed bat \CHIROPTERA \Vespertilionidae \Murina feae \1. Zoonomia \GCA_004026665.1 \LC \\ 96 \Brandt's bat \CHIROPTERA \Vespertilionidae \Myotis brandtii \2. Existing assembly \GCF_000412655.1 \LC \\ 97 \David's myotis bat \CHIROPTERA \Vespertilionidae \Myotis davidii \2. Existing assembly \GCF_000327345.1 \LC \\ 98 \Little brown bat \CHIROPTERA \Vespertilionidae \Myotis lucifugus \2. Existing assembly \GCF_000147115.1 \LC \\ 99 \Greater mouse-eared bat \CHIROPTERA \Vespertilionidae \Myotis myotis \1. Zoonomia \GCA_004026985.1 \LC \\ 100 \Greater bulldog bat \CHIROPTERA \Noctilionidae \Noctilio leporinus \1. Zoonomia \GCA_004026585.1 \LC \\ 101 \Common pipistrelle \CHIROPTERA \Vespertilionidae \Pipistrellus pipistrellus \1. Zoonomia \GCA_004026625.1 \LC \\ 102 \Parnell's mustached bat \CHIROPTERA \Mormoopidae \Pteronotus parnellii \2. Existing assembly \GCA_000465405.1 \LC \\ 103 \Black flying fox \CHIROPTERA \Pteropodidae \Pteropus alecto \2. Existing assembly \GCF_000325575.1 \LC \\ 104 \Large flying fox \CHIROPTERA \Pteropodidae \Pteropus vampyrus \2. Existing assembly \GCF_000151845.1 \NT \\ 105 \Chinese rufous horseshoe bat \CHIROPTERA \Rhinolophidae \Rhinolophus sinicus \2. Existing assembly \GCA_001888835.1 \LC \\ 106 \Egyptian fruit bat \CHIROPTERA \Pteropodidae \Rousettus aegyptiacus \1. Zoonomia \GCA_004024865.1 \LC \\ 107 \Mexican free-tailed bat \CHIROPTERA \Molossidae \Tadarida brasiliensis \1. Zoonomia \GCA_004025005.1 \LC \\ 108 \Stripe-headed round-eared bat \CHIROPTERA \Phyllostomidae \Tonatia saurophila \1. Zoonomia \GCA_004024845.1 \LC \\ 109 \Screaming hairy armadillo \CINGULATA \Dasypodidae \Chaetophractus vellerosus \1. Zoonomia \GCA_004027955.1 \LC \\ 110 \Nine-banded armadillo \CINGULATA \Dasypodidae \Dasypus novemcinctus \2. Existing assembly \GCF_000208655.1 \LC \\ 111 \Southern three-banded armadillo \CINGULATA \Dasypodidae \Tolypeutes matacus \1. Zoonomia \GCA_004025125.1 \NT \\ 112 \Sunda flying lemur \DERMOPTERA \Cynocephalidae \Galeopterus variegatus \1. Zoonomia \GCA_004027255.1 \LC \\ 113 \Star-nosed mole \EULIPOTYPHLA \Talpidae \Condylura cristata \2. Existing assembly \GCF_000260355.1 \LC \\ 114 \Indochinese shrew \EULIPOTYPHLA \Soricidae \Crocidura indochinensis \1. Zoonomia \GCA_004027635.1 \LC \\ 115 \Western european hedgehog \EULIPOTYPHLA \Erinaceidae \Erinaceus europaeus \2. Existing assembly \GCF_000296755.1 \LC \\ 116 \Eastern mole \EULIPOTYPHLA \Talpidae \Scalopus aquaticus \1. Zoonomia \GCA_004024925.1 \LC \\ 117 \Hispaniolan solenodon \EULIPOTYPHLA \Solenodontidae \Solenodon paradoxus \1. Zoonomia \GCA_004363575.1 \EN \\ 118 \European shrew \EULIPOTYPHLA \Soricidae \Sorex araneus \2. Existing assembly \GCF_000181275.1 \LC \\ 119 \Gracile shrew-like mole \EULIPOTYPHLA \Talpidae \Uropsilus gracilis \1. Zoonomia \GCA_004024945.1 \LC \\ 120 \African yellow-spotted rock hyrax \HYRACOIDEA \Procaviidae \Heterohyrax brucei \1. Zoonomia \GCA_004026845.1 \LC \\ 121 \South African rock hyrax \HYRACOIDEA \Procaviidae \Procavia capensis \1. Zoonomia \GCA_004026925.1 \LC \\ 122 \Snowshoe hare \LAGOMORPHA \Leporidae \Lepus americanus \1. Zoonomia \GCA_004026855.1 \LC \\ 123 \American pika \LAGOMORPHA \Ochotonidae \Ochotona princeps \2. Existing assembly \GCF_000292845.1 \LC \\ 124 \Rabbit \LAGOMORPHA \Leporidae \Oryctolagus cuniculus \2. Existing assembly \GCF_000003625.3 \NT \\ 125 \Cape elephant shrew \MACROSCELIDEA \Macroscelididae \Elephantulus edwardii \1. Zoonomia \GCA_004027355.1 \LC \\ 126 \Southern white rhinoceros \PERISSODACTYLA \Rhinocerotidae \Ceratotherium simum \2. Existing assembly \GCF_000283155.1 \NT \\ 127 \Northern white rhino \PERISSODACTYLA \Rhinocerotidae \Ceratotherium simum cottoni \1. Zoonomia \GCA_004027795.1 \CR \\ 128 \Sumatran rhinoceros \PERISSODACTYLA \Rhinocerotidae \Dicerorhinus sumatrensis \2. Existing assembly \GCA_002844835.1 \CR \\ 129 \Black rhinocerous \PERISSODACTYLA \Rhinocerotidae \Diceros bicornis \1. Zoonomia \GCA_004027315.1 \CR \\ 130 \Ass \PERISSODACTYLA \Equidae \Equus asinus \2. Existing assembly \GCF_001305755.1 \LC \\ 131 \Horse \PERISSODACTYLA \Equidae \Equus caballus \2. Existing assembly \GCF_000002305.2 \LC \\ 132 \Przewalski's horse \PERISSODACTYLA \Equidae \Equus przewalskii \2. Existing assembly \GCF_000696695.1 \EN \\ 133 \Malayan tapir \PERISSODACTYLA \Tapiridae \Tapirus indicus \1. Zoonomia \GCA_004024905.1 \EN \\ 134 \South American tapir \PERISSODACTYLA \Tapiridae \Tapirus terrestris \1. Zoonomia \GCA_004025025.1 \VU \\ 135 \Malayan pangolin \PHOLIDOTA \Manidae \Manis javanica \2. Existing assembly \GCF_001685135.1 \CR \\ 136 \Chinese pangolin \PHOLIDOTA \Manidae \Manis pentadactyla \2. Existing assembly \GCA_000738955.1 \CR \\ 137 \Linnaeus's two toed sloth \PILOSA \Megalonychidae \Choloepus didactylus \1. Zoonomia \GCA_004027855.1 \LC \\ 138 \Hoffmann's two-fingered sloth \PILOSA \Megalonychidae \Choloepus hoffmanni \2. Existing assembly \GCA_000164785.2 \LC \\ 139 \Giant anteater \PILOSA \Myrmecophagidae \Myrmecophaga tridactyla \1. Zoonomia \GCA_004026745.1 \VU \\ 140 \Southern tamandua \PILOSA \Myrmecophagidae \Tamandua tetradactyla \1. Zoonomia \GCA_004025105.1 \LC \\ 141 \Mexican howler monkey \PRIMATES \Atelidae \Alouatta palliata mexicana \1. Zoonomia \GCA_004027835.1 \CR \\ 142 \Ma's night monkey \PRIMATES \Aotidae \Aotus nancymaae \2. Existing assembly \GCA_000952055.2 \VU \\ 143 \Geoffroy's spider monkey \PRIMATES \Atelidae \Ateles geoffroyi \1. Zoonomia \GCA_004024785.1 \EN \\ 144 \White-eared titi \PRIMATES \Pitheciidae \Callicebus donacophilus \1. Zoonomia \GCA_004027715.1 \LC \\ 145 \White-tufted-ear marmoset \PRIMATES \Cebidae \Callithrix jacchus \2. Existing assembly \GCA_002754865.1 \LC \\ 146 \White-fronted capuchin \PRIMATES \Cebidae \Cebus albifrons \1. Zoonomia \GCA_004027755.1 \LC \\ 147 \White-faced sapajou \PRIMATES \Cebidae \Cebus capucinus \2. Existing assembly \GCF_001604975.1 \LC \\ 148 \Sooty mangabey \PRIMATES \Cercopithecidae \Cercocebus atys \2. Existing assembly \GCF_000955945.1 \NT \\ 149 \De brazza's monkey \PRIMATES \Cercopithecidae \Cercopithecus neglectus \1. Zoonomia \GCA_004027615.1 \LC \\ 150 \Fat-tailed dwarf lemur \PRIMATES \Cheirogaleidae \Cheirogaleus medius \1. Zoonomia \GCA_004024725.1 \LC \\ 151 \Green monkey \PRIMATES \Cercopithecidae \Chlorocebus sabaeus \2. Existing assembly \GCF_000409795.2 \LC \\ 152 \Angolan colobus \PRIMATES \Cercopithecidae \Colobus angolensis \2. Existing assembly \GCF_000951035.1 \VU \\ 153 \Aye-aye \PRIMATES \Daubentoniidae \Daubentonia madagascariensis \1. Zoonomia \GCA_004027145.1 \EN \\ 154 \Patas monkey \PRIMATES \Cercopithecidae \Erythrocebus patas \1. Zoonomia \GCA_004027335.1 \LC \\ 155 \Sclater's lemur \PRIMATES \Lemuridae \Eulemur flavifrons \2. Existing assembly \GCA_001262665.1 \CR \\ 156 \Common brown lemur \PRIMATES \Lemuridae \Eulemur fulvus \1. Zoonomia \GCA_004027275.1 \NT \\ 157 \Western lowland gorilla \PRIMATES \Hominidae \Gorilla gorilla \2. Existing assembly \GCA_900006655.3 \CR \\ 158 \Human \PRIMATES \Hominidae \Homo sapiens \2. Existing assembly \GCA_000001405.27 \LC \\ 159 \Indri \PRIMATES \Indridae \Indri indri \1. Zoonomia \GCA_004363605.1 \CR \\ 160 \Ring tailed lemur \PRIMATES \Lemuridae \Lemur catta \1. Zoonomia \GCA_004024665.1 \EN \\ 161 \Crab-eating macaque \PRIMATES \Cercopithecidae \Macaca fascicularis \2. Existing assembly \GCF_000364345.1 \DD \\ 162 \Rhesus monkey \PRIMATES \Cercopithecidae \Macaca mulatta \2. Existing assembly \GCF_000772875.2 \LC \\ 163 \Pig-tailed macaque \PRIMATES \Cercopithecidae \Macaca nemestrina \2. Existing assembly \GCF_000956065.1 \VU \\ 164 \Drill \PRIMATES \Cercopithecidae \Mandrillus leucophaeus \2. Existing assembly \GCF_000951045.1 \EN \\ 165 \Gray mouse lemur \PRIMATES \Cheirogaleidae \Microcebus murinus \2. Existing assembly \GCA_000165445.3 \LC \\ 166 \Coquerel's giant mouse lemur \PRIMATES \Cheirogaleidae \Mirza coquereli \1. Zoonomia \GCA_004024645.1 \EN \\ 167 \Proboscis monkey \PRIMATES \Cercopithecidae \Nasalis larvatus \1. Zoonomia \GCA_004027105.1 \EN \\ 168 \Northern white-cheeked gibbon \PRIMATES \Hylobatidae \Nomascus leucogenys \2. Existing assembly \GCF_000146795.2 \CR \\ 169 \Sunda slow loris \PRIMATES \Lorisidae \Nycticebus coucang \1. Zoonomia \GCA_004027815.1 \VU \\ 170 \Small-eared galago \PRIMATES \Galagidae \Otolemur garnettii \2. Existing assembly \GCF_000181295.1 \LC \\ 171 \Pygmy chimpanzee \PRIMATES \Hominidae \Pan paniscus \2. Existing assembly \GCF_000258655.2 \EN \\ 172 \Chimpanzee \PRIMATES \Hominidae \Pan troglodytes \2. Existing assembly \GCA_002880755.3 \EN \\ 173 \Olive baboon \PRIMATES \Cercopithecidae \Papio anubis \2. Existing assembly \GCA_000264685.2 \LC \\ 174 \Ugandan red colobus \PRIMATES \Cercopithecidae \Piliocolobus tephrosceles \2. Existing assembly \GCA_002776525.1 \EN \\ 175 \White-faced saki \PRIMATES \Pitheciidae \Pithecia pithecia \1. Zoonomia \GCA_004026645.1 \LC \\ 176 \Sumatran orangutan \PRIMATES \Hominidae \Pongo abelii \2. Existing assembly \GCA_002880775.3 \CR \\ 177 \Coquerel's sifaka \PRIMATES \Indridae \Propithecus coquereli \2. Existing assembly \GCF_000956105.1 \EN \\ 178 \Red-shanked douc \PRIMATES \Cercopithecidae \Pygathrix nemaeus \1. Zoonomia \GCA_004024825.1 \EN \\ 179 \Black snub-nosed monkey \PRIMATES \Cercopithecidae \Rhinopithecus bieti \2. Existing assembly \GCF_001698545.1 \EN \\ 180 \Golden snub-nosed monkey \PRIMATES \Cercopithecidae \Rhinopithecus roxellana \2. Existing assembly \GCF_000769185.1 \EN \\ 181 \Emperor tamarin \PRIMATES \Cebidae \Saguinus imperator \1. Zoonomia \GCA_004024885.1 \LC \\ 182 \Bolivian squirrel monkey \PRIMATES \Cebidae \Saimiri boliviensis \2. Existing assembly \GCF_000235385.1 \LC \\ 183 \Northern Plains gray langur \PRIMATES \Cercopithecidae \Semnopithecus entellus \1. Zoonomia \GCA_004025065.1 \LC \\ 184 \African savanna elephant \PROBOSCIDEA \Elephantidae \Loxodonta Africana \2. Existing assembly \GCF_000001905.1 \VU \\ 185 \Cairo spiny mouse \RODENTIA \Muridae \Acomys cahirinus \1. Zoonomia \GCA_004027535.1 \LC \\ 186 \Gobi jerboa \RODENTIA \Dipodidae \Allactaga bullata \1. Zoonomia \GCA_004027895.1 \LC \\ 187 \Mountain beaver \RODENTIA \Aplodontiidae \Aplodontia rufa \1. Zoonomia \GCA_004027875.1 \LC \\ 188 \Desmarest's hutia \RODENTIA \Capromyidae \Capromys pilorides \1. Zoonomia \GCA_004027915.1 \LC \\ 189 \North American beaver \RODENTIA \Castoridae \Castor canadensis \1. Zoonomia \GCA_004027675.1 \LC \\ 190 \Brazilian guinea pig \RODENTIA \Caviidae \Cavia aperea \2. Existing assembly \GCA_000688575.1 \LC \\ 191 \Domestic guinea pig \RODENTIA \Caviidae \Cavia porcellus \2. Existing assembly \GCF_000151735.1 \LC \\ 192 \Montane guinea pig \RODENTIA \Caviidae \Cavia tschudii \1. Zoonomia \GCA_004027695.1 \LC \\ 193 \Long-tailed chinchilla \RODENTIA \Chinchillidae \Chinchilla lanigera \2. Existing assembly \GCF_000276665.1 \EN \\ 194 \Gambian pouched rat \RODENTIA \Nesomyidae \Cricetomys gambianus \1. Zoonomia \GCA_004027575.1 \LC \\ 195 \Chinese hamster \RODENTIA \Nesomyidae \Cricetulus griseus \2. Existing assembly \GCA_900186095.1 \LC \\ 196 \Common gundi \RODENTIA \Ctenodactylidae \Ctenodactylus gundi \1. Zoonomia \GCA_004027205.1 \LC \\ 197 \Social tuco-tuco \RODENTIA \Ctenomyidae \Ctenomys sociabilis \1. Zoonomia \GCA_004027165.1 \CR \\ 198 \Lowland paca \RODENTIA \Cuniculidae \Cuniculus paca \1. Zoonomia \GCA_004365215.1 \LC \\ 199 \Central American agouti \RODENTIA \Dasyproctidae \Dasyprocta punctata \1. Zoonomia \GCA_004363535.1 \LC \\ 200 \Pacarana \RODENTIA \Dinomyidae \Dinomys branickii \1. Zoonomia \GCA_004027595.1 \LC \\ 201 \Ord's kangaroo rat \RODENTIA \Heteromyidae \Dipodomys ordii \2. Existing assembly \GCF_000151885.1 \LC \\ 202 \Stephen's kangaroo rat \RODENTIA \Heteromyidae \Dipodomys stephensi \1. Zoonomia \GCA_004024685.1 \VU \\ 203 \Patagonian mara \RODENTIA \Caviidae \Dolichotis patagonum \1. Zoonomia \GCA_004027295.1 \NT \\ 204 \Transcaucasian mole vole \RODENTIA \Cricetidae \Ellobius lutescens \2. Existing assembly \GCA_001685075.1 \LC \\ 205 \Northern mole vole \RODENTIA \Cricetidae \Ellobius talpinus \2. Existing assembly \GCA_001685095.1 \LC \\ 206 \Damara mole-rat \RODENTIA \Bathyergidae \Fukomys damarensis \2. Existing assembly \GCF_000743615.1 \LC \\ 207 \Edible dormouse \RODENTIA \Gliridae \Glis glis \1. Zoonomia \GCA_004027185.1 \LC \\ 208 \Woodland doormouse \RODENTIA \Gliridae \Graphiurus murinus \1. Zoonomia \GCA_004027655.1 \LC \\ 209 \Naked mole-rat \RODENTIA \Bathyergidae \Heterocephalus glaber \2. Existing assembly \GCF_000247695.1 \LC \\ 210 \Capybara \RODENTIA \Caviidae \Hydrochoerus hydrochaeris \1. Zoonomia \GCA_004027455.1 \LC \\ 211 \Northern crested porcupine \RODENTIA \Hystricidae \Hystrix cristata \1. Zoonomia \GCA_004026905.1 \LC \\ 212 \Thirteen-lined ground squirrel \RODENTIA \Sciuridae \Ictidomys tridecemlineatus \2. Existing assembly \GCF_000236235.1 \LC \\ 213 \Lesser egyptian jerboa \RODENTIA \Dipodidae \Jaculus jaculus \2. Existing assembly \GCF_000280705.1 \LC \\ 214 \Alpine marmot \RODENTIA \Sciuridae \Marmota marmota \2. Existing assembly \GCF_001458135.1 \LC \\ 215 \Mongolian jird \RODENTIA \Muridae \Meriones unguiculatus \1. Zoonomia \GCA_004026785.1 \LC \\ 216 \Golden hamster \RODENTIA \Cricetidae \Mesocricetus auratus \2. Existing assembly \GCF_000349665.1 \VU \\ 217 \Prairie vole \RODENTIA \Cricetidae \Microtus ochrogaster \2. Existing assembly \GCF_000317375.1 \LC \\ 218 \Ryukyu mouse \RODENTIA \Muridae \Mus caroli \2. Existing assembly \GCA_900094665.2 \LC \\ 219 \House mouse \RODENTIA \Muridae \Mus musculus \2. Existing assembly \GCF_000001635.26 \LC \\ 220 \Shrew mouse \RODENTIA \Muridae \Mus pahari \2. Existing assembly \GCA_900095145.2 \LC \\ 221 \Western wild mouse \RODENTIA \Muridae \Mus spretus \2. Existing assembly \GCA_001624865.1 \LC \\ 222 \Hazel dormouse \RODENTIA \Gliridae \Muscardinus avellanarius \1. Zoonomia \GCA_004027005.1 \LC \\ 223 \Coypu \RODENTIA \Myocastoridae \Myocastor coypus \1. Zoonomia \GCA_004027025.1 \LC \\ 224 \Upper galilee mountains blind mole rat \RODENTIA \Spalacidae \Nannospalax galili \2. Existing assembly \GCF_000622305.1 \DD \\ 225 \Degu \RODENTIA \Octodontidae \Octodon degus \2. Existing assembly \GCF_000260255.1 \LC \\ 226 \Muskrat \RODENTIA \Cricetidae \Ondatra zibethicus \1. Zoonomia \GCA_004026605.1 \LC \\ 227 \Scorpion mouse \RODENTIA \Cricetidae \Onychomys torridus \1. Zoonomia \GCA_004026725.1 \LC \\ 228 \Pacific pocket mouse \RODENTIA \Heteromyidae \Perognathus longimembris pacificus \1. Zoonomia \GCA_004363475.1 \LC \\ 229 \Prairie deer mouse \RODENTIA \Cricetidae \Peromyscus maniculatus \2. Existing assembly \GCF_000500345.1 \LC \\ 230 \Dassie rat \RODENTIA \Petromuridae \Petromus typicus \1. Zoonomia \GCA_004026965.1 \LC \\ 231 \Fat sand rat \RODENTIA \Muridae \Psammomys obesus \2. Existing assembly \GCA_002215935.1 \LC \\ 232 \Norway rat \RODENTIA \Muridae \Rattus norvegicus \2. Existing assembly \GCF_000001895.5 \LC \\ 233 \Hispid cotton rat \RODENTIA \Cricetidae \Sigmodon hispidus \1. Zoonomia \GCA_004025045.1 \LC \\ 234 \Daurian ground squirrel \RODENTIA \Sciuridae \Spermophilus dauricus \2. Existing assembly \GCA_002406435.1 \LC \\ 235 \Greater cane rat \RODENTIA \Thryonomyidae \Thryonomys swinderianus \1. Zoonomia \GCA_004025085.1 \LC \\ 236 \Cape ground squirrel \RODENTIA \Sciuridae \Xerus inauris \1. Zoonomia \GCA_004024805.1 \LC \\ 237 \Meadow jumping mouse \RODENTIA \Dipodidae \Zapus hudsonius \1. Zoonomia \GCA_004024765.1 \LC \\ 238 \Northern tree shrew \SCANDENTIA \Tupaiidae \Tupaia belangeri chinensis \2. Existing assembly \GCF_000334495.1 \LC \\ 239 \Large treeshrew \SCANDENTIA \Tupaiidae \Tupaia tana \1. Zoonomia \GCA_004365275.1 \LC \\ 240 \Florida manatee \SIRENIA \Trichechidae \Trichechus manatus \2. Existing assembly \GCF_000243295.1 \EN \\ 241 \Aardvark \TUBULIDENTATA \Orycteropodidae \Orycteropus afer \1. Zoonomia \GCA_004365145.1 \LC \
\ Table 1. Genome assemblies included in the 241-way Conservation track.
\ Species status:LC = Least Concern; NT = Near threatened; VU = Vulnerable; EN = Endangered; CR = Critically endangered
\
\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ size of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ The names of selected species are colored according to their clade,\ alternating between blue and green.\ Note that excluding species from the pairwise display does not alter the\ the conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. The following\ conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment. The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\\ Codon translation is available in base-level display mode if the\ displayed region is identified as a coding segment. To display this annotation,\ select the species for translation from the pull-down menu in the Codon\ Translation configuration section at the top of the page. Then, select one of\ the following modes:\
\ Codon translation uses the following gene tracks as the basis for translation:\
\ \\
\ Table 2. Gene tracks used for codon translation.\\ Gene Track Species \ UCSC Genes Human \ Ensembl Genes v104 Brazilian guinea pig, gibbon \ RefSeq Genes Angolan colobus, Balaenoptera acutorostrata, Bison bison, Black flying-fox, Brandt's myotis (bat), Bushbaby, Camelus bactrianus, Camelus ferus, Canis lupus familiaris, Cape elephant shrew, Capra hircus, Cavia porcellus, Ceratotherium simum, Cercocebus atys, Chinchilla, Chinese tree shrew, Chlorocebus sabaeus, Condylura cristata, Damara mole rat, Dasypus novemcinctus, David's myotis (bat), Delphinapterus leucas, Echinops telfairi, Enhydra lutris, Eptesicus fuscus, Equus asinus, Equus przewalskii, Erinaceus europaeus, Felis catus, Heterocephalus glaber, Jaculus jaculus, Kangaroo rat, Killer whale, Leptonychotes weddellii, Lipotes vexillifer, Little brown bat, Loxodonta africana, Macaca fascicularis, Macaca nemestrina, Mandrillus leucophaeus, Manis javanica, Marmota marmota, Mesocricetus auratus, Miniopterus natalensis, Mus musculus, Nannospalax galili, Ochotona princeps, Octodon degus, Oryctolagus cuniculus, Pacific walrus, Pan paniscus, Panthera tigris, Peromyscus maniculatus, Prairie vole, Propithecus coquereli, Pteropus vampyrus, Puma concolor, Rattus norvegicus, Rhinopithecus bieti, Shrew, Squirrel monkey, Squirrel, Sus scrofa, Trichechus manatus, Ursus maritimus, White-faced sapajou, Wild yak \ no annotation Acinonyx jubatus, Acomys cahirinus, Ailuropoda melanoleuca, Ailurus fulgens, Allactaga bullata, Alouatta palliata, Ammotragus lervia, Anoura caudifer, Antilocapra americana, Aotus nancymaae, Aplodontia rufa, Artibeus jamaicensis, Ateles geoffroyi, Balaenoptera bonaerensis, Beatragus hunteri, Bos indicus, Bos taurus, Bubalus bubalis, Callicebus donacophilus, Callithrix jacchus, Camelus dromedarius, Canis lupus, Capra aegagrus, Capromys pilorides, Carollia perspicillata, Castor canadensis, Catagonus wagneri, Cavia tschudii, Cebus albifrons, Ceratotherium simum cottoni, Cercopithecus neglectus, Chaetophractus vellerosus, Cheirogaleus medius, Choloepus didactylus, Choloepus hoffmanni, Chrysochloris asiatica, Craseonycteris thonglongyai, Cricetomys gambianus, Cricetulus griseus, Crocidura indochinensis, Cryptoprocta ferox, Ctenodactylus gundi, Ctenomys sociabilis, Cuniculus paca, Dasyprocta punctata, Daubentonia madagascariensis, Desmodus rotundus, Dicerorhinus sumatrensis, Diceros bicornis, Dinomys branickii, Dipodomys stephensi, Dolichotis patagonum, Elaphurus davidianus, Ellobius lutescens, Ellobius talpinus, Equus caballus, Erythrocebus patas, Eschrichtius robustus, Eubalaena japonica, Eulemur flavifrons, Eulemur fulvus, Felis nigripes, Galeopterus variegatus, Giraffa tippelskirchi, Glis glis, Gorilla gorilla, Graphiurus murinus, Helogale parvula, Hemitragus hylocrius, Heterohyrax brucei, Hippopotamus amphibius, Hipposideros armiger, Hipposideros galeritus, Hyaena hyaena, Hydrochoerus hydrochaeris, Hystrix cristata, Indri indri, Inia geoffrensis, Kogia breviceps, Lasiurus borealis, Lemur catta, Lepus americanus, Lycaon pictus, Macaca mulatta, Macroglossus sobrinus, Manis pentadactyla, Megaderma lyra, Mellivora capensis, Meriones unguiculatus, Mesoplodon bidens, Microcebus murinus, Microgale talazaci, Micronycteris hirsuta, Miniopterus schreibersii, Mirounga angustirostris, Mirza coquereli, Monodon monoceros, Mormoops blainvillei, Moschus moschiferus, Mungos mungo, Murina feae, Mus caroli, Mus pahari, Mus spretus, Muscardinus avellanarius, Mustela putorius, Myocastor coypus, Myotis myotis, Myrmecophaga tridactyla, Nasalis larvatus, Neomonachus schauinslandi, Neophocaena asiaeorientalis, Noctilio leporinus, Nycticebus coucang, Odocoileus virginianus, Okapia johnstoni, Ondatra zibethicus, Onychomys torridus, Orycteropus afer, Ovis aries, Ovis canadensis, Pan troglodytes, Panthera onca, Panthera pardus, Pantholops hodgsonii, Papio anubis, Paradoxurus hermaphroditus
\ The Zoonomia alignment was composed of two sets of mammalian genomes: newly\ assembled DISCOVAR assemblies and GenBank assemblies. The DISCOVAR genomes\ were masked with RepeatMasker (commit 2d947604), using Repbase version\ 20170127 as the repeat library and CrossMatch as the alignment engine. The\ pipeline used is available at\ repeatMaskerPipeline\ (commit a6ad966). The\ guide-tree topology was taken from the TimeTree database (using release\ current in October 2018), and the branch lengths were estimated using the\ least-squares-fit mode of PHYLIP, version\ 3.695. The distance matrix used was largely based on distances from the 4d\ site trees from the UCSC browser. To add those species not present in the\ UCSC tree, approximate distances estimated by Mash (commit 541971b)\ to the closest UCSC species\ were added to the distance between the two closest UCSC species. We used the\ HAL package (commit 68db41d)\ produce the HAL file.\
\\
\\
\ \\ The phyloP are phylogenetic methods that rely\ on a tree model containing the tree topology, branch lengths representing\ evolutionary distance at neutrally evolving sites, the background distribution\ of nucleotides, and a substitution rate matrix.\ The\ all-species tree model for this track was\ generated using the phyloFit program from the PHAST package\ (REV model, EM algorithm, medium precision) using multiple alignments of\ 4-fold degenerate sites extracted from the 241-way alignment\ (msa_view). The 4d sites were derived from the RefSeq (Reviewed+Coding) gene\ set, filtered to select single-coverage long transcripts.\
\\ This same tree model was used in the phyloP calculations; however, the\ background frequencies were modified to maintain reversibility.\ The resulting tree model:\ all species.\
\\ The phyloP program supports several different methods for computing\ p-values of conservation or acceleration, for individual nucleotides or\ larger elements (\ http://compgen.cshl.edu/phast/). Here it was used\ to produce separate scores at each base (--wig-scores option), considering\ all branches of the phylogeny rather than a particular subtree or lineage\ (i.e., the --subtree option was not used). The scores were computed by\ performing a likelihood ratio test at each alignment column (--method LRT),\ and scores for both conservation and acceleration were produced (--mode\ CONACC).\
\ \\ Zoonomia Consortium..\ \ A comparative genomics multitool for scientific discovery and conservation.\ Nature. 2020 Nov;587(7833):240-245.\ PMID: 33177664;\ PMC: PMC7759459;\ DOI: 10.1038/s41586-020-2876-6\
\ \ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q, Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ PMID: 33177663;\ PMC: PMC7673649;\ DOI: 10.1038/s41586-020-2871-y\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ PMID: 21665927;\ PMC: PMC3166836;\ DOI: 10.1101/gr.123356.111\
\ \\ Harris RS.\ Improved pairwise alignment of genomic DNA.\ Ph.D. Thesis. Pennsylvania State University, USA. 2007.\
\ \\ Cooper GM, Stone EA, Asimenos G, NISC Comparative Sequencing Program., Green ED, Batzoglou S, Sidow\ A.\ \ Distribution and intensity of constraint in mammalian genomic sequence.\ Genome Res. 2005 Jul;15(7):901-13.\ PMID: 15965027;\ PMC: PMC1172034;\ DOI: 10.1101/gr.3577405\
\ \\ Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A.\ \ Detection of nonneutral substitution rates on mammalian phylogenies.\ Genome Res. 2010 Jan;20(1):110-21.\ PMID: 19858363;\ PMC: PMC2798823\
\ \\ Siepel A, Haussler D.\ Phylogenetic Hidden Markov Models.\ In: Nielsen R, editor. Statistical Methods in Molecular Evolution.\ New York: Springer; 2005. pp. 325-351.\ DOI: 10.1007/0-387-27733-1_12\
\ \\ Siepel A, Pollard KS, and Haussler D. New methods for detecting\ lineage-specific selection. In Proceedings of the 10th International\ Conference on Research in Computational Molecular Biology (RECOMB 2006), pp. 190-205.\ DOI: 10.1007/11732990_17\
\ compGeno 1 compositeTrack on\ dimensions dimensionX=clade\ dragAndDrop subTracks\ group compGeno\ html cons241way\ longLabel Zoonomia Alignment - 241 Placental Mammal Genomes aligned by the Zoonomia Project with Cactus\ priority 2\ shortLabel Zoonomia 241 Placent\ subGroup1 view Views align=Cactus_Alignments phyloP=Basewise_Conservation_(phyloP) phastcons=Element_Conservation_(phastCons) elements=Conserved_Elements\ subGroup2 clade Clade primate=Primate carnivore=Carnivore cetartiodactyla=Cetartiodactyla chiroptera=Chiroptera rodents=Rodents mammals=Mammals all=All_species\ track cons241way\ type bed 4\ visibility hide\ covidHgiGwas COVID GWAS v3 bigLolly 9 + GWAS meta-analyses from the COVID-19 Host Genetics Initiative 0 2.1 0 0 0 127 127 127 0 0 22 chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22,\ This track set shows GWAS meta-analyses from the \ \ COVID-19 Host Genetics Initiative (HGI): \ a collaborative effort to facilitate \ the generation, analysis and sharing of COVID-19 host genetics research.\ The COVID-19 HGI organizes meta-analyses across multiple studies contributed by \ partners world-wide\ to identify the genetic determinants of SARS-CoV-2 infection susceptibility and disease severity \ and outcomes. Moreover, the COVID-19 HGI also aims to provide a platform for study partners to \ share analytical results in the form of summary statistics and/or individual level data where \ possible.\
\ \\ The specific phenotypes studied by the COVID-19 HGI are those that benefit from maximal sample \ size: primary analysis on disease severity. Two meta-analyses are represented in this track:\
\ \\ Displayed items are colored by GWAS effect: red for positive, blue for negative. \ The height of the item reflects the effect size. The effect size, defined as the \ contribution of a SNP to the genetic variance of the trait, was measured as beta coefficient \ (beta). The higher the absolute value of the beta coefficient, the stronger the effect.\ The color saturation indicates statistical significance: p-values smaller than 1e-5\ are brightly colored (bright red\ \ , bright blue\ \ ),\ those with less significance (p >= 1e-5) are paler (light red\ \ , light blue\ \ ). For better visualization of the data, only SNPs with p-values smaller than 1e-3 are \ displayed by default. \
\ \\ Each track has separate display controls and data can be filtered according to the\ number of studies, minimum -log10 p-value, and the\ effect size (beta coefficient), using the track Configure options.\
\ \\ Mouseover on items shows the rs ID (or chrom:pos if none assigned), both the non-effect \ and effect alleles, the effect size (beta coefficient), the p-value, and the number of \ studies.\ Additional information on each variant can be found on the details page by clicking on the item.\
\ \\ COVID-19 Host Genetics Initiative (HGI) GWAS meta-analysis round 3 (July 2020) results were used \ in this study. Each participating study partner submitted GWAS summary statistics for up to four \ of the COVID-19 phenotype definitions.\
\\ Data were generated from genome-wide SNP array and whole exome and genome\ sequencing, leveraging the impact of both common and rare variants. The statistical analysis\ performed takes into account differences between sex, ancestry, and date of sample collection. \ Alleles were harmonized across studies and reported allele frequencies are based on gnomAD \ version 3.0 reference data. Most study partners used the SAIGE GWAS pipeline in order \ to generate summary statistics used for the COVID-19 HGI meta-analysis. The summary statistics \ of individual studies were manually examined for inflation, \ deflation, and excessive number of false positives. Qualifying summary statistics were filtered for \ INFO > 0.6 and MAF > 0.0001 prior to meta-analyzing the entirety of the data. \ The meta-analysis was done using inverse variance weighting of effects method, accounting for \ strand differences and allele flips in the individual studies. \
\\ The meta-analysis results of variants appearing in at least three studies (analysis C2) or two \ studies (all other analyses) were made publicly available.\ The meta-analysis software and workflow are available here. More information about the \ prospective studies, processing pipeline, results and data sharing can be found \ here.\
\ \ \\ The data underlying these tracks and summary statistics results are publicly available in \ COVID19-hg Release 3 (June 2020).\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. \ Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.\
\ \\ Thanks to the COVID-19 Host Genetics Initiative contributors and project leads for making these \ data available, and in particular to Rachel Liao, Juha Karjalainen, and Kumar Veerapen at the \ Broad Institute for their review and input during browser track development.\
\ \\ COVID-19 Host Genetics Initiative.\ \ The COVID-19 Host Genetics Initiative, a global initiative to elucidate the role of host genetic\ factors in susceptibility and severity of the SARS-CoV-2 virus pandemic.\ Eur J Hum Genet. 2020 Jun;28(6):715-718.\ PMID: 32404885; PMC: PMC7220587\
\ \ \ \ phenDis 1 autoScale on\ bedNameLabel SNP\ chromosomes chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22\ compositeTrack on\ filter._effectSizeAbs 0\ filter.effectSize -13:21\ filter.pValueLog 3\ filter.sourceCount 1\ filterByRange.effectSize on\ filterLabel._effectSizeAbs Minimum effect size +-\ filterLabel.effectSize Effect size range\ filterLabel.sourceCount Minimum number of studies\ filterLimits.effectSize -13:21\ lollyField 21\ longLabel GWAS meta-analyses from the COVID-19 Host Genetics Initiative\ maxHeightPixels 48:75:128\ maxItems 500000\ mouseOver $name $ref/$alt effect $effectSize pval $pValue studies $sourceCount\ noScoreFilter on\ priority 2.1\ shortLabel COVID GWAS v3\ superTrack covid hide\ track covidHgiGwas\ type bigLolly 9 +\ viewLimits -13:21\ wgEncodeReg4RnaSeq RNA-seq (Indiv.) bigWig Signal from individual total RNA-seq experiments from ENCODE 4 0 2.1 0 0 0 127 127 127 0 0 0This track displays genome-wide, strand-specific transcription levels from 523\ individual ENCODE total RNA-seq experiments. The data capture both coding and non-coding RNAs\ profiled across all phases of the ENCODE project. The signal shown is derived from the\ reads per million (RPM) of uniquely mapped reads on each genomic strand. The data are\ processed following the\ ENCODE\ bulk RNA-seq pipeline.
\ \Each subtrack represents a single RNA-seq experiment in a specific biosample, with\ separate signal tracks for the plus and minus strands. These datasets provide the\ underlying experimental signals used to generate the corresponding layered summary tracks.\ Additional datasets measuring transcription levels are available at the\ ENCODE portal.
\ \Click a specific biosample type and organ/tissue combination to view available datasets.\ Subtracks can be further filtered by Organ, Biosample Type, Life Stage, and Strand (plus\ or minus). Each track is colored based on the organ/tissue of origin.
\ \Plus Strand and Minus Strand subtracks are colored by the organ or tissue of origin, as shown below.
\| adipose | \adrenal gland | \blood | \blood vessel | \
| bone | \brain | \breast | \connective tissue | \
| embryo | \epithelium | \esophagus | \eye | \
| gallbladder | \heart | \kidney | \large intestine | \
| liver | \lung | \mouth | \muscle | \
| nerve | \nose | \ovary | \pancreas | \
| penis | \placenta | \prostate | \skin | \
| small intestine | \spinal cord | \spleen | \stomach | \
| testis | \thyroid | \trachea | \urinary bladder | \
| uterus | \vagina | \\ | \ |
\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigWig\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.\ Clicking any accession in the track's configuration table links directly to the\ corresponding file details page on the ENCODE portal.
\ \\
These files may also be locally explored using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain data confined to a given range, e.g.,\
\
bigWigToWig -chrom=chr1 -start=100000 -end=100500 https://encode-public.s3.amazonaws.com/2020/07/14/2b6cd419-144d-49fe-aa19-b5214ca04297/ENCFF102QGV.bigWig stdout
Data were generated by the ENCODE Consortium through the following production labs:\ Drs. Barbara Wold (Caltech) and Thomas Gingeras (CSHL).
\ \The data were further processed for visualization through a collaborative effort between\ the Weng lab and the\ Moore lab at UMass\ Chan Medical School (funded by NIH grant HG012343). Integration and visualization were\ developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\ \\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 0 colorSettingsUrl /gbdb/hg38/encode4/regulation/organ_colors.json\ compositeTrack faceted\ html wgEncodeReg4RnaSeq.html\ longLabel Signal from individual total RNA-seq experiments from ENCODE 4\ maxCheckboxes 50\ metaDataUrl /gbdb/hg38/encode4/regulation/wgEncodeReg4RnaSeq_metadata.tsv\ noInherit on\ primaryKey Accession\ priority 2.1\ shortLabel RNA-seq (Indiv.)\ subtrackUrls Accession=https://www.encodeproject.org/files/$$/\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4RnaSeq\ type bigWig\ visibility hide\ covidMuts COVID Rare Harmful Var bigBed 12 + Rare variants underlying COVID-19 severity and susceptibility from the COVID Human Genetics Effort 3 2.2 179 0 0 217 127 127 0 0 0\ This track shows rare variants associated with monogenic congenital defects of immunity to \ the SARS-CoV-2 virus identified by the \ COVID Human Genetic Effort. \ This international consortium aims to discover truly causative variations: those underlying \ severe forms of COVID-19 in previously healthy individuals, and those that make certain \ individuals resistant to infection by the SARS-CoV2 virus despite repeated exposure.\
\\ The major feature of the small set of variants in this track is that they are functionally tested\ to be deleterious and genetically tested to be disease-causing. \ Specifically, rare variants were predicted to be loss-of-function at human loci known to govern\ interferon (IFN) immunity to influenza virus in patients with life-threatening COVID-19 pneumonia, \ relative to subjects with asymptomatic or benign infection.\ These genetic defects display incomplete penetrance for influenza respiratory distress and only\ appear clinically upon infection with the more virulent SARS-CoV-2.\
\ \\ Only eight genes with 23 variants are contained in this track. \ Use the links below to navigate to the gene of interest or view \ all eight genes together using the following sessions for \ hg38 or\ hg19.\
\ \| Gene Name | \Human GRCh37/hg19 Assembly | \Human GRCh38/hg38 Assembly | \
|---|---|---|
| TLR3 | \\ chr4:186990309-187006252 | \\ chr4:186069152-186088069 | \
| IRF7 | \\ chr11:612555-615999 | \\ chr11:612591-615970 | \
| UNC93B1 | \\ chr11:67758575-67771593 | \\ chr11:67991100-68004097 | \
| TBK1 | \\ chr12:64845840-64895899 | \\ chr12:64452120-64502114 | \
| TICAM1 | \\ chr19:4815936-4831754 | \\ chr19:4815932-4831704 | \
| IRF3 | \\ chr19:50162826-50169132 | \\ chr19:49659570-49665875 | \
| IFNAR1 | \\ chr21:34697214-34732128 | \\ chr21:33324970-33359864 | \
| IFNAR2 | \\ chr21:34602231-34636820 | \\ chr21:33229974-33264525 | \
\ This track uses variant calls in autosomal IFN-related genes from whole exome and genome data \ with a MAF lower than 0.001 (gnomAD v2.1.1) and experimental demonstration of loss-of-function.\ The patient population studied consisted of 659 patients with life-threatening COVID-19 pneumonia \ relative to 534 subjects with asymptomatic or benign infection of varying ethnicities. \ Variants underlying autosomal-recessive or autosomal-dominant deficiencies were identified in \ 23 patients (3.5%) 17 to 77 years of age.\ The proportion of individuals carrying at least one variant was compared between severe cases \ and control cases by means of logistic regression with the likelihood ratio test.\ Principal Component Analysis (PCA) was conducted with Plink v1.9 software on whole exome and \ genome sequencing data with the 1000 Genomes (1kG) Project phase 3 public database as reference.\ Analysis of enrichment in rare synonymous variants of the genes was performed to check the \ calibration of the burden test. \ The odds ratio was also estimated by logistic regression and adjusted for ethnic heterogeneity.\
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator.\ Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.\
\ \\ Thanks to the COVID Human Genetic Effort contributors for making these data available, and in\ particular to Qian Zhang at the Rockefeller University for review and input during browser track\ development.\
\ \\ Zhang Q, Bastard P, Liu Z, Le Pen J, Moncada-Velez M, Chen J, Ogishi M, Sabli IKD, Hodeib S, Korol C\ et al.\ \ Inborn errors of type I IFN immunity in patients with life-threatening COVID-19.\ Science. 2020 Sep 24;.\ PMID: 32972995\
\ \ phenDis 1 bigDataUrl /gbdb/hg38/covidMuts/covidMuts.bb\ color 179,0,0\ defaultLabelFields gene, name\ labelFields gene, name\ longLabel Rare variants underlying COVID-19 severity and susceptibility from the COVID Human Genetics Effort\ mouseOver $gene $name $rsId Genotype: $genotype; Zygosity: $zygo ; Inheritance: $inhMode\ multiRegionsBedUrl /gbdb/hg38/covidMuts/covidMuts.regions.bed\ noScoreFilter on\ priority 2.2\ shortLabel COVID Rare Harmful Var\ superTrack covid pack\ track covidMuts\ type bigBed 12 +\ wgEncodeReg4TfChip TF ChIP-seq (Indiv.) bed 3 Peaks and signal from individual transcription factor ChIP experiments from ENCODE 4 0 2.2 0 0 0 127 127 127 0 0 0This track displays genome-wide binding profiles of DNA-associated proteins from 2,503\ individual ENCODE ChIP-seq experiments, which form the experimental basis for the TF rPeaks\ track. These proteins include transcription factors (TFs),\ RNA polymerase, and chromatin-associated proteins involved in transcriptional regulation.\ Sequence-specific TFs bind directly to DNA motifs via DNA-binding domains, while others\ interact indirectly through protein-protein interactions. ChIP-seq (chromatin\ immunoprecipitation followed by sequencing) enables genome-wide mapping of protein-DNA\ interactions. Each ChIP-seq experiment is shown as two subtracks:
\ \| Color | \Score | \
|---|---|
| \ | 1000 (highest signal) | \
| \ | 750 | \
| \ | 500 | \
| \ | 250 | \
| \ | 1 (lowest signal) | \
Peaks often correspond to protein binding sites in specific biosamples. Additional\ ChIP-seq datasets can be explored through the\ ENCODE portal.
\ \Click a specific protein target and organ/tissue combination to view available datasets.\ Subtracks can be further filtered by TF, Organ, Biosample Type, Life Stage, and Data\ Type (Signal or Peak).
\ \Signal subtracks are colored by the organ or tissue of origin, as shown below.\ Peak subtracks use the grayscale shading by ChIP-seq signal described above and are\ not colored by organ.
\| adipose | \adrenal gland | \blood | \blood vessel | \
| bone | \bone marrow | \brain | \breast | \
| connective tissue | \embryo | \epithelium | \esophagus | \
| eye | \heart | \kidney | \large intestine | \
| liver | \lung | \lymphoid tissue | \mouth | \
| muscle | \nerve | \ovary | \pancreas | \
| parathyroid gland | \penis | \placenta | \prostate | \
| skin | \small intestine | \spinal cord | \spleen | \
| stomach | \testis | \thyroid | \uterus | \
| vagina | \\ | \ | \ |
\ The ENCODE 4 Regulation data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored in bigBed\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.\ The original data files are also available from the\ ENCODE portal.\ Clicking any accession in the track's configuration table links directly to the\ corresponding file details page on the ENCODE portal.
\ \\
These files may also be locally explored using our tool bigBedToBed,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain features confined to a given range, e.g.,\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 https://encode-public.s3.amazonaws.com/2020/12/04/ddd64b54-7aad-4a2d-9270-ce677581b64b/ENCFF492SKF.bigBed stdout
Data were generated by the ENCODE Consortium through the following production labs:\ Drs. Bradley Bernstein (Broad), John Stamatoyannopoulos (UW),\ Kevin Struhl (HMS), Kevin White (UChicago), Michael Snyder (Stanford),\ Peggy Farnham (USC), Richard Myers (HAIB), Sherman Weissman (Yale),\ Tim Reddy (Duke), Vishwanath Iyer (UTA), and Xiang-Dong Fu (UCSD).
\ \The data were further processed for visualization through a collaborative effort between\ the Weng lab and the\ Moore lab at UMass\ Chan Medical School (funded by NIH grant HG012343). Integration and visualization were\ developed by Drs. Mingshi Gao, Jill Moore, and Zhiping Weng at UMass Chan Medical School,\ who were part of the ENCODE Data Analysis Center.
\ \\ ENCODE Project Consortium, Moore JE, Purcaro MJ, Pratt HE, Epstein CB, Shoresh N, Adrian J,\ Kawli T, Davis CA, Dobin A et al.\ \ Expanded encyclopaedias of DNA elements in the human and mouse genomes.\ Nature. 2020 Jul;583(7818):699-710.\ PMID: 32728249; PMC: PMC7410828\
\ \\ Moore JE, Pratt HE, Fan K, Phalke N, Fisher J, Elhajjajy SI, Andrews G, Gao M, Shedd N,\ Fu Y et al.\ \ An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional\ Regulation.\ Nature. 2026 January 7.\ PMID: 39763870; PMC: PMC11703161\
\ regulation 1 colorSettingsUrl /gbdb/hg38/encode4/regulation/organ_colors.json\ compositeTrack faceted\ defaultSortField _Experiment\ html wgEncodeReg4TfChip.html\ longLabel Peaks and signal from individual transcription factor ChIP experiments from ENCODE 4\ maxCheckboxes 50\ metaDataUrl /gbdb/hg38/encode4/regulation/wgEncodeReg4TfChip_metadata.tsv\ noInherit on\ primaryKey Accession\ priority 2.2\ shortLabel TF ChIP-seq (Indiv.)\ subtrackUrls Accession=https://www.encodeproject.org/files/$$/ Experiment=https://www.encodeproject.org/experiments/$$/\ superTrack wgEncodeReg4 hide\ track wgEncodeReg4TfChip\ type bed 3\ visibility hide\ chainGalVar1 Malayan flying lemur Chain chain galVar1 Malayan flying lemur (Jun. 2014 (G_variegatus-3.0.2/galVar1)) Chained Alignments 3 3 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Malayan flying lemur (Jun. 2014 (G_variegatus-3.0.2/galVar1)) Chained Alignments\ otherDb galVar1\ parent placentalChainNetViewchain off\ shortLabel Malayan flying lemur Chain\ subGroups view=chain species=s006 clade=c00\ track chainGalVar1\ type chain galVar1\ chainMelGal5 Turkey Chain chain melGal5 Turkey (Nov. 2014 (Turkey_5.0/melGal5)) Chained Alignments 3 3 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Turkey (Nov. 2014 (Turkey_5.0/melGal5)) Chained Alignments\ otherDb melGal5\ parent vertebrateChainNetViewchain off\ shortLabel Turkey Chain\ subGroups view=chain species=s006 clade=c01\ track chainMelGal5\ type chain melGal5\ chainPanPan3 Bonobo Chain chain panPan3 Bonobo (May 2020 (Mhudiblu_PPA_v0/panPan3)) Chained Alignments 3 3 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Bonobo (May 2020 (Mhudiblu_PPA_v0/panPan3)) Chained Alignments\ otherDb panPan3\ parent primateChainNetViewchain off\ shortLabel Bonobo Chain\ subGroups view=chain species=s007b clade=c00\ track chainPanPan3\ type chain panPan3\ gustafsonSv 1KG ONT 100 SVs bigBed 9 + Structural Variants from 100 1000 Genomes ONT Samples (Gustafson et al. 2024) 0 3 0 0 0 127 127 127 0 0 0\ This track shows structural variants (SVs) from Oxford Nanopore long-read\ whole-genome sequencing of 100 individuals in the 1000 Genomes Project,\ as released by the 1000 Genomes Project ONT Sequencing Consortium and\ described in Gustafson et al. 2024. The cohort spans all five 1000\ Genomes superpopulations and 19 subpopulations. Samples were sequenced\ with ONT R9.4.1 pores at ~37x coverage with median read N50 of ~54 kb.\
\\ The track contains 113,159 SVs (63,177 insertions, 49,700 deletions,\ 211 inversions, 71 duplications; byte-identical duplicate records have been\ removed). Each variant was called by up to five\ independent methods (three alignment-based: Sniffles2, cuteSV, SVIM;\ and assembly-based hapdiff on Flye or Shasta/Hapdup assemblies) and then\ merged across callers and samples with Jasmine to produce a\ cross-sample consensus catalog.\
\\ This 100-sample Gustafson cohort is distinct from the Vienna\ 1000-Genomes-ONT release (1KG ONT SVs),\ which uses different samples, pore chemistry and callers; the two\ releases share neither samples nor calls.\
\ \\ Items are colored by SV type:\
\ Insertions are placed at the insertion site with a width of 1 bp; deletions,\ duplications and inversions span the affected reference interval. Filters\ are available for SV type, SV length and carrier-sample count. The detail\ page also shows the number of per-caller calls supporting each site\ (VARCALLS) and whether the source caller marked the breakpoints as precise.\
\ \\ Gustafson et al. 2024 performed Oxford Nanopore long-read sequencing on\ 100 samples from the 1000 Genomes Project (all five superpopulations and\ 19 subpopulations) using R9.4.1 flow cells, at a median per-sample\ coverage of ~37x and read N50 of ~54 kb. Per-sample SV calls were\ generated through the Napu pipeline with five independent methods: three\ alignment-based callers (Sniffles2, cuteSV and SVIM run on minimap2\ alignments to GRCh38) and two assembly-based callers (hapdiff run on Flye\ and on Shasta/Hapdup assemblies). The five per-sample VCFs were merged\ with Jasmine\ in two stages (intra-sample consensus, then cross-sample merge). The\ released confident site-level callset is defined as variants supported by\ hapdiff and at least two unique alignment-based callers, yielding 113,696\ SVs (63,177 insertions, 49,704 deletions, 744 inversions, 71\ duplications). SV counts per sample and multicaller concordance were\ benchmarked against the HPRC Sniffles2 truth and the GIAB HG002 Tier1\ region with Truvari v4.1.0.\
\\ The source Jasmine-merged VCF was downloaded from the 1000 Genomes ONT S3\ bucket:\ \ 20240423_jasmine_intrasample_noBND_custom_suppvec_alphanumeric_header_JASMINE.vcf.gz.\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator, and accessed\ programmatically through our API,\ track=gustafsonSv.\
\\ The bigBed is available from\ our\ download server as gustafson.bb. Example:\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/gustafson.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\\ The original VCF is available from the 1000 Genomes ONT S3 bucket:\ \ 20240423_jasmine_intrasample_noBND_custom_suppvec_alphanumeric_header_JASMINE.vcf.gz.\
\ \\ Thanks to Gustafson and colleagues and the 1000 Genomes Project ONT\ Sequencing Consortium for releasing this dataset.\
\ \\ Gustafson JA, Gibson SB, Damaraju N, Zalusky MPG, Hoekzema K, Twesigomwe D, Yang L, Snead AA,\ Richmond PA, De Coster W et al.\ \ High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive\ catalog of human genetic variation.\ Genome Res. 2024 Nov 20;34(11):2061-2073.\ PMID: 39358015; PMC: PMC11610458\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/gustafson.bb\ filter.AC 0:200\ filter.insLen 0:25094\ filter.sampleCount 1:100\ filter.svLen 0:98289\ filterByRange.AC on\ filterByRange.insLen on\ filterByRange.sampleCount on\ filterByRange.svLen on\ filterLabel.AC Allele Count (placeholder)\ filterLabel.insLen Insertion Length\ filterLabel.sampleCount Number of Carrier Samples\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterType.svType multipleListOr\ filterValues.svType DEL,INS,DUP,INV\ itemRgb on\ longLabel Structural Variants from 100 1000 Genomes ONT Samples (Gustafson et al. 2024)\ mouseOver Var: $name ($svType)\ The Australian\ Medical Genome Reference Bank (MGRB) collected whole-genome sequencing data of 4,011 healthy\ elderly individuals who lived ≥70 years, so the dataset is depleted of damaging genetic\ variants. Age and sex summary graphs are available from\ the MGRB website.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ VCF access can be requested via a form from\ Sydney Genomics.\
\ \\ The 4,011 MGRB samples underwent whole-genome sequencing on Illumina HiSeq X instruments at KCCG\ under ISO 15189 accreditation, with paired-end TruSeq DNA Nano libraries sequenced one lane per\ sample. Sequence reads were aligned to the hg38 reference genome assembly with bwa 0.7.15-r1140.\ Variants were called with GATK 4.1.4.0 following the Genome Analysis Toolkit (GATK) best practices\ procedure. A sites-only VCF with only passing variants (FILTER=PASS) was made with bcftools 1.20.\
\\ We received VCF files from m.hobbs@garvan.org.au via a transfer link and imported them.\ The makeDoc file of the track documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were needed; these are available from GitHub.\
\ \\ Lacaze P, Pinese M, Kaplan W, Stone A, Brion MJ, Woods RL, McNamara M, McNeil JJ, Dinger ME, Thomas\ DM.\ \ The Medical Genome Reference Bank: a whole-genome data resource of 4000 healthy elderly individuals.\ Rationale and cohort design.\ Eur J Hum Genet. 2019 Feb;27(2):308-316.\ PMID: 30353151; PMC: PMC6336775\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_mgrb/MGRB.phase3.GRCh38.norm.vcf.gz\ dataVersion Phase 3\ longLabel SNV Frequencies: Australia Medical Genome Reference Bank - 4,011 WGS\ parent varFreqs on\ priority 3\ shortLabel Australia MGRB 4k WGS\ tableBrowser off\ track mgrb\ type vcfTabix\ visibility hide\ cons30wayViewphyloP Basewise Conservation (phyloP) bed 4 UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group 2 3 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group\ parent cons30way\ shortLabel Basewise Conservation (phyloP)\ track cons30wayViewphyloP\ view phyloP\ viewLimits -3:1\ viewLimitsMax -20:1.312\ visibility full\ iscaBenignGainCum Benign Gain bedGraph 4 ClinGen CNVs: Benign Gain Coverage 2 3 0 0 200 127 127 227 0 0 0 phenDis 0 color 0,0,200\ longLabel ClinGen CNVs: Benign Gain Coverage\ parent iscaViewTotal\ shortLabel Benign Gain\ subGroups view=cov class=ben level=sub\ track iscaBenignGainCum\ bismap50Pos Bismap S50 + bigBed 6 Single-read mappability with 50-mers after bisulfite conversion (forward strand) 0 3 240 120 80 247 187 167 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k50.C2T-Converted.bb\ color 240,120,80\ longLabel Single-read mappability with 50-mers after bisulfite conversion (forward strand)\ parent bismapBigBed off\ priority 3\ shortLabel Bismap S50 +\ subGroups view=SR\ track bismap50Pos\ visibility hide\ BLCA BLCA bigLolly 12 + Bladder Urothelial Carcinoma 0 3 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/BLCA.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Bladder Urothelial Carcinoma\ parent gdcCancer off\ priority 3\ shortLabel BLCA\ track BLCA\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4DnaseBoneMarrow Bone marrow bigWig Avg. DNase level of 16 bone marrow experiments (tissues and primary cells only) 0 3 184 120 120 219 187 187 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBoneMarrowDNase.bw\ color 184,120,120\ longLabel Avg. DNase level of 16 bone marrow experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 3\ shortLabel Bone marrow\ track wgEncodeReg4DnaseBoneMarrow\ type bigWig\ wgEncodeReg4MarkH3k27acBrain Brain bigWig Avg. H3K27ac level of 68 brain experiments (tissues and primary cells only) 2 3 155 155 18 205 205 136 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBrainH3K27ac.bw\ color 155,155,18\ longLabel Avg. H3K27ac level of 68 brain experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac\ priority 3\ shortLabel Brain\ track wgEncodeReg4MarkH3k27acBrain\ type bigWig\ wgEncodeReg4MarkH3k4me3Brain Brain bigWig Avg. H3K4me3 level of 79 brain experiments (tissues and primary cells only) 0 3 155 155 18 205 205 136 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBrainH3K4me3.bw\ color 155,155,18\ longLabel Avg. H3K4me3 level of 79 brain experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3\ priority 3\ shortLabel Brain\ track wgEncodeReg4MarkH3k4me3Brain\ type bigWig\ lincRNAsCTBrain Brain bed 5 + lincRNAs from brain 1 3 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from brain\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Brain\ subGroups view=lincRNAsRefseqExp tissueType=brain\ track lincRNAsCTBrain\ wgEncodeReg4MarkCtcfBreast Breast bigWig Avg. CTCF level of 6 breast experiments (tissues and primary cells only) 0 3 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBreastCTCF.bw\ color 65,171,173\ longLabel Avg. CTCF level of 6 breast experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf off\ priority 3\ shortLabel Breast\ track wgEncodeReg4MarkCtcfBreast\ type bigWig\ wgEncodeReg4AtacBreast Breast bigWig Avg. ATAC level of 3 breast experiments (tissues and primary cells only) 0 3 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBreastATAC.bw\ color 65,171,173\ longLabel Avg. ATAC level of 3 breast experiments (tissues and primary cells only)\ parent wgEncodeReg4Atac off\ priority 3\ shortLabel Breast\ track wgEncodeReg4AtacBreast\ type bigWig\ clinGenGeneDisease ClinGen Validity bigBed 9 + ClinGen Gene-Disease Validity Classification 3 3 0 0 0 127 127 127 0 0 0 phenDis 1 bedNameLabel Associated Disease\ bigDataUrl /gbdb/hg38/bbi/clinGen/clinGenGeneDisease.bb\ filterLabel.Classification ClinGen Gene-Disease Validity Classification\ filterLabel.Inheritance Inheritance Pattern\ filterLabel.SOPversion ClinGen SOP Version Number\ filterValues.Classification Definitive,Strong,Moderate,Limited,Animal Model Only,No Reported Evidence,Disputed,Refuted\ filterValues.Inheritance Autosomal Dominant,Autosomal Recessive,Semidominant,X-Linked,X-linked recessive,Other\ filterValues.SOPversion SOP4,SOP5,SOP6,SOP7\ itemRgb on\ longLabel ClinGen Gene-Disease Validity Classification\ mouseOver Gene/ISCA ID: $geneSymbolNOTE:
\
While the DECIPHER database is \
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions.\
Because the UCSC Genes mappings for CNVs are based on associations from\ RefSeq and UniProt, they are dependent on any interpretations from those\ sources. Furthermore, because many DECIPHER records refer to multiple gene\ names, or syndromes not tightly mapped to individual genes, the associations\ in this track should be treated with skepticism and any conclusions\ based on them should be carefully scrutinized using independent\ resources.\
\Data Display Agreement Notice
\
The CNV/SNV data are only available for display in the Browser, and not for bulk\
download. Access to bulk data may be obtained directly from DECIPHER\
(https://www.deciphergenomics.org/about/data-sharing) and is subject to a\
Data Access Agreement, in which the user certifies that no attempt to\
identify individual patients will be undertaken. The same restrictions\
apply to the public data displayed at UCSC in the UCSC Genome Browser;\
no one is authorized to attempt to identify patients by any means.\
These data are made available as soon as possible and may be a\ pre-publication release. For information on the proper use of DECIPHER\ data, please see https://www.deciphergenomics.org/about/data-sharing.\
\The DECIPHER consortium provides these data in good faith as a research\ tool, but without verifying the accuracy, clinical validity, or utility of\ the data. The DECIPHER consortium makes no warranty, express or implied,\ nor assumes any legal liability or responsibility for any purpose for\ which the data are used.\
\\ The \ DECIPHER\ database of submicroscopic chromosomal imbalance \ collects clinical information about chromosomal \ microdeletions/duplications/insertions, translocations and inversions, \ and displays this information on the human genome map.\
\ The CNVs and SNVs tracks show genomic regions of reported cases and their \ associated phenotype information. All data have passed the strict\ consent requirements of the DECIPHER project and are approved for\ unrestricted public release. Clicking the Patient View ID link\ brings up a more detailed informational page on the patient at the \ DECIPHER web site.
\ \\ The Population CNVs track shows common copy-number variants (CNVs) and their\ population frequencies, lifted over from the hg19 assembly.
\ \\ The genomic locations of DECIPHER variants are labeled with the DECIPHER variant descriptions. \ Mouseover on items shows variant details, clinical interpretation, and associated conditions. \ Further information on each variant is displayed on the details page by a click onto any variant. \
\ \\ For the CNVs track, the entries are colored by the type of variant:\
\ A light-to-dark color gradient indicates the clinical significance of each variant, with \ the lightest shade being benign, to the darkest shade being pathogenic. Detailed information on the \ CNV color code is described here.\ Items can be filtered according to the size of the variant, variant type, and clinical significance \ using the track Configure options.\
\ \\ For the SNVs track, the entries are colored according to the estimated clinical significance \ of the variant:\
\ For the Population CNVs track, genomic variants are visually differentiated to facilitate quick and\ clear identification. Variants are colored according to their clinical significance and type:\
\\ The Population CNVs track's mouseover tooltip provides the following information\ about the data:\
\\ Data provided by the DECIPHER project group are imported and processed\ to create a simple BED track to annotate the genomic regions associated\ with individual patients.\
\ \ \\ For more information on DECIPHER, please contact\ \ contact@deciphergenomics.\ org\
\ \\ The DECIPHER data access and documentation can be found at\ DECIPHER Downloads.\
\ \\ Firth HV, Richards SM, Bevan AP, Clayton S, Corpas M, Rajan D, Van Vooren S, Moreau Y, Pettett RM,\ Carter NP.\ \ DECIPHER: Database of Chromosomal Imbalance and Phenotype in Humans Using Ensembl Resources.\ Am J Hum Genet. 2009 Apr;84(4):524-33.\ PMID: 19344873; PMC: PMC2667985\
\ phenDis 1 bigDataUrl /gbdb/hg38/decipher/population_cnv.bb\ filter.sample_size 0\ filterValues.type loss,gain,del/dup\ html decipherContainer\ longLabel DECIPHER: Population CNVs\ mouseOver Position: $chrom:${chromStart}-${chromEnd}\ This container comprises various DNA Methylation tracks from different sources.\ Click on the specific subtracks for detailed descriptions of the data.
\\ The two tracks available are:\
\ \\ Loyfer N, Magenheim J, Peretz A, Cann G, Bredno J, Klochendler A, Fox-Fisher I,\ Shabi-Porat S, Hecht M, Pelet T et al.\ \ A DNA methylation atlas of normal human cell types.\ Nature. 2023 Jan;613(7943):355-364.\ PMID: 36599988\
\ \\ Loyfer N, Rosenski J, Kaplan T.\ \ wgbstools: a computational suite for DNA methylation sequencing data analysis.\ Life Sci Alliance. 2026 Apr;9(4):e202503514.\ PMID: 41611450\
\ regulation 0 group regulation\ html dnaMethylation.html\ longLabel DNA Methylation\ priority 3\ shortLabel DNA Methylation\ superTrack on\ track dnaMethylation\ cons30wayViewphastcons Element Conservation (phastCons) bed 4 UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group 2 3 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group\ parent cons30way\ shortLabel Element Conservation (phastCons)\ track cons30wayViewphastcons\ view phastcons\ visibility full\ ENCFF237KCK_ENCFF700TZZ_ENCFF988QAR_ENCFF796PZW ENCFF237KCK_ENCFF700TZZ_ENCFF988QAR_ENCFF796PZW bigBed 9 + 5 Adrenal gland, female adult (41 years): (1) cCREs 4 3 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF237KCK_ENCFF700TZZ_ENCFF988QAR_ENCFF796PZW.bb\ longLabel Adrenal gland, female adult (41 years): (1) cCREs\ mouseOver ID: ${name}\ This track displays the ENCODE Registry of candidate cis-Regulatory Elements (cCREs) \ in the human genome, a total of 926,535 elements identified and classified by the ENCODE Data \ Analysis Center according to biochemical signatures.\ cCREs are the subset of representative DNase hypersensitive sites across ENCODE and\ Roadmap Epigenomics samples that are supported \ by either histone modifications (H3K4me3 and H3K27ac) or CTCF-binding data.\ The Registry of cCREs is one of the core components of the integrative level of the\ ENCODE Encyclopedia of DNA Elements.
\ \\ Additional exploration of the cCRE's and underlying raw ENCODE data is provided by the\ \ SCREEN\ (Search Candidate cis-Regulatory Elements) web tool,\ designed specifically for the Registry, accessible by linkouts from the track details page.\ The cCREs identified in the mouse genome are available in a companion track, \ here.
\ \ \ \\ CCREs are colored and labeled according to classification by regulatory signature:\
\
| Color | \\ | UCSC label | \ENCODE classification | \ENCODE label | \
|---|---|---|---|---|
| red | \prom | \promoter-like signature | \PLS | |
| orange | \enhP | \proximal enhancer-like signature | \pELS | |
| yellow | \enhD | \distal enhancer-like signature | \dELS | |
| pink | \K4m3 | \DNase-H3K4me3 | \DNase-H3K4me3 | |
| blue | \CTCF | \CTCF-only | \CTCF-only |
\ The DNase-H3K4me3 elements are those with promoter-like biochemical signature that\ are not within 200bp of an annotated TSS.\
\ \\ All individual DNase hypersensitive sites (DHSs) identified from 706 DNase-seq experiments\ in humans (a total of 93 million sites from 706 experiments) were iteratively clustered\ and filtered for the highest signal across all experiments, producing \ representative DHSs (rDHSs), with a total of 2.2 million such sites in human.\ The highest signal elements from this set that were also supported by high H3K4me3, H3K27ac \ and/or CTCF ChIP-seq signals were designated cCRE's (a total of 926,535 in human).\
\\ Classification of cCRE's was performed based on the following criteria:\
\
\
\ The GENCODE V24 (Ensembl 33) basic gene annotation set was used in this analysis.\ For further detail about the identification and classification of ENCODE cCREs see \ the About page of the\ SCREEN web tool.\
\ \\ The ENCODE accession numbers of the constituent datasets at the\ ENCODE Portal\ are available from the cCRE details page.\
\\ The data in this track can be interactively explored with the \ Table Browser or the \ Data Integrator. \ The data can be accessed from scripts through our \ API, the track name is "encodeCcreCombined".\ \
\
For automated download and analysis, this annotation is stored in a bigBed file that\
can be downloaded from\
our download server.\
The file for this track is called encodeCcreCombined.bb. \
Individual regions or the whole genome annotation can be obtained using our tool \
bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. \
Instructions for downloading source code and binaries can be found\
here.\
The tool can also be used to obtain only features within a given range, e.g.
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/encode3/ccre/encodeCcreCombined.bb -chrom=chr21 -start=0 -end=100000000 stdout
\ This annotation is based on ENCODE data released on or before September 14, 2018.
\\ Data from the Common fund supported\ Roadmap Epigenomics Mapping Consortium\ (REMC) were included for building the ENCODE cCREs. Please see the 2015 paper on their analysis\ of reference human genomes for more information.
\ \\ This dataset was produced by the\ ENCODE Data Analysis Center\ (ZLab at UMass Medical Center). Please check the\ ZLab ENCODE Public Hubs\ for the most updated data.\ Thanks to Henry Pratt, Jill Moore, Michael Purcaro, and Zhiping Weng, PI for providing\ this data.\ Thanks also to the ENCODE Consortium, the ENCODE production laboratories, \ and the ENCODE Data Coordination Center for generating and processing the datasets used here.\
\ \\ ENCODE Project Consortium.\ \ Expanded Encyclopedias of DNA Elements in the Human and Mouse Genomes.\ Nature. 2020 July 30;583(7818):699-710
\ \\ ENCODE Project Consortium.\ \ An integrated encyclopedia of DNA elements in the human genome.\ Nature. 2012 Sep 6;489(7414):57-74.\ PMID: 22955616; PMC: PMC3439153\
\ \ ENCODE Project Consortium.\ \ A user's guide to the encyclopedia of DNA elements (ENCODE).\ PLoS Biol. 2011 Apr;9(4):e1001046.\ PMID: 21526222; PMC: PMC3079585\ \ \ regulation 1 bedNameLabel ENCODE Accession\ bigDataUrl /gbdb/hg38/encode3/ccre/encodeCcreCombined.bb\ cartVersion 10\ darkerLabels on\ defaultLabelFields accessionLabel,ucscLabel\ filterLabel.ucscLabel cCRE classification\ filterValues.ucscLabel prom|promoter-like signature (PLS/prom),enhP|proximal enhancer-like signature (pELS/enhP),enhD|distal enhancer-like signature (dELS/enhD),CTCF|CTCF only (CTCF/CTCF-only),K4m3|DNase-H3K4me3 (DNase-H3K4me3/k4m3)\ group regulation\ labelFields accessionLabel,ucscLabel,encodeLabel\ longLabel ENCODE3 Registry of candidate Cis-Regulatory Elements (cCREs)\ mouseOver Accession: $name\ The GENCODE Genes track (version 45, January 2024) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The following table provides statistics for the v45 release derived from the GTF file that contains\ annotations only on the main chromosomes. More information on how they were generated can be found\ in the GENCODE site.
\ \\
\ \\
\ GENCODE v45 Release Stats \ Genes Observed Transcripts Observed \ Protein-coding genes 19,395 Protein-coding transcripts 89,110 \ Long non-coding RNA genes 20,424 - full length protein-coding 64,028 \ Small non-coding RNA genes 7,565 - partial length protein-coding 25,082 \ Pseudogenes 14,719 Nonsense mediated decay transcripts 21,427 \ Immunoglobulin/T-cell receptor gene segments 648 Long non-coding RNA loci transcripts 59,719 \ Total No of distinct translations 65,357 Genes that have more than one distinct translations 13,600
\
\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \ \\ Within a gene using the pack display mode, transcripts below a specified rank will be\ condensed into a view similar to squish mode. The transcript ranking approach is\ preliminary and will change in future releases. The transcripts rankings are defined by the\ following criteria for protein-coding and non-coding genes:
\ Protein_coding genes\\
The GENCODE v45 track was built from the GENCODE downloads file \
gencode.v45.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources\
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney. This version of the track was\ generated by Jonathan Casper.
\ \\ Frankish A, Carbonell-Sala S, Diekhans M, Jungreis I, Loveland JE, Mudge JM, Sisu C, Wright JC,\ Arnan C, Barnes I et al.\ \ GENCODE: reference annotation for the human and mouse genomes in 2023.\ Nucleic Acids Res. 2023 Jan 6;51(D1):D942-D949.\ PMID: 36420896; PMC: PMC9825462\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV45.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV45\ group genes\ html knownGeneV45\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V45\ maxItems 50000\ parent knownGeneArchive\ priority 3\ searchIndex name\ shortLabel GENCODE V45\ squishyPackField rank\ squishyPackLabel Number of transcripts shown at full height (ranked by GENCODE transcript ranking)\ squishyPackPoint 1\ track knownGeneV45\ type bigGenePred knownGenePep knownGeneMrna\ visibility pack\ geneHancerInteractionsDoubleElite GH Interactions (DE) bigInteract Interactions between GeneHancer regulatory elements and genes (Double Elite) 2 3 0 0 0 127 127 127 0 0 0 https://www.genecards.org/cgi-bin/carddisp.pl?gene=$\ This container track helps call out sections of the genome that often cause problems or\ confusion when working with the genome. The hg19 genome has a track with the same name, but with\ more subtracks, as the GeT-RM and Genome-in-a-Bottle artifact variants do not exist \ for hg38.\ \
\ The Problematic Regions track contains the following subtracks:\
\ The Highly Reproducible Regions track highlights regions and variants\ from eight samples that can be used to assess variant detection pipelines. The\ "Highly Reproducible Regions" subtrack comprises the intersection of the reproducible\ regions across all eight samples, while the "Variants" subtracks contain the reproducible\ variants from each assayed sample. Both tracks contain data from the following samples:\
\The Genome in a Bottle (GIAB) Problematic Regions tracks provide stratifications of the\ genome to evaluate variant calls in complex regions. It is designed for use with Global Alliance\ for Genomic Health (GA4GH) benchmarking tools like\ hap.py\ and includes regions with low complexity, segmental duplications, functional regions,\ and difficult-to-sequence areas. Developed in collaboration with GA4GH, the\ Genome in a Bottle (GIAB) consortium, and the\ Telomere-to-Telomere Consortium (T2T), the dataset aims to standardize the\ analysis of genetic variation by offering pre-defined BED files for stratifying true and false\ positives in genomic studies, facilitating accurate assessments in complex areas of the genome.
\ \\ The creation of the GIAB Problematic Regions tracks involves using a pipeline and configuration to\ generate stratification BED files that categorize genomic regions based on specific challenges,\ such as low complexity or difficult mapping, to facilitate accurate benchmarking of variant calls.\ For more information on the pipeline and configuration used, please visit the following webpage:\ \ https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/genome-stratifications/v3.5/README.md.\ If you have questions or comments, please write to Justin Zook (jzook@nist.gov).
\ \\ The Panmask Easy 151b Regions subtrack contains a set of sample-agnostic easy regions where\ short-read variant calling reaches high accuracy. Easy regions are derived for variant filtration\ agnostic to individual samples. They are genomic intervals where general variant callers achieve\ high accuracy without sophisticated filtering.
\\ A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that\ could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples\ were inspected to exclude regions with excessively high or low coverage or those enriched with\ low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure\ are referred to as pm151:lenient, where "pm" stands for panmask. In addition, low\ complexity regions identified by SDUST were removed.
\The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and\ other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38,\ 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter\ variant calls for clinical or research human samples. Like the HighRepro track in this container\ (see above), it shows regions that are easy to sequence, not those that are problematic. The data\ was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.
\ \\ Each track contains a set of regions of varying length with no special configuration options. \ The UCSC Unusual Regions track has a mouse-over description, all other tracks have at most\ a name field, which can be shown in pack mode. The tracks are usually kept in dense mode.\
\ \\ The Hide empty subtracks control hides subtracks with no data in the browser window.\ Changing the browser window by zooming or scrolling may result in the display of a different\ selection of tracks.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator.\ \
\
For automated download and analysis, the genome annotation is stored in bigBed files that\
can be downloaded from\
our download server.\
Individual\
regions or the whole genome annotation can be obtained using our tool bigBedToBed\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g. \
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/problematic/comments.bb -chrom=chr21 -start=0 -end=100000000 stdout
\
\ Files were downloaded from the respective databases and converted to bigBed format.\ The procedure is documented in our\ hg38 makeDoc file.\
\ \\ Thanks to Anna Benet-Pagès, Max Haeussler, Angie Hinrichs, Daniel Schmelter, and Jairo\ Navarro at the UCSC Genome Browser for planning, building, and testing these tracks. The\ underlying data comes from the\ ENCODE Blacklist and some parts were copied manually from the HGNC and NCBI\ RefSeq tracks.\
\ \\ Amemiya HM, Kundaje A, Boyle AP.\ \ The ENCODE Blacklist: Identification of Problematic Regions of the Genome.\ Sci Rep. 2019 Jun 27;9(1):9354.\ PMID: 31249361; PMC: PMC6597582\
\ \\ Dwarshuis N, Kalra D, McDaniel J, Sanio P, Alvarez Jerez P, Jadhav B, Huang WE, Mondal R, Busby B,\ Olson ND et al.\ \ The GIAB genomic stratifications resource for human reference genomes.\ Nat Commun. 2024 Oct 19;15(1):9029.\ PMID: 39424793; PMC: PMC11489684\
\ \\ Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, Gonzalez-Porta M, Eberle MA,\ Tezak Z, Lababidi S et al.\ \ Best practices for benchmarking germline small-variant calls in human genomes.\ Nat Biotechnol. 2019 May;37(5):555-560.\ PMID: 30858580; PMC: PMC6699627\
\ \\ Li H.\ \ Finding easy regions for short-read variant calling from pangenome data.\ ArXiv. 2025 Aug 8;.\ PMID: 40799803; PMC: PMC12340882\
\ \\ Pan B, Ren L, Onuchic V, Guan M, Kusko R, Bruinsma S, Trigg L, Scherer A, Ning B, Zhang C et\ al.\ \ Assessing reproducibility of inherited variants detected with short-read whole genome\ sequencing.\ Genome Biol. 2022 Jan 3;23(1):2.\ PMID: 34980216; PMC: PMC8722114\
\ map 1 compositeTrack on\ hideEmptySubtracks off\ html problematic\ longLabel Difficult regions from GIAB via NCBI\ parent problematicSuper\ priority 3\ shortLabel GIAB Problematic Regions\ track problematicGIAB\ type bigBed 3\ visibility hide\ grcExclusions GRC Exclusions bigBed 4 GRC Exclusion list: contaminations or false duplications 3 3 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/grcExclusions.bb\ longLabel GRC Exclusion list: contaminations or false duplications\ parent problematic off\ priority 3\ shortLabel GRC Exclusions\ track grcExclusions\ type bigBed 4\ gtexImmuneAtlasFullDetails GTEx Immune Atlas bigBarChart GTEx single nuclei immune expression 3 3 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$\
This track collection shows data from \
Single-nucleus cross-tissue molecular reference maps toward\
understanding disease gene function. The dataset covers ~200,000 single nuclei\
from a total of 16 human donors across 25 samples, using 4 different sample preparation\
protocols followed by droplet based single-cell RNA-seq. The samples were obtained from\
frozen tissue as part of the Genotype-Tissue Expression (GTEx) project.\
Samples were taken from the esophagus, skeletal muscle, heart, lung, prostate, breast,\
and skin. The dataset includes 43 broad cell classes, some specific to certain tissues\
and some shared across all tissue types.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ This track collection contains three bar chart tracks of RNA expression. The first track,\ Cross Tissue Nuclei, allows\ cells to be grouped together and faceted on up to 4 categories: tissue, cell class, cell subclass,\ and cell type. The second track,\ Cross Tissue Details, allows\ cells to be grouped together and faceted on up to 7 categories: tissue, cell class, cell subclass,\ cell type, granular cell type, sex, and donor. The third track,\ GTEx Immune Atlas,\ allows cells to be grouped together and faceted on up to 5 categories: tissue, cell type, cell\ class, sex, and donor.\
\ \\ Please see the\ GTEx portal\ for further interactive displays and additional data.
\ \\ Tissue-cell type combinations in the Full and Combined tracks are\ colored by which cell type they belong to in the below table:\
\
| Color | \Cell Type | \
|---|---|
| Endothelial | |
| Epithelial | |
| Glia | |
| Immune | |
| Neuron | |
| Stromal | |
| Other |
\ Tissue-cell type combinations in the Immune Atlas track are shaded according\ to the below table:\
| Color | \Cell Type | \
|---|---|
| Inflammatory Macrophage | |
| Lung Macrophage | |
| Monocyte/Macrophage FCGR3A High | |
| Monocyte/Macrophage FCGR3A Low | |
| Macrophage HLAII High | |
| Macrophage LYVE1 High | |
| Proliferating Macrophage | |
| Dendritic Cell 1 | |
| Dendritic Cell 2 | |
| Mature Dendritic Cell | |
| Langerhans | |
| CD14+ Monocyte | |
| CD16+ Monocyte | |
| LAM-like | |
| Other |
\ Using the previously collected tissue samples from the Genotype-Tissue Expression\ project, nuclei were isolated using four different protocols and sequenced\ using droplet based single cell RNA-seq. CellBender v2.1 and other standard quality\ control techniques were applied, resulting in 209,126 nuclei profiles across eight\ tissues, with a mean of 918 genes and 1519 transcripts per profile.\
\ \\ Data from all samples was integrated with a conditional variation autoencoder\ in order to correct for multiple sources of variation like sex, and protocol\ while preserving tissue and cell type specific effects.\
\ \\ For detailed methods, please refer to Eraslan et al, or the\ \ GTEx portal website.\
\ \\
The gene expression files were downloaded from the\
\
GTEx portal. The UCSC command line utilities matrixClusterColumns,\
matrixToBarChartBed, and bedToBigBed were used to transform\
these into a bar chart format bigBed file that can be visualized.\
The UCSC utilities can be found on\
our download server.\
\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions or our Data Access FAQ for more\ information.
\ \\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the GTEx Consortium for creating and analyzing these data.
\ \\ Eraslan G, Drokhlyansky E, Anand S, Fiskin E, Subramanian A, Slyper M, Wang J, Van Wittenberghe N,\ Rouhana JM, Waldman J et al.\ \ Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function.\ Science. 2022 May 13;376(6594):eabl4290.\ PMID: 35549429; PMC: PMC9383269\
\ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/gtexImmuneAtlas/facet_detailed.categories\ barChartFacets tissue,cell_type,cell_class,sex,donor\ barChartMerge on\ barChartMetric gene/genome\ barChartStatsUrl /gbdb/hg38/bbi/gtexImmuneAtlas/facet_detailed_class.facets\ barChartStretchToItem on\ barChartUnit parts per million\ bigDataUrl /gbdb/hg38/bbi/gtexImmuneAtlas/facet_detailed_class.bb\ defaultLabelFields name\ html crossTissueMaps\ labelFields name,name2\ longLabel GTEx single nuclei immune expression\ parent crossTissueMaps\ priority 3\ shortLabel GTEx Immune Atlas\ track gtexImmuneAtlasFullDetails\ type bigBarChart\ url https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$\ This track displays regions that are likely to be useful as microsatellite\ markers. These are sequences of at least 15 perfect di-nucleotide and \ tri-nucleotide repeats and tend to be highly polymorphic in the\ population.\
\ \\ The data shown in this track are a subset of the Simple Repeats track, \ selecting only those \ repeats of period 2 and 3, with 100% identity and no indels and with\ at least 15 copies of the repeat. The Simple Repeats track is\ created using the \ Tandem Repeats Finder. For more information about this \ program, see Benson (1999).
\ \\ Tandem Repeats Finder was written by \ Gary Benson.
\ \\ Benson G.\ \ Tandem repeats finder: a program to analyze DNA sequences.\ Nucleic Acids Res. 1999 Jan 15;27(2):573-80.\ PMID: 9862982; PMC: PMC148217\
\ rep 1 group rep\ longLabel Microsatellites - Di-nucleotide and Tri-nucleotide Repeats\ priority 3\ shortLabel Microsatellite\ track microsat\ type bed 4\ visibility hide\ dbSnp155Mult Mult. dbSNP(155) bigDbSnp Short Genetic Variants from dbSNP Release 155 that Map to Multiple Genomic Loci 1 3 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$ varRep 1 bigDataUrl /gbdb/hg38/snp/dbSnp155Mult.bb\ defaultGeneTracks knownGene\ longLabel Short Genetic Variants from dbSNP Release 155 that Map to Multiple Genomic Loci\ parent dbSnp155ViewVariants off\ priority 3\ shortLabel Mult. dbSNP(155)\ subGroups view=variants\ track dbSnp155Mult\ cons30wayViewalign Multiz Alignments bed 4 UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group 3 3 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group\ parent cons30way\ shortLabel Multiz Alignments\ track cons30wayViewalign\ view align\ viewUi on\ visibility pack\ caddG Mutation: G bigWig CADD 1.6 Score: Mutation is G 1 3 100 130 160 177 192 207 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/cadd/g.bw\ longLabel CADD 1.6 Score: Mutation is G\ maxHeightPixels 128:20:8\ parent cadd on\ shortLabel Mutation: G\ track caddG\ type bigWig\ viewLimits 10:50\ viewLimitsMax 0:100\ visibility dense\ promoterAiG Mutation: G bigWig PromoterAI: Mutation is G 1 3 200 0 0 0 0 200 0 0 0 phenDis 0 altColor 0,0,200\ alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/_promoterAi/g.bw\ color 200,0,0\ longLabel PromoterAI: Mutation is G\ maxHeightPixels 128:40:8\ maxWindowToDraw 10000000\ maxWindowToQuery 500000\ mouseOverFunction noAverage\ parent promoterAi on\ shortLabel Mutation: G\ track promoterAiG\ type bigWig\ viewLimits -1:1\ viewLimitsMax -1:1\ visibility dense\ cadd1_7_G Mutation: G bigWig CADD 1.7 Score: Mutation is G 1 3 100 130 160 177 192 207 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/cadd1.7/g.bw\ longLabel CADD 1.7 Score: Mutation is G\ maxHeightPixels 128:20:8\ parent cadd1_7 on\ setColorWith /gbdb/hg38/cadd1.7/g.color.bb\ shortLabel Mutation: G\ track cadd1_7_G\ type bigWig\ viewLimits 10:50\ viewLimitsMax 0:100\ visibility dense\ revelG Mutation: G bigWig REVEL: Mutation is G 1 3 150 80 200 202 167 227 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/revel/g.bw\ longLabel REVEL: Mutation is G\ maxHeightPixels 128:20:8\ maxWindowToDraw 10000000\ maxWindowToQuery 500000\ mouseOverFunction noAverage\ parent revel on\ setColorWith /gbdb/hg38/revel/g.color.bb\ shortLabel Mutation: G\ track revelG\ type bigWig\ viewLimits 0:1.0\ viewLimitsMax 0:1.0\ visibility dense\ alphaMissense_G Mutation: G bigWig AlphaMissense Score: Mutation is G 1 3 100 130 160 177 192 207 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/alphaMissense/g.bw\ longLabel AlphaMissense Score: Mutation is G\ maxHeightPixels 128:20:8\ parent alphaMissense on\ setColorWith /gbdb/hg38/alphaMissense/g.color.bb\ shortLabel Mutation: G\ track alphaMissense_G\ type bigWig\ viewLimits 0:1\ visibility dense\ mutScoreG Mutation: G bigWig MutScore: Mutation is G 2 3 50 80 200 152 167 227 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/mutscore/mutscoreG.bw\ longLabel MutScore: Mutation is G\ maxHeightPixels 128:20:8\ maxWindowToDraw 10000000\ maxWindowToQuery 500000\ mouseOverFunction noAverage\ parent mutScore on\ shortLabel Mutation: G\ track mutScoreG\ type bigWig\ viewLimits 0:1.0\ viewLimitsMax 0:1.0\ visibility full\ platinumNA12878 NA12878 vcfTabix Platinum genome variant NA12878 3 3 0 0 0 127 127 127 0 0 23 chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22,chrX, varRep 1 bigDataUrl /gbdb/hg38/platinumGenomes/NA12878.vcf.gz\ chromosomes chr1,chr2,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr20,chr21,chr22,chrX\ configureByPopup off\ group varRep\ longLabel Platinum genome variant NA12878\ maxWindowToDraw 200000\ parent platinumGenomes\ shortLabel NA12878\ showHardyWeinberg on\ track platinumNA12878\ type vcfTabix\ vcfDoFilter off\ vcfDoMaf off\ visibility pack\ nmdDetectiveB NMDetective-B bigWig NMDetective-B: Decision tree prediction of NMD efficiency (Lindeboom 2016) 0 3 0 128 255 127 191 255 0 0 0\ The NMDetective tracks display genome-wide predictions of nonsense-mediated mRNA\ decay (NMD) efficiency from\ Lindeboom et al. 2016.\ NMDetective scores predict whether a premature termination codon (PTC) at a given position\ will trigger NMD and mRNA degradation, or whether the transcript will escape NMD and\ potentially produce a truncated protein.\
\ \\ Scores range from approximately −1 to +1. Positive values indicate that a PTC at\ that position is predicted to trigger NMD (the mRNA is degraded). Negative values indicate\ that the PTC is predicted to escape NMD (the truncated mRNA may be translated into an\ aberrant protein). Values near zero indicate intermediate or uncertain NMD efficiency.\
\ \| Track | Description |
|---|---|
| NMDetective-A | \Random forest model predicting NMD efficiency for all possible PTCs introduced\ by single-nucleotide variants. Explains ~71% of systematic variance in NMD\ efficiency. |
| NMDetective-B | \Simplified decision tree model for all possible PTCs. Slightly lower accuracy\ (~68% variance explained) but more interpretable, making it suitable for\ clinical applications. |
| NMDetective-A PTC | \Random forest model predicting NMD efficiency specifically for the first\ out-of-frame PTC introduced by frameshifting indel mutations. |
| NMDetective-B PTC | \Decision tree model for the first out-of-frame PTC from frameshifting\ indels. |
\ Each subtrack is displayed as a signal (bigWig) track. By default, the vertical axis\ ranges from −1 to +1. Regions with positive values (predicted NMD-triggering) are\ shown above the baseline; regions with negative values (predicted NMD escape) are shown\ below.\
\\ The NMDetective models were trained on somatic nonsense mutation data from 9,769 cancer\ patients and validated with frameshift mutations and germline variants\ (Lindeboom et al. 2019).\ The models incorporate the following features to predict NMD efficiency:\
\\ NMDetective-A (random forest regression) captures non-linear interactions among\ these features and achieves the highest predictive accuracy.\ NMDetective-B (decision tree) applies a simpler rule-based classification that\ is more transparent, with a modest reduction in accuracy.\
\ \\ The predictions were generated for every possible PTC-introducing single-nucleotide\ variant and for the first out-of-frame PTC from every possible single-nucleotide\ frameshifting indel across all human protein-coding transcripts. The original bedGraph\ custom track files were downloaded from the\ NMDetective Figshare page\ resource and converted to bigWig format at UCSC.\
\ \\ The data underlying these tracks can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Rik Lindeboom for providing custom tracks and the original NMDetective data\ on Figshare.\
\ \\ Lindeboom RG, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Lindeboom RGH, Vermeulen M, Lehner B, Supek F.\ \ The impact of nonsense-mediated mRNA decay on genetic disease, gene editing and cancer\ immunotherapy.\ Nat Genet. 2019 Nov;51(11):1645-1651.\ PMID: 31659324; PMC: PMC6858879\
\ \ genes 0 autoScale off\ bigDataUrl /gbdb/hg38/nmd/NMDetectiveB.bw\ color 0,128,255\ html nmdDetective\ longLabel NMDetective-B: Decision tree prediction of NMD efficiency (Lindeboom 2016)\ maxHeightPixels 128:32:8\ parent nmd off\ priority 3\ shortLabel NMDetective-B\ track nmdDetectiveB\ type bigWig\ viewLimits -1:1\ visibility hide\ notinalldifficultregions Not difficult regions bigBed 3 Genome In a Bottle: not difficult regions 1 3 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/GIAB/notinalldifficultregions.bb\ longLabel Genome In a Bottle: not difficult regions\ parent problematicGIAB on\ shortLabel Not difficult regions\ track notinalldifficultregions\ type bigBed 3\ visibility dense\ nuMtSeq NuMTs Sequence bigBed 6 Nuclear mitochondrial DNA segments 0 3 0 0 0 127 127 127 1 0 0\ Nuclear mitochondrial DNA segments (NUMTs) are a kind of insertion from the mitochondrion to the\ nucleus, which is an ongoing and frequent process that happens in all eukaryotes. In previous\ studies, NUMTs have been reported to increase genetic diversity, promote gene and genome evolution,\ and generate novel nuclear exons. NUMTs can also affect the accuracy when nuclear genomes are\ assembled.
\\ This track is a collection of Nuclear mitochondrial DNA segments, provided in BED format.
\Notice: Alignments to incompletely assembled or unmapped chromosome locations are omitted\ in this track.
\\ In this track, the BED score is calculated by -10log10(E-value), representing the alignment\ confidence and is reflected in the level of gray. Scores >=100 (E-values <= 1e-10) are\ colored black. It is important to note that when a NUMT is a merged result, the score is taken as the\ highest score among all results.
\ \ \\ This dataset identifies nuclear mitochondrial genome segments (NUMTs) by comparing nuclear and\ mitochondrial genomes and proteins using LAST alignment tools. The method involves several steps:\ nuclear genome-mitochondrial genome comparison, nuclear genome-mitochondrial protein comparison,\ and exclusion of overlapping nuclear ribosomal RNA regions using maf-Bed and seg-suite tools.\ Results are merged if alignments are consistent across both comparisons, with sequences under 30bp\ excluded. Bedtools and LAST are used throughout the process for efficient alignment and merging.\
\\ For more detailed information on the methods used for detecting NUMTs, please visit the following\ webpage:
\ https://github.com/Koumokuyou/NUMTs\ \If you have questions or comments, please write to:\
Huang Muyao, \ \ 2171272903@edu.k.u-tokyo.ac.jp\ \
\ \\ Kleine T, Maier UG, Leister D.\ \ DNA transfer from organelles to the nucleus: the idiosyncratic genetics of endosymbiosis.\ Annu Rev Plant Biol. 2009;60:115-38.\ DOI: 10.1146/annurev.arplant.043008.092119; PMID: 19014347\
\\ Zhang GJ, Dong R, Lan LN, Li SF, Gao WJ, Niu HX.\ \ Nuclear Integrants of Organellar DNA Contribute to Genome Structure and Evolution in Plants.\ Int J Mol Sci. 2020 Jan 21;21(3).\ DOI: 10.3390/ijms21030707; PMID:\ 31973163; PMC: PMC7037861\
\\ Yao Y, Frith MC.\ \ Improved DNA-Versus-Protein Homology Search for Protein Fossils.\ IEEE/ACM Trans Comput Biol Bioinform. 2023 May-Jun;20(3):1691-1699.\ DOI: 10.1109/TCBB.2022.3177855; PMID: 35617174\
\\ Frith MC.\ \ A simple method for finding related sequences by adding probabilities of alternative alignments.\ Genome Res. 2024 Sep 13;.\ DOI: 10.1101/gr.279464.124;\ PMID: 39152037\
\ rep 1 bigDataUrl /gbdb/hg38/bbi/nuMtSeq/nuMtSeq_hg38.bb\ group rep\ longLabel Nuclear mitochondrial DNA segments\ priority 3\ scoreMax 100\ shortLabel NuMTs Sequence\ spectrum on\ track nuMtSeq\ type bigBed 6\ omimLocation OMIM Cyto Loci bed 4 OMIM Cytogenetic Loci Phenotypes - Gene Unknown 0 3 0 80 0 127 167 127 0 0 0 http://www.omim.org/entry/NOTE:
\
OMIM is intended for use primarily by physicians and other\
professionals concerned with genetic disorders, by genetics researchers, and\
by advanced students in science and medicine. While the OMIM database is\
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions. Further, please be\
sure to click through to omim.org for the very latest, as they are continually \
updating data.
NOTE ABOUT DOWNLOADS:
\
OMIM is the property \
of Johns Hopkins University and is not available for download or mirroring \
by any third party without their permission. Please see \
OMIM\
for downloads.
OMIM is a compendium of human genes and genetic phenotypes. The full-text,\ referenced overviews in OMIM contain information on all known Mendelian\ disorders and over 12,000 genes. OMIM is authored and edited at the\ McKusick-Nathans Institute of Genetic Medicine, Johns Hopkins University\ School of Medicine, under the direction of Dr. Ada Hamosh. This database\ was initiated in the early 1960s by Dr. Victor A. McKusick as a catalog\ of Mendelian traits and disorders, entitled Mendelian Inheritance\ in Man (MIM).\
\ \\ The OMIM data are separated into three separate tracks:\
\ \OMIM Alleles \
Variants in the OMIM database that have associated \
dbSNP identifiers. This track is currently unavailable on the hg38 assembly,\
as it depends on dbSNP data that has not been released yet.\
\
OMIM Genes\
The genomic positions of gene entries in the OMIM \
database. The coloring indicates the associated OMIM phenotype map key.\
OMIM Phenotypes - Gene Unknown\
Regions known to be associated with a phenotype, \
but for which no specific gene is known to be causative. This track \
also includes known multi-gene syndromes.\
\ This track shows the cytogenetic locations of phenotype entries in the Online Mendelian\ Inheritance in Man (OMIM) database for which\ the gene is unknown.\
\ \Cytogenetic locations of OMIM entries are displayed as solid\ blocks. The entries are colored according to the OMIM phenotype map key of associated disorders:\ \
Gene symbols and disease information, when available, are displayed on the details pages.\
\The descriptions of OMIM entries are shown on the main browser display when Full display\ mode is chosen. In Pack mode, the descriptions are shown when mousing over each entry. Items\ displayed can be filtered according to phenotype map key on the track controls page.\
\ \\ This track was constructed as follows: \
\ Because OMIM has only allowed Data queries within individual chromosomes, no download files are\ available from the Genome Browser. Full genome datasets can be downloaded directly from the\ OMIM Downloads page.\ All genome-wide downloads are freely available from OMIM after registration.
\\ If you need the OMIM data in exactly the format of the UCSC Genome Browser,\ for example if you are running a UCSC Genome Browser local installation (a partial "mirror"),\ please create a user account on omim.org and contact OMIM via\ https://omim.org/contact. Send them your OMIM\ account name and request access to the UCSC Genome Browser 'entitlement'. They will\ then grant you access to a MySQL/MariaDB data dump that contains all UCSC\ Genome Browser OMIM tables.
\\ UCSC offers queries within chromosomes from\ Table Browser that include a variety\ of filtering options and cross-referencing other datasets using our\ Data Integrator tool.\ UCSC also has an API\ that can be used to retrieve data in JSON format from a particular chromosome range.
\\ Please refer to our searchable\ mailing list archives\ for more questions and example queries, or our\ Data Access FAQ\ for more information.
\ \\ Thanks to OMIM and NCBI for the use of their data. This track was constructed by Fan Hsu,\ Robert Kuhn, and Brooke Rhead of the UCSC Genome Bioinformatics Group.
\ \Amberger J, Bocchini CA, Scott AF, Hamosh A.\ McKusick's Online Mendelian Inheritance in Man (OMIM®).\ Nucleic Acids Res. 2009 Jan;37(Database issue):D793-6. Epub 2008 Oct 8.\
\\ Hamosh A, Scott AF, Amberger JS, Bocchini CA, McKusick VA.\ Online Mendelian Inheritance in Man (OMIM), a knowledgebase of\ human genes and genetic disorders.\ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D514-7.\
\ phenDis 1 color 0, 80, 0\ hgsid on\ longLabel OMIM Cytogenetic Loci Phenotypes - Gene Unknown\ noGenomeReason Distribution restrictions by OMIM. See the track documentation for details. You can download the complete OMIM dataset for free from omim.org\ parent omimContainer\ priority 3\ shortLabel OMIM Cyto Loci\ tableBrowser noGenome\ track omimLocation\ type bed 4\ url http://www.omim.org/entry/\ visibility hide\ panelAppTandRep PanelApp GE STRs bigBed 9 + Genomics England PanelApp Short Tandem Repeats 3 3 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/panelApp/tandRep.bb\ filter.version 1\ filterLabel.version Minimum panel version to display\ filterValues.confidenceLevel 3,2,1,0\ itemRgb on\ labelFields hgncSymbol\ longLabel Genomics England PanelApp Short Tandem Repeats\ mouseOver Gene name: $geneName\ The recombination rate track represents calculated rates of recombination based\ on the genetic maps from deCODE (Halldorsson et al., 2019) and 1000 Genomes\ (2013 Phase 3 release, lifted from hg19). The deCODE map is more recent, has a higher \ resolution and was natively created on hg38 and therefore recommended. \ For the Recomb. deCODE average track, the recombination rates for chrX represent the female rate.\
\ \This track also includes a subtrack with all the\ individual deCODE recombination events and another subtrack with several thousand\ de-novo mutations found in the deCODE sequencing data. These two tracks are hidden by\ default and have to be switched on explicitly on the configuration page.\
\ \\ This is a super track that contains different subtracks, three with the deCODE\ recombination rates (paternal, maternal and average) and one with the 1000\ Genomes recombination rate (average). These tracks are in \ signal graph\ (wiggle) format. By default, to show most recombination hotspots, their maximum\ value is set to 100 cM, even though many regions have values higher than 100.\ The maximum value can be changed on the configuration pages of the tracks.\
\ \\ There are two more tracks that show additional details provided by deCODE: one\ subtrack with the raw data of all cross-overs tagged with their proband ID and\ another one with around 8000 human de-novo mutation variants that are linked to\ cross-over changes.\
\ \\ The deCODE genetic map was created at \ deCODE Genetics. It is based \ on microarrays assaying 626,828 SNP markers that allowed to identify 1,476,140 crossovers in\ 56,321 paternal meioses and 3,055,395 crossovers in 70,086 maternal meioses.\ In total, the data is based on 4,531,535 crossovers in 126,427 meioses. By\ using WGS data with 9,305,070 SNPs, the boundaries for 761,981 crossovers were\ refined: 247,942 crossovers in 9423 paternal meioses and 514,039 crossovers in\ 11,750 maternal meioses. The average resolution of the genetic map is 682 base\ pairs (bp): 655 and 708 bp for the paternal and maternal maps, respectively.\
\ \The 1000 Genomes genetic map is based on the IMPUTE genetic map based on 1000 Genomes Phase 3, on hg19 coordinates. It\ was converted to hg38 by Po-Ru Loh at the Broad Institute. After a run of \ liftOver, he post-processed the data to deal with situations in which\ consecutive map locations became much closer/farther after lifting. The\ heuristic used is sufficient for statistical phasing but may not be optimal for\ other analyses. For this reason, and because of its higher resolution, the DeCODE\ map is therefore recommended for hg38.\
\ \As with all other tracks, the data conversion commands and pointers to the\ original data files are documented in the \ makeDoc file of this track.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr17 -start=45941345 -end=45942345 http://hgdownload.soe.ucsc.edu/gbdb/hg38/recombRate/recombAvg.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ This track was produced at UCSC using data that are freely available for\ the deCODE\ and 1000 Genomes genetic maps. Thanks to Po-Ru Loh at the\ Broad Institute for providing the code to lift the hg19 1000 Genomes map data to hg38.\
\ \\ 1000 Genomes Project Consortium., Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA,\ Hurles ME, McVean GA.\ \ A map of human genome variation from population-scale sequencing.\ Nature. 2010 Oct 28;467(7319):1061-73.\ PMID: 20981092; PMC: PMC3042601\
\ \\ Halldorsson BV, Palsson G, Stefansson OA, Jonsson H, Hardarson MT, Eggertsson HP, Gunnarsson B,\ Oddsson A, Halldorsson GH, Zink F et al.\ \ Characterizing mutagenic effects of recombination through a sequence-level genetic map.\ Science. 2019 Jan 25;363(6425).\ PMID: 30679340\
\ map 0 bigDataUrl /gbdb/hg38/recombRate/recombMat.bw\ html recombRate2.html\ longLabel Recombination rate: deCODE Genetics, maternal\ maxHeightPixels 128:60:8\ parent recombRate2\ priority 3\ shortLabel Recomb. deCODE Mat\ track recombMat\ type bigWig\ viewLimits 0.0:100\ viewLimitsMax 0:150000\ visibility full\ ncbiRefSeqPredicted RefSeq Predicted genePred NCBI RefSeq genes, predicted subset (XM_* or XR_*) 1 3 12 12 120 133 133 187 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 12,12,120\ idXref ncbiRefSeqLink mrnaAcc name\ longLabel NCBI RefSeq genes, predicted subset (XM_* or XR_*)\ parent refSeqComposite off\ priority 3\ shortLabel RefSeq Predicted\ track ncbiRefSeqPredicted\ chainHg19ReMapAxtChain ReMap + axtChain hg19 chain hg19 NCBI ReMap alignments to hg19/GRCh37, joined by axtChain 0 3 0 0 0 127 127 127 0 0 0 map 1 chainLinearGap medium\ chainMinScore 3000\ longLabel NCBI ReMap alignments to hg19/GRCh37, joined by axtChain\ matrix 16 91,-114,-31,-123,-114,100,-125,-31,-31,-125,100,-114,-123,-31,-114,91\ matrixHeader A, C, G, T\ otherDb hg19\ parent liftHg19\ priority 3\ shortLabel ReMap + axtChain hg19\ track chainHg19ReMapAxtChain\ type chain hg19\ gnomad310XPercentage Sample % > 10X bigWig gnomAD Percentage of Genome Samples with at least 10X Coverage v3.0.1 2 3 195 0 60 225 127 157 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_10.bw\ color 195,0,60\ longLabel gnomAD Percentage of Genome Samples with at least 10X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 3\ shortLabel Sample % > 10X\ track gnomad310XPercentage\ viewLimits 0:1\ gnomad4Exome10XPercentage Sample % > 10X bigWig gnomAD Percentage of Exome Samples with at least 10X Coverage v4.0 2 3 195 0 60 225 127 157 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_10.bw\ color 195,0,60\ longLabel gnomAD Percentage of Exome Samples with at least 10X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 3\ shortLabel Sample % > 10X\ track gnomad4Exome10XPercentage\ viewLimits 0:1\ unipLocSignal Signal Peptide bigBed 12 + UniProt Signal Peptides 1 3 255 0 150 255 127 202 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipLocSignal.bb\ color 255,0,150\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ itemRgb off\ longLabel UniProt Signal Peptides\ mouseOver UniProt record: $uniProtId\ The FANTOM5 track shows mapped transcription start sites (TSS) and their usage in primary cells,\ cell lines, and tissues to produce a comprehensive overview of gene expression across the human\ body by using single molecule sequencing.\
\ \Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \Individual biological states are profiled by HeliScopeCAGE, which is a variation of the CAGE\ (Cap Analysis Gene Expression) protocol based on a single molecule sequencer. The standard protocol\ requiring 5 µg of total RNA as a starting material is referred to as hCAGE, and an\ optimized version for a lower quantity (~ 100 ng) is referred to as LQhCAGE (Kanamori-Katyama\ et al. 2011).\
Transcription start sites (TSSs) were mapped and their usage in human and mouse primary cells,\ cell lines, and tissues was to produce a comprehensive overview of mammalian gene expression across the\ human body. 5′-end of the mapped CAGE reads are counted at a single base pair resolution\ (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the\ sample. Individual samples shown in "TSS activity" tracks are grouped as below.\
TSS (CAGE) peaks across the panel of the biological states (samples) are identified by DPI\ (decomposition based peak identification, Forrest et al. 2014), where each of the peaks consists of\ neighboring and related TSSs. The peaks are used as anchors to define promoters and units of\ promoter-level expression analysis. Two subsets of the peaks are defined based on evidence of read\ counts, depending on scopes of subsequent analyses, and the first subset (referred as a\ robust set of the peaks, thresholded for expression analysis is shown as TSS peaks. They are\ named "p#@GENE_SYMBOL" if associated with 5'-end of known genes, or "p@CHROM:START..END,STRAND"\ otherwise. The summary tracks consist of the TSS (CAGE) peaks and summary profiles of TSS\ activities (total and maximum values). The summary track consists of the following tracks.\
\ 5′-end of the mapped CAGE reads are counted at a single base pair resolution (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the sample. The read counts tracks indicate raw counts of CAGE reads, and the TPM tracks indicate normalized counts as TPM (tags per million).\
\ \\ FANTOM5 data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ The FANTOM5 reprocessed data can be found and downloaded on the FANTOM website.
\ \\ Thanks to the FANTOM5 consortium,\ the Large Scale Data Managing Unit and Preventive Medicine and\ Applied Genomics Unit, the Center for Integrative Medical Sciences (IMS), and\ RIKEN for providing this data\ and its analysis.
\ \\ FANTOM Consortium and the RIKEN PMI and CLST (DGT), Forrest AR, Kawaji H, Rehli M, Baillie JK, de\ Hoon MJ, Haberle V, Lassmann T, Kulakovskiy IV, Lizio M et al.\ \ A promoter-level mammalian expression atlas.\ Nature. 2014 Mar 27;507(7493):462-70.\ PMID: 24670764; PMC: PMC4529748\
\ \\ Kanamori-Katayama M, Itoh M, Kawaji H, Lassmann T, Katayama S, Kojima M, Bertin N, Kaiho A, Ninomiya\ N, Daub CO et al.\ \ Unamplified cap analysis of gene expression on a single-molecule sequencer.\ Genome Res. 2011 Jul;21(7):1150-9.\ PMID: 21596820; PMC: PMC3129257\
\ \\ Lizio M, Harshbarger J, Shimoji H, Severin J, Kasukawa T, Sahin S, Abugessaisa I, Fukuda S, Hori F,\ Ishikawa-Kato S et al.\ \ Gateways to the FANTOM5 promoter level mammalian expression atlas.\ Genome Biol. 2015 Jan 5;16(1):22.\ PMID: 25723102; PMC: PMC4310165\
\ regulation 0 boxedCfg on\ compositeTrack on\ dataVersion FANTOM5 reprocessed7\ dimensions dimX=sequenceTech dimY=category dimA=strand\ html fantom5.html\ longLabel FANTOM5: TSS activity per sample (TPM)\ priority 3\ shortLabel TSS activity (TPM)\ showSubtrackColorOnUi off\ sortOrder category=+ sequenceTech=+\ subGroup1 sequenceTech Sequence_Tech hCAGE=hCAGE LQhCAGE=LQhCAGE\ subGroup2 category Category cellLine=cellLine fractionation=fractionation primaryCell=primaryCell tissue=tissue AoSMC_response_to_FGF2=AoSMC_response_to_FGF2_timecourse AoSMC_response_to_IL1b=AoSMC_response_to_IL1b_timecourse ES_to_cardiomyocyte=ES_to_cardiomyocyte_timecourse Embryoid_body_to_melanocyte=Embryoid_body_to_melanocyte_timecourse Epithelial_to_mesenchymal=Epithelial_to_mesenchymal_timecourse Human_iPS_to_neuron_Downs_syndrome_1=Human_iPS_to_neuron_Downs_syndrome_1_timecourse Human_iPS_to_neuron_Downs_syndrome_2=Human_iPS_to_neuron_Downs_syndrome_2_timecourse Human_iPS_to_neuron_wt_1=Human_iPS_to_neuron_wt_1_timecourse Human_iPS_to_neuron_wt_2=Human_iPS_to_neuron_wt_2_timecourse Lymphatic_EC_response_to_VEGFC=Lymphatic_EC_response_to_VEGFC_timecourse MCF7_response_to_EGF=MCF7_response_to_EGF_timecourse MCF7_response_to_HRG=MCF7_response_to_HRG_timecourse MSC_to_adipocyte_human=MSC_to_adipocyte_human_timecourse Macrophage_influenza_infection=Macrophage_influenza_infection_timecourse Macrophage_response_to_LPS=Macrophage_response_to_LPS_timecourse Myoblast_to_myotube_wt_and_DMD=Myoblast_to_myotube_wt_and_DMD_timecourse Preadipocyte_to_adipocyte=Preadipocyte_to_adipocyte_timecourse Rinderpest_infection_series=Rinderpest_infection_series_timecourse Saos_calcification=Saos_calcification_timecourse timecourse=other_samples_in_timecourse\ subGroup3 strand Strand forward=forward reverse=reverse\ superTrack fantom5\ track TSS_activity_TPM\ type bigWig\ visibility dense\ cons30way UCSC 30 Primates bed 4 UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group 0 3 0 0 0 127 127 127 0 0 0\ This track shows multiple alignments of 30 species and measurements of\ evolutionary conservation using\ two methods (phastCons and phyloP) from the\ \ PHAST package, for all thirty species.\ The multiple alignments were generated using multiz and\ other tools in the UCSC/Penn State Bioinformatics\ comparative genomics alignment pipeline.\ Conserved elements identified by phastCons are also displayed in\ this track.\
\\ PhastCons (which has been used in previous Conservation tracks) is a hidden\ Markov model-based method that estimates the probability that each\ nucleotide belongs to a conserved element, based on the multiple alignment.\ It considers not just each individual alignment column, but also its\ flanking columns. By contrast, phyloP separately measures conservation at\ individual columns, ignoring the effects of their neighbors. As a\ consequence, the phyloP plots have a less smooth appearance than the\ phastCons plots, with more "texture" at individual sites. The two methods\ have different strengths and weaknesses. PhastCons is sensitive to "runs"\ of conserved sites, and is therefore effective for picking out conserved\ elements. PhyloP, on the other hand, is more appropriate for evaluating\ signatures of selection at particular nucleotides or classes of nucleotides\ (e.g., third codon positions, or first positions of miRNA target sites).\
\\ Another important difference is that phyloP can measure acceleration\ (faster evolution than expected under neutral drift) as well as\ conservation (slower than expected evolution). In the phyloP plots, sites\ predicted to be conserved are assigned positive scores (and shown in blue),\ while sites predicted to be fast-evolving are assigned negative scores (and\ shown in red). The absolute values of the scores represent -log p-values\ under a null hypothesis of neutral evolution. The phastCons scores, by\ contrast, represent probabilities of negative selection and range between 0\ and 1.\
\\ Both phastCons and phyloP treat alignment gaps and unaligned nucleotides as\ missing data.\
\\ See also: lastz parameters and other details \ and chain minimum score and gap parameters used in these alignments.\
\ \\ Missing sequence in the assemblies is highlighted in the track display\ by regions of yellow when zoomed out and Ns displayed at base\ level (see Gap Annotation, below).
\\
\ \ Downloads for data in this track are available:\\
\ Organism Species Release date UCSC version alignment type \ Human Homo sapiens \ Dec. 2013 (GRCh38/hg38) Dec. 2013 (GRCh38/hg38) MAF Net \ Chimp Pan troglodytes \ May 2016 (Pan_tro 3.0/panTro5) May 2016 (Pan_tro 3.0/panTro5) MAF Net \ Bonobo Pan paniscus \ Aug. 2015 (MPI-EVA panpan1.1/panPan2) Aug. 2015 (MPI-EVA panpan1.1/panPan2) MAF Net \ Gorilla Gorilla gorilla gorilla \ Mar. 2016 (GSMRT3/gorGor5) Mar. 2016 (GSMRT3/gorGor5) MAF Net \ Orangutan Pongo pygmaeus abelii \ July 2007 (WUGSC 2.0.2/ponAbe2) July 2007 (WUGSC 2.0.2/ponAbe2) MAF Net \ Gibbon Nomascus leucogenys \ Oct. 2012 (GGSC Nleu3.0/nomLeu3) Oct. 2012 (GGSC Nleu3.0/nomLeu3) MAF Net \ Rhesus Macaca mulatta \ Nov. 2015 (BCM Mmul_8.0.1/rheMac8) Nov. 2015 (BCM Mmul_8.0.1/rheMac8) MAF Net \ Crab-eating macaque Macaca fascicularis \ Jun. 2013 (Macaca_fascicularis_5.0/macFas5) Jun. 2013 (Macaca_fascicularis_5.0/macFas5) MAF Net \ Pig-tailed macaque Macaca nemestrina \ Mar. 2015 (Mnem_1.0/macNem1) Mar. 2015 (Mnem_1.0/macNem1) MAF Net \ Sooty mangabey Cercocebus atys \ Mar. 2015 (Caty_1.0/cerAty1) Mar. 2015 (Caty_1.0/cerAty1) MAF Net \ Baboon Papio anubis \ Feb. 2013 (Baylor Panu_2.0/papAnu3) Feb. 2013 (Baylor Panu_2.0/papAnu3) MAF Net \ Green monkey Chlorocebus sabaeus \ Mar. 2014 (Chlorocebus_sabeus 1.1/chlSab2) Mar. 2014 (Chlorocebus_sabeus 1.1/chlSab2) MAF Net \ Drill Mandrillus leucophaeus \ Mar. 2015 (Mleu.le_1.0/manLeu1) Mar. 2015 (Mleu.le_1.0/manLeu1) MAF Net \ Proboscis monkey Nasalis larvatus \ Nov. 2014 (Charlie1.0/nasLar1) Nov. 2014 (Charlie1.0/nasLar1) MAF Net \ Angolan colobus Colobus angolensis palliatus \ Mar. 2015 (Cang.pa_1.0/colAng1) Mar. 2015 (Cang.pa_1.0/colAng1) MAF Net \ Golden snub-nosed monkey Rhinopithecus roxellana \ Oct. 2014 (Rrox_v1/rhiRox1) Oct. 2014 (Rrox_v1/rhiRox1) MAF Net \ Black snub-nosed monkey Rhinopithecus bieti \ Aug. 2016 (ASM169854v1/rhiBie1) Aug. 2016 (ASM169854v1/rhiBie1) MAF Net \ Marmoset Callithrix jacchus \ March 2009 (WUGSC 3.2/calJac3) March 2009 (WUGSC 3.2/calJac3) MAF Net \ Squirrel monkey Saimiri boliviensis \ Oct. 2011 (Broad/saiBol1) Oct. 2011 (Broad/saiBol1) MAF Net \ White-faced sapajou Cebus capucinus imitator \ Apr. 2016 (Cebus_imitator-1.0/cebCap1) Apr. 2016 (Cebus_imitator-1.0/cebCap1) MAF Net \ Ma's night monkey Aotus nancymaae \ Jun. 2017 (Anan_2.0/aotNan1) Jun. 2017 (Anan_2.0/aotNan1) MAF Net \ Tarsier Tarsius syrichta \ Sep. 2013 (Tarsius_syrichta-2.0.1/tarSyr2) Sep. 2013 (Tarsius_syrichta-2.0.1/tarSyr2) MAF Net \ Mouse lemur Microcebus murinus \ Feb. 2017 (Mmur_3.0/micMur3) Feb. 2017 (Mmur_3.0/micMur3) MAF Net \ Coquerel's sifaka Propithecus coquereli \ Mar. 2015 (Pcoq_1.0/proCoq1) Mar. 2015 (Pcoq_1.0/proCoq1) MAF Net \ Black lemur Eulemur macaco \ Aug. 2015 (Emacaco_refEf_BWA_oneround/eulMac1) Aug. 2015 (Emacaco_refEf_BWA_oneround/eulMac1) MAF Net \ Sclater's lemur Eulemur flavifrons \ Aug. 2015 (Eflavifronsk33QCA/eulFla1) Aug. 2015 (Eflavifronsk33QCA/eulFla1) MAF Net \ Bushbaby Otolemur garnettii \ Mar. 2011 (Broad/otoGar3) Mar. 2011 (Broad/otoGar3) MAF Net \ Mouse Mus musculus \ Dec. 2011 (GRCm38/mm10) Dec. 2011 (GRCm38/mm10) MAF Net \ Dog Canis lupus familiaris \ Sep. 2011 (Broad CanFam3.1/canFam3) Sep. 2011 (Broad CanFam3.1/canFam3) MAF Net \ Armadillo Dasypus novemcinctus \ Dec. 2011 (Baylor/dasNov3) Dec. 2011 (Baylor/dasNov3) MAF Net
\ Table 1. Genome assemblies included in the 30-way Conservation track.\
\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ value of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ The names of selected species are colored according to their clade,\ alternating between blue and green.\ Configuration buttons are available to select all of the species\ (Set all), deselect all of the species (Clear all), or\ use the default settings (Set defaults).\ Note that excluding species from the pairwise display does not alter the\ the conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. The following\ conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment.\ The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\\ Codon translation is available in base-level display mode if the\ displayed region is identified as a coding segment. To display this annotation,\ select the species for translation from the pull-down menu in the Codon\ Translation configuration section at the top of the page. Then, select one of\ the following modes:\
\ Codon translation uses the following gene tracks as the basis for\ translation, depending on the species chosen (Table 2).\ \
\ \\
\ Table 2. Gene tracks used for codon translation.\\ Gene Track Species \ Known Genes human, mouse \ Ensembl Genes v78 baboon, bushbaby, chimp, dog, gorilla, marmoset, mouse lemur, orangutan, tree shrew \ RefSeq crab-eating macaque, rhesus \ no annotation bonobo, green monkey, gibbon, proboscis monkey, golden snub-nosed monkey, squirrel monkey, tarsier
\ Pairwise alignments with the human genome were generated for\ each species using lastz from repeat-masked genomic sequence.\ Pairwise alignments were then linked into chains using a dynamic programming\ algorithm that finds maximally scoring chains of gapless subsections\ of the alignments organized in a kd-tree.\ The scoring matrix and parameters for pairwise alignment and chaining\ were tuned for each species based on phylogenetic distance from the reference.\ High-scoring chains were then placed along the genome, with\ gaps filled by lower-scoring chains, to produce an alignment net.\ For more information about the chaining and netting process and\ parameters for each species, see the description pages for the Chain and Net\ tracks.
\\ An additional filtering step was introduced in the generation of the 30-way\ conservation track to reduce the number of paralogs and pseudogenes from the\ high-quality assemblies and the suspect alignments from the low-quality\ assemblies.\
\\
\ \\
\ Table 3. Type of Net alignment\\ type of net alignment Species \ Syntenic Net baboon, chimp, dog, gibbon, green monkey, crab-eating macaque, marmoset, mouse, orangutan, rhesus \ Reciprocal best Net bushbaby, bonobo, gorilla, golden snub-nosed monkey, mouse lemur, proboscis monkey, squirrel monkey, tarsier, tree shrew
\ The resulting best-in-genome pairwise alignments\ were progressively aligned using multiz/autoMZ,\ following the tree topology diagrammed above, to produce multiple alignments.\ The multiple alignments were post-processed to\ add annotations indicating alignment gaps, genomic breaks,\ and base quality of the component sequences.\ The annotated multiple alignments, in MAF format, are available for\ bulk download.\ An alignment summary table containing an entry for each\ alignment block in each species was generated to improve\ track display performance at large scales.\ Framing tables were constructed to enable\ visualization of codons in the multiple alignment display.
\ \\ Both phastCons and phyloP are phylogenetic methods that rely\ on a tree model containing the tree topology, branch lengths representing\ evolutionary distance at neutrally evolving sites, the background distribution\ of nucleotides, and a substitution rate matrix.\ The\ all species tree model for this track was\ generated using the phyloFit program from the PHAST package\ (REV model, EM algorithm, medium precision) using multiple alignments of\ 4-fold degenerate sites extracted from the 30-way alignment\ (msa_view). The 4d sites were derived from the Xeno RefSeq gene set,\ filtered to select single-coverage long transcripts.\
\\ This same tree model was used in the phyloP calculations, however their\ background frequencies were modified to maintain reversibility.\ The resulting tree model for\ all species.\
\\ The phastCons program computes conservation scores based on a phylo-HMM, a\ type of probabilistic model that describes both the process of DNA\ substitution at each site in a genome and the way this process changes from\ one site to the next (Felsenstein and Churchill 1996, Yang 1995, Siepel and\ Haussler 2005). PhastCons uses a two-state phylo-HMM, with a state for\ conserved regions and a state for non-conserved regions. The value plotted\ at each site is the posterior probability that the corresponding alignment\ column was "generated" by the conserved state of the phylo-HMM. These\ scores reflect the phylogeny (including branch lengths) of the species in\ question, a continuous-time Markov model of the nucleotide substitution\ process, and a tendency for conservation levels to be autocorrelated along\ the genome (i.e., to be similar at adjacent sites). The general reversible\ (REV) substitution model was used. Unlike many conservation-scoring programs,\ phastCons does not rely on a sliding window\ of fixed size; therefore, short highly-conserved regions and long moderately\ conserved regions can both obtain high scores.\ More information about\ phastCons can be found in Siepel et al. (2005).
\\ The phastCons parameters used were: expected-length=45,\ target-coverage=0.3, rho=0.3.
\ \\ The phyloP program supports several different methods for computing\ p-values of conservation or acceleration, for individual nucleotides or\ larger elements\ (http://compgen.cshl.edu/phast/).\ Here it was used\ to produce separate scores at each base (--wig-scores option), considering\ all branches of the phylogeny rather than a particular subtree or lineage\ (i.e., the --subtree option was not used). The scores were computed by\ performing a likelihood ratio test at each alignment column (--method LRT),\ and scores for both conservation and acceleration were produced (--mode CONACC).\
\\ A second phyloP track, Cons 30 Mam (SSREV), computes single-base\ conservation scores using the strand-symmetric reversible (SSREV)\ substitution model rather than the standard REV model. The default REV\ model is not strand-symmetric, which can bias single-base conservation\ scores depending on the strand of the underlying transcript -- most\ visible at splice sites and other strand-specific motifs (Pollard\ et al. 2010, supplementary section 2.4). The SSREV model\ enforces equal substitution rates between complementary base pairs, so\ a splice donor on the plus strand (GT...) and on the minus\ strand (...AC) receive equivalent conservation scores.\
\\
The SSREV track uses the same alignment, tree topology, score range, and\
phyloP options (--method LRT --mode CONACC --wig-scores) as\
the REV phyloP track; only the substitution model differs. Use the SSREV\
track for analyses sensitive to transcript strand (splice sites, miRNA\
seed regions, antisense regulatory features); the REV track remains\
appropriate for general genome-wide conservation analysis.\
\ The conserved elements were predicted by running phastCons with the\ --viterbi option. The predicted elements are segments of the alignment\ that are likely to have been "generated" by the conserved state of the\ phylo-HMM. Each element is assigned a log-odds score equal to its log\ probability under the conserved model minus its log probability under the\ non-conserved model. The "score" field associated with this track contains\ transformed log-odds scores, taking values between 0 and 1000. (The scores\ are transformed using a monotonic function of the form a * log(x) + b.) The\ raw log odds scores are retained in the "name" field and can be seen on the\ details page or in the browser when the track's display mode is set to\ "pack" or "full".\
\ \This track was created using the following programs:\
The phylogenetic tree is based on Murphy et al. (2001) and general\ consensus in the vertebrate phylogeny community as of March 2007.\
\ \\ Felsenstein J, Churchill GA.\ A Hidden Markov Model approach to\ variation among sites in rate of evolution.\ Mol Biol Evol. 1996 Jan;13(1):93-104.\ PMID: 8583911\
\ \\ Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A.\ \ Detection of nonneutral substitution rates on mammalian phylogenies.\ Genome Res. 2010 Jan;20(1):110-21.\ PMID: 19858363; PMC: PMC2798823\
\ \\ Siepel A, Bejerano G, Pedersen JS, Hinrichs AS, Hou M, Rosenbloom K,\ Clawson H, Spieth J, Hillier LW, Richards S, et al.\ Evolutionarily conserved elements in vertebrate, insect, worm,\ and yeast genomes.\ Genome Res. 2005 Aug;15(8):1034-50.\ PMID: 16024819; PMC: PMC1182216\
\ \\ Siepel A, Haussler D.\ Phylogenetic Hidden Markov Models.\ In: Nielsen R, editor. Statistical Methods in Molecular Evolution.\ New York: Springer; 2005. pp. 325-351\
\ \\ Yang Z.\ A space-time process model for the evolution of DNA\ sequences.\ Genetics. 1995 Feb;139(2):993-1005.\ PMID: 7713447; PMC: PMC1306396\
\ \\ Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D.\ Evolution's cauldron:\ duplication, deletion, and rearrangement in the mouse and human genomes.\ Proc Natl Acad Sci U S A. 2003 Sep 30;100(30):11484-9.\ PMID: 14500911; PMC: PMC308784\
\ \\ Blanchette M, Kent WJ, Riemer C, Elnitski L, Smit AF, Roskin KM,\ Baertsch R, Rosenbloom K, Clawson H, Green ED, et al.\ Aligning multiple genomic sequences with the threaded blockset aligner.\ Genome Res. 2004 Apr;14(4):708-15.\ PMID: 15060014; PMC: PMC383327\
\ \\ Harris RS.\ Improved pairwise alignment of genomic DNA.\ Ph.D. Thesis. Pennsylvania State University, USA. 2007.\
\ \\ Chiaromonte F, Yap VB, Miller W.\ Scoring pairwise genomic sequence alignments.\ Pac Symp Biocomput. 2002:115-26.\ PMID: 11928468\
\ \\ Schwartz S, Kent WJ, Smit A, Zhang Z, Baertsch R, Hardison RC,\ Haussler D, Miller W.\ Human-mouse alignments with BLASTZ.\ Genome Res. 2003 Jan;13(1):103-7.\ PMID: 12529312; PMC: PMC430961\
\ \\ Murphy WJ, Eizirik E, O'Brien SJ, Madsen O, Scally M, Douady CJ, Teeling E,\ Ryder OA, Stanhope MJ, de Jong WW, Springer MS.\ Resolution of the early placental mammal radiation using Bayesian phylogenetics.\ Science. 2001 Dec 14;294(5550):2348-51.\ PMID: 11743200\
\ compGeno 1 compositeTrack on\ dragAndDrop subTracks\ group compGeno\ longLabel UCSC 30 Primates - 30 primate genomes aligned with MultiZ by the UCSC Browser Group\ priority 3\ shortLabel UCSC 30 Primates\ subGroup1 view Views align=Multiz_Alignments phyloP=Basewise_Conservation_(phyloP) phastcons=Element_Conservation_(phastCons) elements=Conserved_Elements\ track cons30way\ type bed 4\ visibility hide\ umap50 Umap S50 bigBed 6 Single-read mappability with 50-mers 0 3 80 120 240 167 187 247 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k50.Unique.Mappability.bb\ color 80,120,240\ longLabel Single-read mappability with 50-mers\ parent umapBigBed off\ priority 3\ shortLabel Umap S50\ subGroups view=SR\ track umap50\ visibility hide\ gnomadGenomesVariantsV3_1_1 gnomAD v3.1.1 bigBed 9 + Genome Aggregation Database (gnomAD) Genome Variants v3.1.1 0 3.1 0 0 0 127 127 127 0 0 0 https://gnomad.broadinstitute.org/variant/$s-$<_startPos>-$-$\ gnomAD 3 was a genomes-only release. The gnomAD v3.1.1 track is the current version of gnomAD 3\ and shows variants from 76,156 whole genomes (and no exomes), all mapped to the GRCh38/hg38\ reference sequence. 4,454 genomes were added to the number of genomes in the previous v3 release.\ For more detailed information on gnomAD v3.1, see the related blog post.\ A bugfix to v3.1 resulted in gnomAD v3.1.1, see\ changelog.\ Do not use gnomAD v3.1 anymore, we will remove the 3.1 track soon.\
\ \\ The gnomAD v3.1 track is deprecated. Please use v3.1.1 instead.\
\ \\ The gnomAD v3 track shows variants from 71,702 whole genomes (and no exomes), all mapped to the\ GRCh38/hg38 reference sequence. For more detailed\ information on gnomAD v3, see the related blog post.
\ \\ For questions on the gnomAD data, also see the gnomAD FAQ.
\\ More details on the Variant type(s) can be found on the Sequence Ontology page.
\ \\ The gnomAD v3.1.1 track version follows the same conventions and configuration as the v3.1 track,\ except as noted below.
\ \\ By default, a maximum of 50,000 variants can be displayed at a time (before applying the filters\ described below), before the track switches to dense display mode.\
\ \\ Mouse hover on an item will display many details about each variant, including the affected gene(s),\ the variant type, and annotation (missense, synonymous, etc).\
\ \\ Clicking on an item will display additional details on the variant, including a population frequency\ table showing allele count in each sub-population.\
\ \\ Following the conventions on the gnomAD browser, items are shaded according to their Annotation\ type:\
| pLoF | |
| Missense | |
| Synonymous | |
| Other |
\ To maintain consistency with the gnomAD website, variants are by default labeled according\ to their chromosomal start position followed by the reference and alternate alleles,\ for example "chr1-1234-T-CAG". dbSNP rsID's are also available as an additional\ label, if the variant is present in dbSnp.\
\ \\ Three filters are available for these tracks:\
\\ The gnomAD v3.1.1 data is unfiltered.
\ \\ For the deprecated v3.1 update only, in order to cut\ down on the amount of displayed data, the following variant\ types have been filtered out, but are still viewable in the gnomAD browser:\
\ For the full steps used to create the gnomAD tracks at UCSC, please see the\ hg38 gnomad makedoc.\
\ \ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below). The\ v3.1 and\ v3.1.1 variants can\ be found in a special directory as they have been transformed from the underlying VCF.
\ \\ For the v3.1.1 variants in particular, the underlying bigBed only contains enough information\ necessary to use the track in the browser. The extra data like VEP annotations and CADD scores are\ available in the same directory\ as the bigBed but in the files gnomad.v3.1.1.details.tab.gz and\ gnomad.v3.1.1.details.tab.gz.gzi. The gnomad.v3.1.1.details.tab.gz contains the gzip\ compressed extra data in JSON format, and the .gzi file is available to speed searching of\ this data. Each variant has an associated md5sum in the name field of the bigBed which can be\ used along with the _dataOffset and _dataLen fields to get the associated external data, as show\ below:\
\
# find item of interest:\
bigBedToBed genomes.bb stdout | head -4 | tail -1\
chr1 12416 12417 854246d79dc5d02dcdbd5f5438542b6e [..omitted for brevity..] chr1-12417-G-A 67293 902\
\
# use the final two fields, _dataOffset and _dataLen (add one to _dataLen to include a newline), to get the extra data:\
bgzip -b 67293 -s 903 gnomad.v3.1.1.details.tab.gz\
854246d79dc5d02dcdbd5f5438542b6e {"DDX11L1": {"cons": ["non_coding_transcript_variant", [..omitted for brevity..]\
\
\
\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/v3.1.1/genomes.bb\ dataVersion Release v3.1.1 (March 20, 2021) and v3.1 chrM Release (November 17, 2020)\ defaultLabelFields _displayName\ detailsDynamicTable _jsonVep|Variant Effect Predictor,_jsonPopTable|Population Frequencies,_jsonHapTable|Haplotype Frequencies\ detailsTabUrls _dataOffset=/gbdb/hg38/gnomAD/v3.1.1/gnomad.v3.1.1.details.tab.gz\ filter.AF 0.0\ filterLabel.AF Minor Allele Frequency Filter\ filterType.AC_non_cancer single\ filterType.FILTER multipleListAnd\ filterType.variation_type multipleListOr\ filterValues.AC_non_cancer Non-Cancer\ filterValues.FILTER PASS,InbreedingCoeff,RF,AC0,AS_VQSR,indel_stack (chrM only),npg (chrM only)\ filterValues.annot pLoF,missense,synonymous,other\ filterValues.variation_type 3_prime_UTR_variant,5_prime_UTR_variant,NMD_transcript_variant,coding_sequence_variant,frameshift_variant,incomplete_terminal_codon_variant,inframe_deletion,inframe_insertion,intron_variant,mature_miRNA_variant,missense_variant,non_coding_transcript_exon_variant,non_coding_transcript_variant,protein_altering_variant,splice_acceptor_variant,splice_donor_variant,splice_region_variant,start_lost,start_retained_variant,stop_gained,stop_lost,stop_retained_variant,synonymous_variant,transcript_ablation\ filterValuesDefault.AC_non_cancer Non-Cancer\ filterValuesDefault.FILTER PASS\ filterValuesDefault.annot pLoF,missense,synonymous\ html gnomadV3.html\ itemRgb on\ labelFields rsId,_displayName\ longLabel Genome Aggregation Database (gnomAD) Genome Variants v3.1.1\ maxItems 50000\ mouseOver Position: $chrom:${chromStart}-${chromEnd} ($ref/$alt)\ gnomAD 3 was a genomes-only release. The gnomAD v3.1.1 track is the current version of gnomAD 3\ and shows variants from 76,156 whole genomes (and no exomes), all mapped to the GRCh38/hg38\ reference sequence. 4,454 genomes were added to the number of genomes in the previous v3 release.\ For more detailed information on gnomAD v3.1, see the related blog post.\ A bugfix to v3.1 resulted in gnomAD v3.1.1, see\ changelog.\ Do not use gnomAD v3.1 anymore, we will remove the 3.1 track soon.\
\ \\ The gnomAD v3.1 track is deprecated. Please use v3.1.1 instead.\
\ \\ The gnomAD v3 track shows variants from 71,702 whole genomes (and no exomes), all mapped to the\ GRCh38/hg38 reference sequence. For more detailed\ information on gnomAD v3, see the related blog post.
\ \\ For questions on the gnomAD data, also see the gnomAD FAQ.
\\ More details on the Variant type(s) can be found on the Sequence Ontology page.
\ \\ The gnomAD v3.1.1 track version follows the same conventions and configuration as the v3.1 track,\ except as noted below.
\ \\ By default, a maximum of 50,000 variants can be displayed at a time (before applying the filters\ described below), before the track switches to dense display mode.\
\ \\ Mouse hover on an item will display many details about each variant, including the affected gene(s),\ the variant type, and annotation (missense, synonymous, etc).\
\ \\ Clicking on an item will display additional details on the variant, including a population frequency\ table showing allele count in each sub-population.\
\ \\ Following the conventions on the gnomAD browser, items are shaded according to their Annotation\ type:\
| pLoF | |
| Missense | |
| Synonymous | |
| Other |
\ To maintain consistency with the gnomAD website, variants are by default labeled according\ to their chromosomal start position followed by the reference and alternate alleles,\ for example "chr1-1234-T-CAG". dbSNP rsID's are also available as an additional\ label, if the variant is present in dbSnp.\
\ \\ Three filters are available for these tracks:\
\\ The gnomAD v3.1.1 data is unfiltered.
\ \\ For the deprecated v3.1 update only, in order to cut\ down on the amount of displayed data, the following variant\ types have been filtered out, but are still viewable in the gnomAD browser:\
\ For the full steps used to create the gnomAD tracks at UCSC, please see the\ hg38 gnomad makedoc.\
\ \ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below). The\ v3.1 and\ v3.1.1 variants can\ be found in a special directory as they have been transformed from the underlying VCF.
\ \\ For the v3.1.1 variants in particular, the underlying bigBed only contains enough information\ necessary to use the track in the browser. The extra data like VEP annotations and CADD scores are\ available in the same directory\ as the bigBed but in the files gnomad.v3.1.1.details.tab.gz and\ gnomad.v3.1.1.details.tab.gz.gzi. The gnomad.v3.1.1.details.tab.gz contains the gzip\ compressed extra data in JSON format, and the .gzi file is available to speed searching of\ this data. Each variant has an associated md5sum in the name field of the bigBed which can be\ used along with the _dataOffset and _dataLen fields to get the associated external data, as show\ below:\
\
# find item of interest:\
bigBedToBed genomes.bb stdout | head -4 | tail -1\
chr1 12416 12417 854246d79dc5d02dcdbd5f5438542b6e [..omitted for brevity..] chr1-12417-G-A 67293 902\
\
# use the final two fields, _dataOffset and _dataLen (add one to _dataLen to include a newline), to get the extra data:\
bgzip -b 67293 -s 903 gnomad.v3.1.1.details.tab.gz\
854246d79dc5d02dcdbd5f5438542b6e {"DDX11L1": {"cons": ["non_coding_transcript_variant", [..omitted for brevity..]\
\
\
\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/v3.1/variants/genomes.bb\ dataVersion Release 3.1 (October 29, 2020)\ defaultLabelFields name\ detailsStaticTable Population Frequencies|/gbdb/hg38/gnomAD/v3.1/variants/v3.1.genomes.popTable.txt\ filter.AF 0.0\ filterLabel.AF Minor Allele Frequency Filter\ filterType.FILTER multipleListAnd\ filterType.annot multiple\ filterType.variation_type multipleListOr\ filterValues.FILTER PASS,InbreedingCoeff,RF,AC0\ filterValues.annot pLoF,missense,synonymous,other\ filterValues.variation_type 3_prime_UTR_variant,5_prime_UTR_variant,NMD_transcript_variant,TFBS_ablation,TF_binding_site_variant,coding_sequence_variant,frameshift_variant,incomplete_terminal_codon_variant,inframe_deletion,inframe_insertion,intergenic_variant,intron_variant,mature_miRNA_variant,missense_variant,non_coding_transcript_exon_variant,non_coding_transcript_variant,protein_altering_variant,splice_acceptor_variant,splice_donor_variant,splice_region_variant,start_lost,stop_gained,stop_lost,stop_retained_variant,synonymous_variant,transcript_ablation\ filterValuesDefault.FILTER PASS\ html gnomadV3.html\ itemRgb on\ labelFields name,rsId\ longLabel Deprecated: Genome Aggregation Database (gnomAD) Genome Variants v3.1\ maxItems 50000\ mouseOver Position: $chrom:${chromStart}-${chromEnd} ($ref/$alt)\ gnomAD 3 was a genomes-only release. The gnomAD v3.1.1 track is the current version of gnomAD 3\ and shows variants from 76,156 whole genomes (and no exomes), all mapped to the GRCh38/hg38\ reference sequence. 4,454 genomes were added to the number of genomes in the previous v3 release.\ For more detailed information on gnomAD v3.1, see the related blog post.\ A bugfix to v3.1 resulted in gnomAD v3.1.1, see\ changelog.\ Do not use gnomAD v3.1 anymore, we will remove the 3.1 track soon.\
\ \\ The gnomAD v3.1 track is deprecated. Please use v3.1.1 instead.\
\ \\ The gnomAD v3 track shows variants from 71,702 whole genomes (and no exomes), all mapped to the\ GRCh38/hg38 reference sequence. For more detailed\ information on gnomAD v3, see the related blog post.
\ \\ For questions on the gnomAD data, also see the gnomAD FAQ.
\\ More details on the Variant type(s) can be found on the Sequence Ontology page.
\ \\ The gnomAD v3.1.1 track version follows the same conventions and configuration as the v3.1 track,\ except as noted below.
\ \\ By default, a maximum of 50,000 variants can be displayed at a time (before applying the filters\ described below), before the track switches to dense display mode.\
\ \\ Mouse hover on an item will display many details about each variant, including the affected gene(s),\ the variant type, and annotation (missense, synonymous, etc).\
\ \\ Clicking on an item will display additional details on the variant, including a population frequency\ table showing allele count in each sub-population.\
\ \\ Following the conventions on the gnomAD browser, items are shaded according to their Annotation\ type:\
| pLoF | |
| Missense | |
| Synonymous | |
| Other |
\ To maintain consistency with the gnomAD website, variants are by default labeled according\ to their chromosomal start position followed by the reference and alternate alleles,\ for example "chr1-1234-T-CAG". dbSNP rsID's are also available as an additional\ label, if the variant is present in dbSnp.\
\ \\ Three filters are available for these tracks:\
\\ The gnomAD v3.1.1 data is unfiltered.
\ \\ For the deprecated v3.1 update only, in order to cut\ down on the amount of displayed data, the following variant\ types have been filtered out, but are still viewable in the gnomAD browser:\
\ For the full steps used to create the gnomAD tracks at UCSC, please see the\ hg38 gnomad makedoc.\
\ \ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below). The\ v3.1 and\ v3.1.1 variants can\ be found in a special directory as they have been transformed from the underlying VCF.
\ \\ For the v3.1.1 variants in particular, the underlying bigBed only contains enough information\ necessary to use the track in the browser. The extra data like VEP annotations and CADD scores are\ available in the same directory\ as the bigBed but in the files gnomad.v3.1.1.details.tab.gz and\ gnomad.v3.1.1.details.tab.gz.gzi. The gnomad.v3.1.1.details.tab.gz contains the gzip\ compressed extra data in JSON format, and the .gzi file is available to speed searching of\ this data. Each variant has an associated md5sum in the name field of the bigBed which can be\ used along with the _dataOffset and _dataLen fields to get the associated external data, as show\ below:\
\
# find item of interest:\
bigBedToBed genomes.bb stdout | head -4 | tail -1\
chr1 12416 12417 854246d79dc5d02dcdbd5f5438542b6e [..omitted for brevity..] chr1-12417-G-A 67293 902\
\
# use the final two fields, _dataOffset and _dataLen (add one to _dataLen to include a newline), to get the extra data:\
bgzip -b 67293 -s 903 gnomad.v3.1.1.details.tab.gz\
854246d79dc5d02dcdbd5f5438542b6e {"DDX11L1": {"cons": ["non_coding_transcript_variant", [..omitted for brevity..]\
\
\
\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/vcf/gnomad.genomes.r3.0.sites.vcf.gz\ configureByPopup off\ dataVersion Release 3.0 (October 16, 2019)\ html gnomadV3.html\ longLabel Genome Aggregation Database (gnomAD) Genome Variants v3\ maxWindowToDraw 200000\ parent gnomadVariants\ priority 3.3\ shortLabel gnomAD v3\ showHardyWeinberg on\ track gnomadGenomesVariantsV3\ type vcfTabix\ url http://gnomad.broadinstitute.org/variant/$s-$\ The arrays listed in this track are probes from the\ Agilent Catalog Oligonucleotide Microarrays.\
\Please note that more microarray tracks are available on the hg19 genome assembly. \ To view those tracks, please \ click this link for hg19 microarrays.\ Microarrays that are not listed can be added as Custom Tracks with data from the companies.\
\ \\ Agilent's oligonucleotide CGH (Comparative Genomic Hybridization) platform enables the\ study of genome-wide DNA copy number changes at a high resolution. The CGH probes on Agilent\ CGH microarrays are 60-mer oligonucleotides synthesized in situ using Agilent's inkjet\ SurePrint technology. The probes represented on the Agilent CGH microarrays have been\ selected using algorithms developed specifically for the CGH application, assuring optimal\ performance of these probes in detecting DNA copy number changes.\
\ \\ With the Infinium MethylationEPIC BeadChip Kit, researchers can interrogate over 850,000\ methylation sites quantitatively across the genome at single-nucleotide resolution. Multiple\ samples, including FFPE, can be analyzed in parallel to deliver high-throughput power while\ minimizing the cost per sample. These tracks show positions being measured on the Illumina 450k and\ 850k (EPIC) microarray tracks, not the probe locations themselves. Contact us\ or Illumina if you need the probe locations directly. More information about\ the arrays can be found on the\ Infinium MethylationEPIC Kit website.\
\ Note: The 450k track on hg38 contains 128,989 regions representing the target regions, not the probes\ themselves.
\ \\ The Infinium CytoSNP-850K v1.2 BeadChip provides comprehensive coverage of\ cytogenetically relevant genes on a proven platform, helping researchers find valuable information\ that may be missed by other technologies. It contains approximately 850,000 empirically selected\ single nucleotide polymorphisms (SNPs) spanning the entire genome with enriched coverage for 3,262\ genes of known cytogenetics relevance in both constitutional and cancer applications. \
\ \\ The CytoScan HD Array, which is included in the\ CytoScan HD Suite, provides the broadest coverage and highest performance for\ detecting chromosomal aberrations. CytoScan HD Suite has greater than 99% sensitivity and can\ reliably detect 25-50kb copy number changes across the genome at high specificity with\ single-nucleotide polymorphism (SNP) allelic corroboration. With more than 2.6 million copy number\ markers, CytoScan HD Suite covers all OMIM and RefSeq genes.\
\ \\ Bionano Laboratories provides access to Optical Genome Mapping (OGM) data for projects across a variety of\ applications for researchers, clinicians, and pharmaceutical companies.
\This track shows the CTTAAG sites used by the \ Bionano Optical Genome Mapping system,\ an assay to detect structural variants.\
\ \\ Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \ \\ The Agilent arrays were downloaded from their \ Agilent SureDesign website tool on March 2022.
\\ The Illumina 450k and 850k (EPIC) tracks were created using a few columns from the\ Infinium MethylationEPIC v1.0 B5 Manifest File (CSV Format)\ and was then converted into a bigBed.
\\ The Illumina CytoSNP-850K track was created by downloading the\ CytoSNP-850K v1.2 Manifest File (CSV Format) (GRCh38) file and then converted\ into a bigBed file.\
\\ The Affymetrix Cytoscan HD GeneChip Array track was created by converting the \ CytoScanHD_Accel_Array.na36.bed.zip\ into a bigBed file.\
\\ The Bionano track was created by receiving the BED files from\ \ apang@bionano.\ com\ \ and converted to bigBed files using the bedToBigBed tool.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API \ or downloaded from our \ Downloads site. Please refer to our\ \ mailing list archives for questions, or our\ \ Data Access FAQ for more information.\
\ \\ Thanks to the Agilent and Illumina support teams for sharing the data and the UCSC Genome Browser\ engineers for configuring the data.
\\ Thanks to Andy Pang from Bionano Genomics for providing the BED data file.
\ varRep 1 bigDataUrl /gbdb/hg38/bbi/illumina/illuminaEPICv2.bb\ colorByStrand 255,0,0 0,0,255\ html genotypeArrays\ longLabel Illumina EPIC v2 Methylation Array\ mouseOver Probe ID: $IlluminaName\ The Arquivo Brasileiro Online de\ Mutações (ABraOM) provides genomic variants obtained with whole-genome sequencing\ from SABE, a census-based sample of elderly individuals from São Paulo, Brazil's largest\ city. The Brazilian population reflects ~500 years of admixture between Africans,\ Europeans, and Native Americans. About 3% of the cohort has non-admixed Japanese ancestry\ (early 20th century migration). Coverage is 38.6x. TEs, HLAs and new sequence are also available.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is abraom.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original data can also be downloaded from the ABraOM website.\
\ \\ For academic use only. Licensing for commercial use might be available under request and agreement.\ By using this resource you agree to cite the flagship paper (Naslavsky et al. Nat Comm 2022).\
\\ Whole-genome sequencing was performed at Human Longevity Inc. using TruSeq Nano DNA HT libraries\ sequenced on Illumina HiSeqX instruments with 150 bp paired-end reads targeting 30x coverage, and\ reads were mapped to GRCh38 using ISIS software. Sample sex was validated by comparing CPMs of X\ chromosome and male-specific Y (MSY) reads relative to autosomes, yielding the expected female\ (~55,000 X CPM, <200 MSY CPM) and male (~27,500 X CPM, >550 MSY CPM) patterns. Germline SNVs\ and indels were called following GATK Best Practices (GATK v3.7) via per-sample GVCFs\ (HaplotypeCaller), joint genotyping (CombineGVCFs, GenotypeGVCFs), and Variant Quality Score\ Recalibration (VQSR-AS); multiallelic variants were split with an in-house script, left-aligned with\ BCFtools, and annotated using Annovar and custom scripts against dbSNP, 1000 Genomes, and gnomAD,\ with putative loss-of-function variants identified using LOFTEE v0.3-beta irrespective of confidence\ labels. Variant and genotype quality was further assessed using the in-house CEGH-Filter two-step\ algorithm based on depth and allele balance, and analyses retained only GATK VQSR-AS PASS variants\ and higher-confidence CEGH-Filter calls. Relatedness was assessed using KING and PC-Relate\ (GENESIS), retaining a single proband per related pair and excluding one contaminated sample\ (>3% by verifyBAMID), resulting in a final dataset of 1,171 unrelated individuals. Final samples\ achieved mean coverages ranging from 31.3x to 64.8x, with an average of 38.65x and a median of\ 36.6x.\ The makeDoc file documents how the source files of the varFreqs track were converted.\ For some tracks, python scripts were necessary and are also available from GitHub.\
\ \\ Naslavsky MS, Scliar MO, Yamamoto GL, Wang JYT, Zverinova S, Karp T, Nunes K, Ceroni JRM, de\ Carvalho DL, da Silva Simões CE et al.\ \ Whole-genome sequencing of 1,171 elderly admixed individuals from São Paulo, Brazil.\ Nat Commun. 2022 Mar 4;13(1):1004.\ PMID: 35246524; PMC: PMC8897431\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/abraom/abraom.vcf.gz\ dataVersion SABE-WGS-1171 Sep 2020\ longLabel SNV Frequencies: ABraOM Brazil - 1,171 unrelated individuals\ parent varFreqs on\ priority 4\ shortLabel Brazil ABraOM 1k WGS\ track abraom\ type vcfTabix\ visibility hide\ BRCA BRCA bigLolly 12 + Breast invasive carcinoma 0 4 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/BRCA.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Breast invasive carcinoma\ parent gdcCancer off\ priority 4\ shortLabel BRCA\ track BRCA\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4MarkH3k27acBreast Breast bigWig Avg. H3K27ac level of 5 breast experiments (tissues and primary cells only) 2 4 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBreastH3K27ac.bw\ color 65,171,173\ longLabel Avg. H3K27ac level of 5 breast experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 4\ shortLabel Breast\ track wgEncodeReg4MarkH3k27acBreast\ type bigWig\ wgEncodeReg4MarkH3k4me3Breast Breast bigWig Avg. H3K4me3 level of 11 breast experiments (tissues and primary cells only) 0 4 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBreastH3K4me3.bw\ color 65,171,173\ longLabel Avg. H3K4me3 level of 11 breast experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 4\ shortLabel Breast\ track wgEncodeReg4MarkH3k4me3Breast\ type bigWig\ recount3_ccle CCLE bigBed 9 + recount3 CCLE introns 0 4 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/recount3/ccle.bb\ filter.readcount 10000:2000000000\ filter.size 30:100000\ filterByRange.readcount on\ filterByRange.size on\ filterLabel.readcount Filter by supporting split reads\ filterLabel.size Filter by intron size\ filterLabel.sjPair splice junctions (format GT/AG)\ filterLabel.strand Strand\ filterLimits.readcount 0:2000000000\ filterText.sjPair *\ filterType.sjPair wildcard\ filterType.strand multiple\ filterValues.strand +,-,.\ iframeOptions height='300' width='1000' scrolling='yes'\ iframeUrl https://snaptron.cs.jhu.edu/snaptron-studies/jxn2studies?compilation=ccle&jid=$$&coords=$S:${-$}\ itemRgb on\ labelFields none\ longLabel recount3 CCLE introns\ mouseOver Split read count: $readcount\ The GENCODE Genes track (version 44, July 2023) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The following table provides statistics for the v44 release derived from the GTF file that contains\ annotations only on the main chromosomes. More information on how they were generated can be found\ in the GENCODE site.
\ \\
\ \\
\ GENCODE v44 Release Stats \ Genes Observed Transcripts Observed \ Protein-coding genes 19,396 Protein-coding transcripts 89,067 \ Long non-coding RNA genes 19,922 - full length protein-coding 63,968 \ Small non-coding RNA genes 7,566 - partial length protein-coding 25,099 \ Pseudogenes 14,735 Nonsense mediated decay transcripts 21,384 \ Immunoglobulin/T-cell receptor gene segments 647 Long non-coding RNA loci transcripts 58,246 \ Total No of distinct translations 65,342 Genes that have more than one distinct translations 13,594
\
\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \ \\ Within a gene using the pack display mode, transcripts below a specified rank will be\ condensed into a view similar to squish mode. The transcript ranking approach is\ preliminary and will change in future releases. The transcripts rankings are defined by the\ following criteria for protein-coding and non-coding genes:
\ Protein_coding genes\\
The GENCODE v44 track was built from the GENCODE downloads file \
gencode.v44.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources\
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney. This version of the track was\ generated by Jonathan Casper.
\ \\ Frankish A, Carbonell-Sala S, Diekhans M, Jungreis I, Loveland JE, Mudge JM, Sisu C, Wright JC,\ Arnan C, Barnes I et al.\ \ GENCODE: reference annotation for the human and mouse genomes in 2023.\ Nucleic Acids Res. 2023 Jan 6;51(D1):D942-D949.\ PMID: 36420896; PMC: PMC9825462\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV44.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV44\ group genes\ html knownGeneV44\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V44\ maxItems 50000\ parent knownGeneArchive\ priority 4\ searchIndex name\ shortLabel GENCODE V44\ squishyPackField rank\ squishyPackLabel Number of transcripts shown at full height (ranked by GENCODE transcript ranking)\ squishyPackPoint 1\ track knownGeneV44\ type bigGenePred knownGenePep knownGeneMrna\ visibility hide\ geneHancerClusteredInteractionsDoubleElite GH Clusters (DE) bigInteract Clustered interactions of GeneHancer regulatory elements and genes (Double Elite) 3 4 0 0 0 127 127 127 0 0 0 https://www.genecards.org/cgi-bin/carddisp.pl?gene=$\ The gnomAD v2 tracks show variants from 125,748 exomes and 15,708 whole genomes, all mapped to\ the GRCh37/hg19 reference sequence and lifted to the GRCh38/hg38 assembly. The data originate\ from 141,456 unrelated individuals sequenced as part of various population-genetic and\ disease-specific studies\ collected by the Genome Aggregation Database (gnomAD), release 2.1.1.\ Raw data from all studies have been reprocessed through a unified pipeline and jointly\ variant-called to increase consistency across projects. For more information on the processing\ pipeline and population annotations, see the following blog post\ and the 2.1.1 README.
\\ gnomAD v2 data are based on the GRCh37/hg19 assembly. These tracks display the\ GRCh38/hg38 lift-over provided by gnomAD on their downloads site.\
\ \The gnomAD MPC score (Missense Deleteriousness Prediction by Constraint) is available for now only on hg19.
\ \\ For questions on the gnomAD data, also see the gnomAD FAQ.
\ \\ The gnomAD v2.1.1 track follows the standard display and configuration options available for\ VCF tracks, briefly explained below.\
\\ Four filters are available for these tracks, the same as the underlying VCF:\
\ There are two additional filters available, one for the minimum minor allele frequency, and a configurable filter on the QUAL score.\
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below). Variant VCFs can be found in the\ vcf/ subdirectory.
\ \\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\ varRep 1 compositeTrack on\ configureByPopup off\ dataVersion Release 2.1.1 (March 6, 2019)\ html gnomadV2.html\ longLabel Genome Aggregation Database (gnomAD) Genome and Exome Variants v2.1\ maxWindowToDraw 200000\ parent gnomadVariants\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 4\ shortLabel gnomAD v2\ showHardyWeinberg on\ track gnomadVariantsV2\ type vcfTabix\ visibility hide\ gnomad4ExomeCoverage gnomAD v4 Exome Coverage bigWig Genome Aggregation Database (gnomAD) Exome Sample Coverage v4.0 2 4 0 0 0 127 127 127 0 0 0\ The Genome Aggregation Database (gnomAD) v4 - Exome Coverage track shows how many\ times regions of the genomes were sequenced. This track includes several subtracks of average\ coverage metrics and sample percentage of coverage.\
\\ There is no gnomAD v4 genome coverage track because the genomes were unchanged from V3. There is no\ gnomAD v3 exomes track because v3 was a genome-only release.
\ \\ The Average/Median Sample Coverage tracks display the mean and median read depth of the\ samples at each base position. The details page shows calculated sample percentages for the range\ of sequence within the browser window.\
\ \\ The nX Coverage Percentage tracks display the percentage of samples whose read\ depth is at least 1X, 5X, 10X, 15X, 20X, 25X, 30X, 50X, and 100X at each base position. The details\ page shows calculated sample percentages for the range of sequence within the browser window.\
\ \\ Coverage was computed using all gnomAD 4 exome samples from their gVCFs. The gVCFs were produced\ using a 3-bin blocking scheme:\
\\ The coverage was binned by quality using the thresholds above and the median coverage value for each\ of the resulting coverage blocks was used to compute the coverage metrics presented in the browser.\ Coverage was computed for all callable bases in the genome (all non-N bases, minus telomeres and\ centromeres).\
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that \ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below). Coverage values\ for the genome are in bigWig files in\ the coverage/ subdirectory. Variant VCFs can be found in the vcf/ subdirectory.
\\ The data can also be found directly from the gnomAD downloads page. Please refer to \ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ More information about using and understanding the gnomAD data can be found in the\ gnomAD FAQ site.\
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the ODC Open Database License\ (ODbL) as described here.\
\ \\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ \ Analysis of protein-coding genetic variation in 60,706 humans.\ Nature. 2016 Aug 18;536(7616):285-91.\ PMID: 27535533; PMC: PMC5018207\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM,\ Ganna A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\ \\ Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, Khera AV, Lowther C,\ Gauthier LD, Wang H et al.\ \ A structural variation reference for medical and population genetics.\ Nature. 2020 May;581(7809):444-451.\ PMID: 32461652; PMC: PMC7334194\
\ \\ Cummings BB, Karczewski KJ, Kosmicki JA, Seaby EG, Watts NA, Singer-Berk M, Mudge JM, Karjalainen J,\ Satterstrom FK, O'Donnell-Luria AH et al.\ \ Transcript expression-aware annotation improves rare variant interpretation.\ Nature. 2020 May;581(7809):452-458.\ PMID: 32461655; PMC: PMC7334198\
\ \ varRep 0 compositeTrack on\ dataVersion Release 4.0\ group varRep\ longLabel Genome Aggregation Database (gnomAD) Exome Sample Coverage v4.0\ maxHeightPixels 100:24:8\ parent gnomadVariants\ priority 4\ shortLabel gnomAD v4 Exome Coverage\ track gnomad4ExomeCoverage\ type bigWig\ visibility full\ wgEncodeRegTxnCaltechRnaSeqHepg2R2x75Il200SigPooled HepG2 bigWig 0 65535 Transcription of HepG2 cells from ENCODE 0 4 128 255 149 191 255 202 0 0 0 regulation 1 color 128,255,149\ longLabel Transcription of HepG2 cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegTxn\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 4\ shortLabel HepG2\ track wgEncodeRegTxnCaltechRnaSeqHepg2R2x75Il200SigPooled\ type bigWig 0 65535\ hffc6MicroC HFFc6 Micro-C hic Micro-C Chromatin Structure on HFFc6 0 4 0 0 0 127 127 127 0 0 0 regulation 1 bigDataUrl /gbdb/hg38/bbi/hic/4DNFI18Q799K.hic\ longLabel Micro-C Chromatin Structure on HFFc6\ parent hicAndMicroC off\ shortLabel HFFc6 Micro-C\ track hffc6MicroC\ type hic\ netHprcGCA_018466985v1 HG02559.mat netAlign GCA_018466985.1 chainHprcGCA_018466985v1 HG02559.mat HG02559.pri.mat.f1_v2 (May 2021 GCA_018466985.1_HG02559.pri.mat.f1_v2) HPRC project computed Chain Nets 1 4 0 0 0 255 255 0 0 0 0 hprc 0 longLabel HG02559.mat HG02559.pri.mat.f1_v2 (May 2021 GCA_018466985.1_HG02559.pri.mat.f1_v2) HPRC project computed Chain Nets\ otherDb GCA_018466985.1\ parent hprcChainNetViewnet off\ priority 20\ shortLabel HG02559.mat\ subGroups view=net sample=s020 population=afr subpop=acb hap=mat\ track netHprcGCA_018466985v1\ type netAlign GCA_018466985.1 chainHprcGCA_018466985v1\ highReproRegions Highly Reproducible Regions bigBed 9 + Highly Reproducible Regions 1 4 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/highRepro/highRepro.bb\ filterType.sampleNames multipleListOr\ filterValues.sampleNames CQ-56,CQ-7,CQ-8,HR_NA10835,HR_NA12248,HR_NA12249,HR_NA12878\ longLabel Highly Reproducible Regions\ parent highReproBeds\ shortLabel Highly Reproducible Regions\ subGroups view=beds\ track highReproRegions\ type bigBed 9 +\ wgEncodeRegDnaseUwHmecPeak HMEC Pk narrowPeak HMEC mammary epithelium DNaseI Peaks from ENCODE 1 4 255 112 85 255 183 170 1 0 0 regulation 1 color 255,112,85\ longLabel HMEC mammary epithelium DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel HMEC Pk\ subGroups view=a_Peaks cellType=HMEC treatment=n_a tissue=breast cancer=normal\ track wgEncodeRegDnaseUwHmecPeak\ wgEncodeRegDnaseUwHmecWig HMEC Sg bigWig 0 32097.2 HMEC mammary epithelium DNaseI Signal from ENCODE 0 4 255 112 85 255 183 170 0 0 0 regulation 1 color 255,112,85\ longLabel HMEC mammary epithelium DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.02847\ shortLabel HMEC Sg\ subGroups cellType=HMEC treatment=n_a tissue=breast cancer=normal\ table wgEncodeRegDnaseUwHmecSignal\ track wgEncodeRegDnaseUwHmecWig\ type bigWig 0 32097.2\ wgEncodeRegMarkH3k27acHuvec HUVEC bigWig 0 3721 H3K27Ac Mark (Often Found Near Regulatory Elements) on HUVEC Cells from ENCODE 2 4 128 212 255 191 233 255 0 0 0 regulation 1 color 128,212,255\ longLabel H3K27Ac Mark (Often Found Near Regulatory Elements) on HUVEC Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k27ac\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel HUVEC\ table wgEncodeBroadHistoneHuvecH3k27acStdSig\ track wgEncodeRegMarkH3k27acHuvec\ type bigWig 0 3721\ wgEncodeRegMarkH3k4me1Huvec HUVEC bigWig 0 4666 H3K4Me1 Mark (Often Found Near Regulatory Elements) on HUVEC Cells from ENCODE 0 4 128 212 255 191 233 255 0 0 0 regulation 1 color 128,212,255\ longLabel H3K4Me1 Mark (Often Found Near Regulatory Elements) on HUVEC Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k4me1\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel HUVEC\ table wgEncodeBroadHistoneHuvecH3k4me1StdSig\ track wgEncodeRegMarkH3k4me1Huvec\ type bigWig 0 4666\ wgEncodeRegMarkH3k4me3Huvec HUVEC bigWig 0 7852 H3K4Me3 Mark (Often Found Near Promoters) on HUVEC Cells from ENCODE 0 4 128 212 255 191 233 255 0 0 0 regulation 1 color 128,212,255\ longLabel H3K4Me3 Mark (Often Found Near Promoters) on HUVEC Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k4me3\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel HUVEC\ table wgEncodeBroadHistoneHuvecH3k4me3StdSig\ track wgEncodeRegMarkH3k4me3Huvec\ type bigWig 0 7852\ xGen_Research_Targets_V2 IDT xGen V2 T bigBed IDT - xGen Exome Research Panel V2 Target Regions 1 4 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/xgen-exome-research-panel-v2-targets-hg38.bb\ color 100,143,255\ longLabel IDT - xGen Exome Research Panel V2 Target Regions\ parent exomeProbesets on\ shortLabel IDT xGen V2 T\ track xGen_Research_Targets_V2\ type bigBed\ visibility dense\ snpArrayIllumina450k Illumina 450k bigBed 6 Illumina 450k Methylation Array 3 4 0 0 0 127 127 127 0 0 0\ The arrays listed in this track are probes from the\ Agilent Catalog Oligonucleotide Microarrays.\
\Please note that more microarray tracks are available on the hg19 genome assembly. \ To view those tracks, please \ click this link for hg19 microarrays.\ Microarrays that are not listed can be added as Custom Tracks with data from the companies.\
\ \\ Agilent's oligonucleotide CGH (Comparative Genomic Hybridization) platform enables the\ study of genome-wide DNA copy number changes at a high resolution. The CGH probes on Agilent\ CGH microarrays are 60-mer oligonucleotides synthesized in situ using Agilent's inkjet\ SurePrint technology. The probes represented on the Agilent CGH microarrays have been\ selected using algorithms developed specifically for the CGH application, assuring optimal\ performance of these probes in detecting DNA copy number changes.\
\ \\ With the Infinium MethylationEPIC BeadChip Kit, researchers can interrogate over 850,000\ methylation sites quantitatively across the genome at single-nucleotide resolution. Multiple\ samples, including FFPE, can be analyzed in parallel to deliver high-throughput power while\ minimizing the cost per sample. These tracks show positions being measured on the Illumina 450k and\ 850k (EPIC) microarray tracks, not the probe locations themselves. Contact us\ or Illumina if you need the probe locations directly. More information about\ the arrays can be found on the\ Infinium MethylationEPIC Kit website.\
\ Note: The 450k track on hg38 contains 128,989 regions representing the target regions, not the probes\ themselves.
\ \\ The Infinium CytoSNP-850K v1.2 BeadChip provides comprehensive coverage of\ cytogenetically relevant genes on a proven platform, helping researchers find valuable information\ that may be missed by other technologies. It contains approximately 850,000 empirically selected\ single nucleotide polymorphisms (SNPs) spanning the entire genome with enriched coverage for 3,262\ genes of known cytogenetics relevance in both constitutional and cancer applications. \
\ \\ The CytoScan HD Array, which is included in the\ CytoScan HD Suite, provides the broadest coverage and highest performance for\ detecting chromosomal aberrations. CytoScan HD Suite has greater than 99% sensitivity and can\ reliably detect 25-50kb copy number changes across the genome at high specificity with\ single-nucleotide polymorphism (SNP) allelic corroboration. With more than 2.6 million copy number\ markers, CytoScan HD Suite covers all OMIM and RefSeq genes.\
\ \\ Bionano Laboratories provides access to Optical Genome Mapping (OGM) data for projects across a variety of\ applications for researchers, clinicians, and pharmaceutical companies.
\This track shows the CTTAAG sites used by the \ Bionano Optical Genome Mapping system,\ an assay to detect structural variants.\
\ \\ Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \ \\ The Agilent arrays were downloaded from their \ Agilent SureDesign website tool on March 2022.
\\ The Illumina 450k and 850k (EPIC) tracks were created using a few columns from the\ Infinium MethylationEPIC v1.0 B5 Manifest File (CSV Format)\ and was then converted into a bigBed.
\\ The Illumina CytoSNP-850K track was created by downloading the\ CytoSNP-850K v1.2 Manifest File (CSV Format) (GRCh38) file and then converted\ into a bigBed file.\
\\ The Affymetrix Cytoscan HD GeneChip Array track was created by converting the \ CytoScanHD_Accel_Array.na36.bed.zip\ into a bigBed file.\
\\ The Bionano track was created by receiving the BED files from\ \ apang@bionano.\ com\ \ and converted to bigBed files using the bedToBigBed tool.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API \ or downloaded from our \ Downloads site. Please refer to our\ \ mailing list archives for questions, or our\ \ Data Access FAQ for more information.\
\ \\ Thanks to the Agilent and Illumina support teams for sharing the data and the UCSC Genome Browser\ engineers for configuring the data.
\\ Thanks to Andy Pang from Bionano Genomics for providing the BED data file.
\ varRep 1 bigDataUrl /gbdb/hg38/bbi/illumina/illumina450K.bb\ colorByStrand 255,0,0 0,0,255\ html genotypeArrays\ longLabel Illumina 450k Methylation Array\ noScoreFilter on\ parent genotypeArrays on\ priority 4\ shortLabel Illumina 450k\ track snpArrayIllumina450k\ type bigBed 6\ urls refGeneAccession="https://www.ncbi.nlm.nih.gov/nuccore/$$" rsID="https://www.ncbi.nlm.nih.gov/snp/?term=$$"\ visibility pack\ unipInterest Interest bigBed 12 + UniProt Regions of Interest 1 4 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipInterest.bb\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ itemRgb off\ longLabel UniProt Regions of Interest\ mouseOver UniProt record: $uniProtId\ The NMDetective tracks display genome-wide predictions of nonsense-mediated mRNA\ decay (NMD) efficiency from\ Lindeboom et al. 2016.\ NMDetective scores predict whether a premature termination codon (PTC) at a given position\ will trigger NMD and mRNA degradation, or whether the transcript will escape NMD and\ potentially produce a truncated protein.\
\ \\ Scores range from approximately −1 to +1. Positive values indicate that a PTC at\ that position is predicted to trigger NMD (the mRNA is degraded). Negative values indicate\ that the PTC is predicted to escape NMD (the truncated mRNA may be translated into an\ aberrant protein). Values near zero indicate intermediate or uncertain NMD efficiency.\
\ \| Track | Description |
|---|---|
| NMDetective-A | \Random forest model predicting NMD efficiency for all possible PTCs introduced\ by single-nucleotide variants. Explains ~71% of systematic variance in NMD\ efficiency. |
| NMDetective-B | \Simplified decision tree model for all possible PTCs. Slightly lower accuracy\ (~68% variance explained) but more interpretable, making it suitable for\ clinical applications. |
| NMDetective-A PTC | \Random forest model predicting NMD efficiency specifically for the first\ out-of-frame PTC introduced by frameshifting indel mutations. |
| NMDetective-B PTC | \Decision tree model for the first out-of-frame PTC from frameshifting\ indels. |
\ Each subtrack is displayed as a signal (bigWig) track. By default, the vertical axis\ ranges from −1 to +1. Regions with positive values (predicted NMD-triggering) are\ shown above the baseline; regions with negative values (predicted NMD escape) are shown\ below.\
\\ The NMDetective models were trained on somatic nonsense mutation data from 9,769 cancer\ patients and validated with frameshift mutations and germline variants\ (Lindeboom et al. 2019).\ The models incorporate the following features to predict NMD efficiency:\
\\ NMDetective-A (random forest regression) captures non-linear interactions among\ these features and achieves the highest predictive accuracy.\ NMDetective-B (decision tree) applies a simpler rule-based classification that\ is more transparent, with a modest reduction in accuracy.\
\ \\ The predictions were generated for every possible PTC-introducing single-nucleotide\ variant and for the first out-of-frame PTC from every possible single-nucleotide\ frameshifting indel across all human protein-coding transcripts. The original bedGraph\ custom track files were downloaded from the\ NMDetective Figshare page\ resource and converted to bigWig format at UCSC.\
\ \\ The data underlying these tracks can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Rik Lindeboom for providing custom tracks and the original NMDetective data\ on Figshare.\
\ \\ Lindeboom RG, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Lindeboom RGH, Vermeulen M, Lehner B, Supek F.\ \ The impact of nonsense-mediated mRNA decay on genetic disease, gene editing and cancer\ immunotherapy.\ Nat Genet. 2019 Nov;51(11):1645-1651.\ PMID: 31659324; PMC: PMC6858879\
\ \ genes 0 autoScale off\ bigDataUrl /gbdb/hg38/nmd/nmdDectA-ptc.bw\ color 0,153,102\ html nmdDetective\ longLabel NMDetective-A: Random forest NMD efficiency for first out-of-frame PTC\ maxHeightPixels 128:32:8\ parent nmd off\ priority 4\ shortLabel NMDetective-A PTC\ track nmdDetectiveA_ptc\ type bigWig\ viewLimits -1:1\ visibility hide\ notinalllowmapandsegdupregions Not lowMap+SegDup bigBed 3 Genome In a Bottle: not lowMap+SegDup mapping regions 1 4 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/GIAB/notinalllowmapandsegdupregions.bb\ longLabel Genome In a Bottle: not lowMap+SegDup mapping regions\ parent problematicGIAB on\ shortLabel Not lowMap+SegDup\ track notinalllowmapandsegdupregions\ type bigBed 3\ visibility dense\ panelAppAusGenes PanelApp Australia Genes bigBed 9 + PanelApp Australia Genes Panels 3 4 0 0 0 127 127 127 0 0 0 https://panelapp-aus.org/panels/$\ The recombination rate track represents calculated rates of recombination based\ on the genetic maps from deCODE (Halldorsson et al., 2019) and 1000 Genomes\ (2013 Phase 3 release, lifted from hg19). The deCODE map is more recent, has a higher \ resolution and was natively created on hg38 and therefore recommended. \ For the Recomb. deCODE average track, the recombination rates for chrX represent the female rate.\
\ \This track also includes a subtrack with all the\ individual deCODE recombination events and another subtrack with several thousand\ de-novo mutations found in the deCODE sequencing data. These two tracks are hidden by\ default and have to be switched on explicitly on the configuration page.\
\ \\ This is a super track that contains different subtracks, three with the deCODE\ recombination rates (paternal, maternal and average) and one with the 1000\ Genomes recombination rate (average). These tracks are in \ signal graph\ (wiggle) format. By default, to show most recombination hotspots, their maximum\ value is set to 100 cM, even though many regions have values higher than 100.\ The maximum value can be changed on the configuration pages of the tracks.\
\ \\ There are two more tracks that show additional details provided by deCODE: one\ subtrack with the raw data of all cross-overs tagged with their proband ID and\ another one with around 8000 human de-novo mutation variants that are linked to\ cross-over changes.\
\ \\ The deCODE genetic map was created at \ deCODE Genetics. It is based \ on microarrays assaying 626,828 SNP markers that allowed to identify 1,476,140 crossovers in\ 56,321 paternal meioses and 3,055,395 crossovers in 70,086 maternal meioses.\ In total, the data is based on 4,531,535 crossovers in 126,427 meioses. By\ using WGS data with 9,305,070 SNPs, the boundaries for 761,981 crossovers were\ refined: 247,942 crossovers in 9423 paternal meioses and 514,039 crossovers in\ 11,750 maternal meioses. The average resolution of the genetic map is 682 base\ pairs (bp): 655 and 708 bp for the paternal and maternal maps, respectively.\
\ \The 1000 Genomes genetic map is based on the IMPUTE genetic map based on 1000 Genomes Phase 3, on hg19 coordinates. It\ was converted to hg38 by Po-Ru Loh at the Broad Institute. After a run of \ liftOver, he post-processed the data to deal with situations in which\ consecutive map locations became much closer/farther after lifting. The\ heuristic used is sufficient for statistical phasing but may not be optimal for\ other analyses. For this reason, and because of its higher resolution, the DeCODE\ map is therefore recommended for hg38.\
\ \As with all other tracks, the data conversion commands and pointers to the\ original data files are documented in the \ makeDoc file of this track.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr17 -start=45941345 -end=45942345 http://hgdownload.soe.ucsc.edu/gbdb/hg38/recombRate/recombAvg.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ This track was produced at UCSC using data that are freely available for\ the deCODE\ and 1000 Genomes genetic maps. Thanks to Po-Ru Loh at the\ Broad Institute for providing the code to lift the hg19 1000 Genomes map data to hg38.\
\ \\ 1000 Genomes Project Consortium., Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA,\ Hurles ME, McVean GA.\ \ A map of human genome variation from population-scale sequencing.\ Nature. 2010 Oct 28;467(7319):1061-73.\ PMID: 20981092; PMC: PMC3042601\
\ \\ Halldorsson BV, Palsson G, Stefansson OA, Jonsson H, Hardarson MT, Eggertsson HP, Gunnarsson B,\ Oddsson A, Halldorsson GH, Zink F et al.\ \ Characterizing mutagenic effects of recombination through a sequence-level genetic map.\ Science. 2019 Jan 25;363(6425).\ PMID: 30679340\
\ map 1 bigDataUrl /gbdb/hg38/recombRate/events.bb\ html recombRate2.html\ longLabel Recombination events in deCODE Genetic Map (zoom to < 10kbp to see the events)\ parent recombRate2\ priority 4\ shortLabel Recomb. deCODE Evts\ track recombEvents\ type bigBed 4 +\ visibility hide\ ncbiRefSeqOther RefSeq Other bigBed 12 + NCBI RefSeq Other Annotations (not NM_*, NR_*, XM_*, XR_*, NP_* or YP_*) 1 4 32 32 32 143 143 143 0 0 0 genes 1 bigDataUrl /gbdb/hg38/ncbiRefSeq/ncbiRefSeqOther.bb\ color 32,32,32\ labelFields gene\ longLabel NCBI RefSeq Other Annotations (not NM_*, NR_*, XM_*, XR_*, NP_* or YP_*)\ parent refSeqComposite off\ priority 4\ searchIndex name\ searchTrix /gbdb/hg38/ncbiRefSeq/ncbiRefSeqOther.ix\ shortLabel RefSeq Other\ skipEmptyFields on\ track ncbiRefSeqOther\ type bigBed 12 +\ urls GeneID="https://www.ncbi.nlm.nih.gov/gene/$$" MIM="https://www.ncbi.nlm.nih.gov/omim/612091" HGNC="https://www.genenames.org/data/gene-symbol-report/#!/hgnc_id/$$" FlyBase="https://flybase.org/reports/$$" WormBase="http://www.wormbase.org/db/gene/gene?name=$$" RGD="https://rgd.mcw.edu/rgdweb/search/search.html?term=$$" SGD="https://www.yeastgenome.org/locus/$$" miRBase="http://www.mirbase.org/cgi-bin/mirna_entry.pl?acc=$$" ZFIN="https://zfin.org/$$" MGI="https://www.informatics.jax.org//marker/$$"\ joinedRmsk RepeatMasker Viz. bed 3 + Detailed Visualization of RepeatMasker Annotations 0 4 0 0 0 127 127 127 1 0 0\ This track was created using Arian Smit's\ RepeatMasker\ program, which screens DNA sequences\ for interspersed repeats and low complexity DNA sequences. The program\ outputs a detailed annotation of the repeats that are present in the\ query sequence (represented by this track), as well as a modified version\ of the query sequence in which all the annotated repeats have been masked\ (generally available on the\ Downloads page). RepeatMasker uses a separately curated version of the \ Repbase Update repeat library from the\ Genetic \ Information Research Institute (GIRI).\ Repbase Update is described in Jurka (2000) in the References section below.
\\ Alternatively, RepeatMasker can use the new\ Dfam database of repeat profile HMMs.\ Profile HMMs provide a richer description of the repeat families and when used with\ RepeatMasker + nhmmer provide a more\ sensitive approach to identifying repeats. Dfam is described in Wheeler et al. (2012)\ in the References section below.\
\ \\ In dense display mode, a single line is displayed denoting the coverage of repeats using a series\ of black boxes. \
\\ In full display mode, the track view is controlled by the scale of the view. At scales between 10 Mb\ and 30 kb, this track displays up to ten different classes of repeats (see below) one class per\ line. The repeat ranges are denoted as grayscale boxes, reflecting both the size of the repeat and\ the amount of base mismatch, base deletion, and base insertion associated with a repeat element.\ The higher the combined number of these, the lighter the shading.\
\\ In full display mode and at scales less than 30 kb, a new detailed display mode is used. Repeats\ are displayed as arrow boxes, indicating the size and orientation of the repeat. The interior\ grayscale shading represents the divergence of the repeat (see above) while the outline color\ represents the class of the repeat. Dotted lines above the repeat and extending left or right\ indicate the length of unaligned repeat consensus sequence. If the length of the unaligned sequence\ is large, a double interruption line is used to indicate that the unaligned sequence is not to scale. \
\\ For example, the following repeat is a SINE element in the forward orientation with average\ divergence. Only the 5' proximal fragment of the consensus sequence is aligned to the genome.\ The 3' unaligned length (384bp) is not drawn to scale and is instead displayed using a set of\ interruption lines along with the length of the unaligned sequence.\
\ \ \ \\ Repeats that have been fragmented by insertions or large internal deletions are now represented\ by join lines. In the example below, a LINE element is found as two fragments. The solid\ connection lines indicate that there are no unaligned consensus bases between the two fragments.\ Also note these fragments represent the end of the repeat, as there is no unaligned consensus\ sequence following the last fragment.\
\ \ \ \\ In cases where there is unaligned consensus sequence between the fragments, the repeat will look like\ the following. The dotted line indicates the length of the unaligned sequence between the two\ fragments. In this case the unaligned consensus is longer than the actual genomic distance between\ these two fragments.\
\ \ \ \\ If there is consensus overlap between the two fragments, the joining lines will be drawn to indicate\ how much of the left fragment is repeated in the right fragment. \
\ \ \ \\ The following table lists the repeat class colors:\
\ \| Color | \Repeat Class | \
|---|---|
| \ | SINE - Short Interspersed Nuclear Element | \
| \ | LINE - Long Interspersed Nuclear Element | \
| \ | LTR - Long Terminal Repeat | \
| \ | DNA - DNA Transposon | \
| \ | Simple - Single Nucleotide Stretches and Tandem Repeats | \
| \ | Low_complexity - Low Complexity DNA | \
| \ | Satellite - Satellite Repeats | \
| \ | RNA - RNA Repeats (including RNA, tRNA, rRNA, snRNA, scRNA, srpRNA) | \
| \ | Other - Other Repeats (including class RC - Rolling Circle) | \
| \ | Unknown - Unknown Classification | \
\ A "?" at the end of the "Family" or "Class" (for example, DNA?)\ signifies that the curator was unsure of the classification. At some point in the future,\ either the "?" will be removed or the classification will be changed.
\ \\ UCSC has used the most current versions of the RepeatMasker software\ and repeat libraries available to generate these data. Note that these\ versions may be newer than those that are publicly available on the Internet.\
\\ Data are generated using the RepeatMasker -s flag. Additional flags\ may be used for certain organisms. Repeats are soft-masked. Alignments may\ extend through repeats, but are not permitted to initiate in them.\ See the FAQ for more information.\
\ \\ Thanks to Arian Smit, Robert Hubley and GIRI for providing the tools and\ repeat libraries used to generate this track.\
\ \\ Smit AFA, Hubley R, Green P. RepeatMasker Open-3.0.\ \ https://www.repeatmasker.org/. 1996-2010.\
\ \\ Dfam is described in:\
\\ Wheeler TJ, Clements J, Eddy SR, Hubley R, Jones TA, Jurka J, Smit AF, Finn RD.\ \ Dfam: a database of repetitive DNA based on profile hidden Markov models.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D70-82.\ PMID: 23203985; PMC: PMC3531169\
\ \\ Repbase Update is described in:\
\\ Jurka J.\ \ Repbase Update: a database and an electronic journal of repetitive elements.\ Trends Genet. 2000 Sep;16(9):418-420.\ PMID: 10973072\
\ \\ For a discussion of repeats in mammalian genomes, see:\
\\ Smit AF.\ \ Interspersed repeats and other mementos of transposable elements in mammalian genomes.\ Curr Opin Genet Dev. 1999 Dec;9(6):657-63.\ PMID: 10607616\
\ \\ Smit AF.\ \ The origin of interspersed repeats in the human genome.\ Curr Opin Genet Dev. 1996 Dec;6(6):743-8.\ PMID: 8994846\
\ rep 0 allButtonPair on\ canPack off\ compositeTrack on\ group rep\ html joinedRmsk\ longLabel Detailed Visualization of RepeatMasker Annotations\ maxWindowToDraw 10000000\ priority 4\ shortLabel RepeatMasker Viz.\ spectrum on\ track joinedRmsk\ type bed 3 +\ visibility hide\ rmskJoinedBaseline RepeatMasker Viz. bed 3 + RepeatMasker v3.0.1 db20100302 : Browser Baseline Dataset 0 4 0 0 0 127 127 127 1 0 0 rep 0 group rep\ longLabel RepeatMasker v3.0.1 db20100302 : Browser Baseline Dataset\ parent joinedRmsk on\ priority 4\ shortLabel RepeatMasker Viz.\ track rmskJoinedBaseline\ visibility hide\ gnomad315XPercentage Sample % > 15X bigWig gnomAD Percentage of Genome Samples with at least 15X Coverage v3.0.1 2 4 165 0 90 210 127 172 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_15.bw\ color 165,0,90\ longLabel gnomAD Percentage of Genome Samples with at least 15X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 4\ shortLabel Sample % > 15X\ track gnomad315XPercentage\ viewLimits 0:1\ gnomad4Exome15XPercentage Sample % > 15X bigWig gnomAD Percentage of Exome Samples with at least 15X Coverage v4.0 2 4 165 0 90 210 127 172 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_15.bw\ color 165,0,90\ longLabel gnomAD Percentage of Exome Samples with at least 15X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 4\ shortLabel Sample % > 15X\ track gnomad4Exome15XPercentage\ viewLimits 0:1\ shorthcondels Short hConDels bigBed 4 + short hConDels: 10032 Short Human hCondels - Human Conserved Deletions < 40bp 0 4 0 0 0 127 127 127 0 0 0 compGeno 1 bigDataUrl /gbdb/hg38/unusualcons/hcondels.bb\ longLabel short hConDels: 10032 Short Human hCondels - Human Conserved Deletions < 40bp\ parent unusualcons on\ shortLabel Short hConDels\ track shorthcondels\ type bigBed 4 +\ sgdp Simons Genome Diversity Project 0.3k WGS vcfTabix Phased Variants: Simons Genome Diversity Project - 279 samples, unmixed populations 3 4 0 0 0 127 127 127 0 0 0\ This tracks contains variants of individual genotypes, usually phased, from the projects\ Human Diversity Genome Project, Simons Genome Diversity Project, gnomad's HGDP+1000 Genomes callset,\ and the Mexico Biobank.\ The original release of 1000 Genomes has its own, separate track.\ Projects where the released variants are not phased can be found in the container track "SNV Frequencies".\
\ \\ Available on hg19 and hg38:
\\ Available only on hg38:
\\ Full haplotype display:\ In "pack" mode, this track sorts the haplotypes. This can be\ useful for determining the similarity between the samples and inferring\ inheritance at a particular locus.\ Each sample's phased and/or homozygous genotypes are split into haplotypes,\ clustered by similarity around a central variant (in pink), and sorted for\ display by their position in the clustering tree. Click a variant to center on it.\ The tree (as space allows) is drawn in the label area next to the track image.\ Leaf clusters, in which all haplotypes are identical (at least for the variants\ used in clustering), are colored purple. \
\\ For a full description of how the display works, please see our \ Haplotype Display help page.\ \
\ MXB: Allele frequencies by geographical state and ancestry are available via\ the MexVar platform.\ Raw genotype data are available under controlled access at the\ EGA (Study: EGAS00001005797; Dataset: EGAD00010002361). For the VCFs, email\ andres.moreno@cinvestav.mx.\
\ \\ SGDP: The version used was\ https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/,\ merged with bcftools and lifted to hg38 with CrossMap. \
\ \\ MXB: We thank the Center for Research and Advanced Studies (Cinvestav) of Mexico for\ generating and providing the frequency data, the National Institute of Medical\ Sciences and Nutrition (INCMNSZ) for DNA extraction, and the Ministry of Health\ together with the National Institute of Public Health (INSP) for the design and\ implementation of the National Health Survey 2000 (ENSA 2000). We also thank\ the ENSA-Genomics Consortium for their contributions to sample collection and\ data processing that made possible the construction of the MXB genomic\ resource.\
\\ SGDP: This project was funded by the Simons Foundation. Thanks to David Reich and Swapan \ Mallick for help with importing the data.\
\ \\ Barberena-Jonas C, Medina-Muñoz SG, Cedillo-Castelán V, Sepúlveda-Morales T,\ Gonzaga-Jáuregui C, ENSA Genomics Consortium, García-García L, Ioannidis AG,\ Moreno-Estrada A.\ \ Clinical genetic variation across Hispanic populations in the Mexican Biobank.\ Nat Med. 2026 Jan 21;.\ DOI: 10.1038/s41591-025-04100-z; PMID: 41566040\
\ \\ Sohail M, Moreno-Estrada A.\ \ The Mexican Biobank Project promotes genetic discovery, inclusive science and local capacity\ building.\ Dis Model Mech. 2024 Jan 1;17(1).\ PMID: 38299665; PMC: PMC10855211\
\ \\ Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Corés CD, Barberena-Jonas C, Medina-Muñoz SG,\ Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L et al.\ \ Mexican Biobank advances population and medical genomics of diverse ancestries.\ Nature. 2023 Oct;622(7984):775-783.\ PMID: 37821706; PMC: PMC10600006\
\ \\ Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J\ et al.\ \ Insights into human genetic variation and population history from 929 diverse genomes.\ Science. 2020 Mar 20;367(6484).\ PMID: 32193295; PMC: PMC7115999\
\ \\ Koenig Z, Yohannes MT, Nkambule LL, Zhao X, Goodrich JK, Kim HA, Wilson MW, Tiao G, Hao SP, Sahakian\ N et al.\ \ A harmonized public resource of deeply sequenced diverse human genomes.\ Genome Res. 2024 Jun 25;34(5):796-809.\ PMID: 38749656; PMC: PMC11216312\
\ \\ Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S,\ Tandon A et al.\ \ The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.\ Nature. 2016 Oct 13;538(7624):201-206.\ PMID: 27654912; PMC: PMC5161557\
\ \ varRep 1 bigDataUrl /gbdb/hg38/phasedVars/sgdp/SGDP.nh2.vcf.gz\ dataVersion 2016-12-07 public (hg38 lift)\ html phasedVars.html\ longLabel Phased Variants: Simons Genome Diversity Project - 279 samples, unmixed populations\ parent phasedVars on\ priority 4\ shortLabel Simons Genome Diversity Project 0.3k WGS\ track sgdp\ type vcfTabix\ visibility pack\ spliceAiDonorMinus SpliceAI Donor Minus bigWig 0 1 SpliceAI Splice Donor Sites, Minus Strand 2 4 0 0 0 127 127 127 0 0 0 phenDis 0 bigDataUrl /gbdb/hg38/bbi/spliceAi/wildtype/spliceAiDonorMinus.bw\ longLabel SpliceAI Splice Donor Sites, Minus Strand\ parent spliceAIWt on\ priority 4\ shortLabel SpliceAI Donor Minus\ track spliceAiDonorMinus\ type bigWig 0 1\ umap100 Umap S100 bigBed 6 Single-read mappability with 100-mers 0 4 80 170 240 167 212 247 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k100.Unique.Mappability.bb\ color 80,170,240\ longLabel Single-read mappability with 100-mers\ parent umapBigBed off\ priority 4\ shortLabel Umap S100\ subGroups view=SR\ track umap100\ visibility hide\ chainMm39 Mouse Chain chain mm39 Mouse (Jun. 2020 (GRCm39/mm39)) Chained Alignments 3 5 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Mouse (Jun. 2020 (GRCm39/mm39)) Chained Alignments\ otherDb mm39\ parent placentalChainNetViewchain off\ shortLabel Mouse Chain\ subGroups view=chain species=s012a clade=c00\ track chainMm39\ type chain mm39\ chainGalGal6 Chicken Chain chain galGal6 Chicken (Mar. 2018 (GRCg6a/galGal6)) Chained Alignments 3 5 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Chicken (Mar. 2018 (GRCg6a/galGal6)) Chained Alignments\ otherDb galGal6\ parent vertebrateChainNetViewchain off\ shortLabel Chicken Chain\ subGroups view=chain species=s008a clade=c01\ track chainGalGal6\ type chain galGal6\ chainGorGor6 Gorilla Chain chain gorGor6 Gorilla (Aug. 2019 (Kamilah_GGO_v0/gorGor6)) Chained Alignments 3 5 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Gorilla (Aug. 2019 (Kamilah_GGO_v0/gorGor6)) Chained Alignments\ otherDb gorGor6\ parent primateChainNetViewchain off\ shortLabel Gorilla Chain\ subGroups view=chain species=s009a clade=c00\ track chainGorGor6\ type chain gorGor6\ encTfChipPkENCFF047UIF A549 CEBPB narrowPeak Transcription Factor ChIP-seq Peaks of CEBPB in A549 from ENCODE 3 (ENCFF047UIF) 0 5 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of CEBPB in A549 from ENCODE 3 (ENCFF047UIF)\ parent encTfChipPk off\ shortLabel A549 CEBPB\ subGroups cellType=A549 factor=CEBPB\ track encTfChipPkENCFF047UIF\ cloneEndABC14 ABC14 bed 12 Agencourt fosmid library 14 0 5 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 14\ parent cloneEndSuper off\ priority 5\ shortLabel ABC14\ subGroups source=agencourt\ track cloneEndABC14\ type bed 12\ visibility hide\ AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep1LK4_CNhs13340_ctss_fwd AorticSmsToFgf2_00hr15minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep1 (LK4)_CNhs13340_12643-134G6_forward 0 5 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12643-134G6 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr15min%2c%20biol_rep1%20%28LK4%29.CNhs13340.12643-134G6.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep1 (LK4)_CNhs13340_12643-134G6_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12643-134G6 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr15minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep1LK4_CNhs13340_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12643-134G6\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep1LK4_CNhs13340_tpm_fwd AorticSmsToFgf2_00hr15minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep1 (LK4)_CNhs13340_12643-134G6_forward 1 5 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12643-134G6 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr15min%2c%20biol_rep1%20%28LK4%29.CNhs13340.12643-134G6.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep1 (LK4)_CNhs13340_12643-134G6_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12643-134G6 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr15minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep1LK4_CNhs13340_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12643-134G6\ urlLabel FANTOM5 Details:\ gtexCovArteryCoronary Artery Coron bigWig Artery Coronary 0 5 238 106 80 246 180 167 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-1GMR3-0626-SM-9WYT3.Artery_Coronary.RNAseq.bw\ color 238,106,80\ longLabel Artery Coronary\ parent gtexCov\ shortLabel Artery Coron\ track gtexCovArteryCoronary\ bismap24Neg Bismap S24 - bigBed 6 Single-read mappability with 24-mers after bisulfite conversion (reverse strand) 1 5 240 20 80 247 137 167 0 0 0 map 1 bigDataUrl /gbdb/hg38/hoffmanMappability/k24.G2A-Converted.bb\ color 240,20,80\ longLabel Single-read mappability with 24-mers after bisulfite conversion (reverse strand)\ parent bismapBigBed on\ priority 5\ shortLabel Bismap S24 -\ subGroups view=SR\ track bismap24Neg\ visibility dense\ wgEncodeReg4TxnBloodPlus Blood + bigWig Avg. + strand total RNA-seq level of 68 blood experiments (tissues and primary cells only) 0 5 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBloodPlus.bw\ color 254,75,173\ longLabel Avg. + strand total RNA-seq level of 68 blood experiments (tissues and primary cells only)\ parent wgEncodeReg4Txn\ priority 5\ shortLabel Blood +\ track wgEncodeReg4TxnBloodPlus\ type bigWig\ wgEncodeReg4DnaseBreast Breast bigWig Avg. DNase level of 5 breast experiments (tissues and primary cells only) 0 5 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBreastDNase.bw\ color 65,171,173\ longLabel Avg. DNase level of 5 breast experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 5\ shortLabel Breast\ track wgEncodeReg4DnaseBreast\ type bigWig\ lincRNAsCTBreast Breast bed 5 + lincRNAs from breast 1 5 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from breast\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Breast\ subGroups view=lincRNAsRefseqExp tissueType=breast\ track lincRNAsCTBreast\ CESC CESC bigLolly 12 + Cervical squamous cell carcinoma and endocervical adenocarcinoma 0 5 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/CESC.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Cervical squamous cell carcinoma and endocervical adenocarcinoma\ parent gdcCancer off\ priority 5\ shortLabel CESC\ track CESC\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ chinamap China ChinaMAP 10.5k WGS vcfTabix SNV Frequencies: ChinaMAP phase 1 - 10,588 WGS at ~40x, Chinese natural population 0 5 0 0 0 127 127 127 0 0 0\ This track shows allele frequencies for 147.4 million variants (136.7\ million SNPs and 10.7 million short indels, autosomes only) from\ 10,588 Chinese individuals deep-whole-genome-sequenced at a mean depth\ of about 40x by the China Metabolic Analytics Project (ChinaMAP).\ Participants come from three large Chinese cohort studies (the China\ Noncommunicable Disease Surveillance, the REACTION study and the\ Community-based Cardiovascular Risk During Urbanization in Shanghai\ study) and span 27 provinces of China and eight ethnic populations\ (Han, Hui, Manchu, Miao, Mongolian, Yi, Tibetan and Zhuang). For\ each variant the track records the cohort allele count, allele number\ and allele frequency. The original release also ships the matched 1000\ Genomes Project (1KGP) allele frequencies (global, EAS, AMR, AFR, EUR\ and SAS) as INFO fields, which are kept verbatim in the VCF.\
\ \\ The track uses the standard UCSC VCF display. When you hover over a\ variant, the popup shows the cohort allele frequency and count, the\ total number of called alleles, and the 1KGP frequencies that the\ ChinaMAP release ships alongside each site.\
\ \\ DNA from each participant was prepared with the QIAGEN DNeasy\ Blood & Tissue Kit, sheared by Covaris, ligated to BGISEQ-500\ adapters and rolling-circle amplified into DNA nanoballs for\ 100 bp paired-end sequencing on the BGISEQ-500 platform at BGI\ Genomics. Reads were quality-filtered with SOAPnuke v1.5.6, aligned\ to GRCh38 (GENCODE release) with BWA-MEM v0.7.16a, coordinate-sorted\ with Picard SortSam v2.13.2, and duplicate-marked and base-quality\ recalibrated with GATK v4.beta.4. Samples were required to pass six\ QC criteria (base quality Q30 > 80%, mean depth > 30x, mapping\ rate ≥ 95%, mismatch rate < 1%, duplicate rate < 10% and\ 20x coverage > 80%) and a 21-SNP mass spectrometric fingerprint\ check; 10,588 WGS samples passed. Germline variants were called\ per-sample as GVCFs with GATK HaplotypeCaller v4.0.4.0, combined\ with GATK CombineGVCFs and joint-called with GATK GenotypeGVCFs\ (v4.0.4.0), ignoring low-complexity regions. Variants were filtered\ with GATK VariantFiltration and restricted to length ≤ 10 bp and a\ maximum of 10 alt alleles. Multi-allelic sites were split, and the\ final callset was annotated with SnpEff v4.3. See Cao et al.\ 2020 (in References below) for the full pipeline.\
\\ The bgzipped sites-only VCF\ (mbiobank_ChinaMAP.phase1.vcf.gz) was downloaded from the\ ChinaMAP / mBiobank distribution site\ (http://chinamapwgs.mbiobank.com/download/),\ renamed locally to chinamap.vcf.gz and tabix-indexed. We did\ not need to lift over coordinates or reformat the file: the upstream\ file is already on GRCh38 with chr-prefixed chromosome names,\ autosomes only, and ships standard AC, AF and\ AN INFO fields. The pipeline is recorded in the\ makeDoc\ file of the track.\
\ \\ Only autosomes (chr1-22) are present; chrX, chrY and chrM are not\ in the ChinaMAP phase 1 release. The 1KGP frequency fields\ (1KGP_AF, 1KGP_EAS_AF, 1KGP_AMR_AF,\ 1KGP_AFR_AF, 1KGP_EUR_AF, 1KGP_SAS_AF) are\ carried over verbatim from the ChinaMAP VCF and only populate the\ small fraction of ChinaMAP sites that are also catalogued in the\ matched 1KGP release.\
\ \\ The ChinaMAP Limitations on Use (see the\ ChinaMAP\ download page) prohibit redistribution of the data, so the\ ChinaMAP VCF is not available from the UCSC Table Browser, Data\ Integrator, REST API or the public download server. The track can be\ browsed interactively in the Genome Browser; for bulk access please\ register with the ChinaMAP project at\ http://chinamapwgs.mbiobank.com/\ and download the original VCF directly from them.\
\ \\ Thanks to the ChinaMAP participants and to the National Clinical\ Research Center for Metabolic Diseases (Shanghai Jiao Tong\ University School of Medicine, Ruijin Hospital) and BGI Genomics, who\ produced and released the ChinaMAP phase 1 sites VCF.\
\ \\ Cao Y, Li L, Xu M, Feng Z, Sun X, Lu J, Xu Y, Du P, Wang T, Hu R et al.\ \ The ChinaMAP analytics of deep whole genome sequences in 10,588 individuals.\ Cell Res. 2020 Sep;30(9):717-731.\ PMID: 32355288; PMC: PMC7609296\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_chinamap/chinamap.vcf.gz\ dataVersion Phase 1 (v2020-03.beta)\ longLabel SNV Frequencies: ChinaMAP phase 1 - 10,588 WGS at ~40x, Chinese natural population\ parent varFreqs on\ priority 5\ shortLabel China ChinaMAP 10.5k WGS\ tableBrowser off\ track chinamap\ type vcfTabix\ visibility hide\ wgEncodeReg4MarkH3k27acConnectiveTissue Connective tissue bigWig H3K27ac level of 1 connective tissue experiment (tissues and primary cells only) 2 5 138 135 169 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpConnectiveTissueH3K27ac.bw\ color 138,135,169\ longLabel H3K27ac level of 1 connective tissue experiment (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 5\ shortLabel Connective tissue\ track wgEncodeReg4MarkH3k27acConnectiveTissue\ type bigWig\ wgEncodeReg4MarkH3k4me3ConnectiveTissue Connective tissue bigWig Avg. H3K4me3 level of 2 connective tissue experiments (tissues and primary cells only) 0 5 138 135 169 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpConnectiveTissueH3K4me3.bw\ color 138,135,169\ longLabel Avg. H3K4me3 level of 2 connective tissue experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 5\ shortLabel Connective tissue\ track wgEncodeReg4MarkH3k4me3ConnectiveTissue\ type bigWig\ cortexNeuron42F Cortex - Neuron - Z0000042F bigWig Methylation Atlas: Cortex - Neuron - Z0000042F 2 5 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/cortexNeuron42F.bw\ color 138,43,226\ longLabel Methylation Atlas: Cortex - Neuron - Z0000042F\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 5\ shortLabel Cortex - Neuron - Z0000042F\ subGroups cellType=Neuron dataType=Replicate\ track cortexNeuron42F\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ dbVar_common_abel dbVar Curated Abel SVs bigBed 9 + . NCBI dbVar Curated Common SVs: all populations from Abel 3 5 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/variants/$$ varRep 1 bigDataUrl /gbdb/hg38/bbi/dbVar/common_abel.bb\ longLabel NCBI dbVar Curated Common SVs: all populations from Abel\ parent dbVar_common off\ priority 5\ shortLabel dbVar Curated Abel SVs\ track dbVar_common_abel\ type bigBed 9 + .\ url https://www.ncbi.nlm.nih.gov/dbvar/variants/$$\ urlLabel NCBI Variant Page:\ ENCFF801REE_ENCFF053KMZ_ENCFF860MMV_ENCFF804PBU ENCFF801REE_ENCFF053KMZ_ENCFF860MMV_ENCFF804PBU bigBed 9 + 5 Adrenal gland, male adult (37 years): (1) cCREs 4 5 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF801REE_ENCFF053KMZ_ENCFF860MMV_ENCFF804PBU.bb\ longLabel Adrenal gland, male adult (37 years): (1) cCREs\ mouseOver ID: ${name}\ The GENCODE Genes track (version 43, February 2023) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The following table provides statistics for the v43 release derived from the GTF file that contains\ annotations only on the main chromosomes. More information on how they were generated can be found\ in the GENCODE site.
\ \\
\ \\
\ GENCODE v43 Release Stats \ Genes Observed Transcripts Observed \ Protein-coding genes 19,393 Protein-coding transcripts 89,411 \ Long non-coding RNA genes 19,928 - full length protein-coding 64,004 \ Small non-coding RNA genes 7,566 - partial length protein-coding 25,407 \ Pseudogenes 14,737 Nonsense mediated decay transcripts 21,354 \ Immunoglobulin/T-cell receptor gene segments 410 Long non-coding RNA loci transcripts 58,023 \ Total No of distinct translations 65,519 Genes that have more than one distinct translations 13,618
\
\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \ \\ Within a gene using the pack display mode, transcripts below a specified rank will be\ condensed into a view similar to squish mode. The transcript ranking approach is\ preliminary and will change in future releases. The transcripts rankings are defined by the\ following criteria for protein-coding and non-coding genes:
\ Protein_coding genes\\
The GENCODE v43 track was built from the GENCODE downloads file \
gencode.v43.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources \
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney.
\ \\ Harrow J, Frankish A, Gonzalez JM, Tapanari E, Diekhans M, Kokocinski F, Aken BL, Barrell D, Zadissa\ A, Searle S et al.\ \ GENCODE: the reference human genome annotation for The ENCODE Project.\ Genome Res. 2012 Sep;22(9):1760-74.\ PMID: 22955987; PMC: PMC3431492\
\ \\ Harrow J, Denoeud F, Frankish A, Reymond A, Chen CK, Chrast J, Lagarde J, Gilbert JG, Storey R,\ Swarbreck D et al.\ \ GENCODE: producing a reference annotation for ENCODE.\ Genome Biol. 2006;7 Suppl 1:S4.1-9.\ PMID: 16925838; PMC: PMC1810553\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV43.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV43\ group genes\ html knownGeneV43\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V43\ maxItems 50000\ parent knownGeneArchive\ priority 5\ searchIndex name\ shortLabel GENCODE V43\ squishyPackField rank\ squishyPackPoint 2\ track knownGeneV43\ type bigGenePred\ visibility hide\ geneHancerRegElements GH Reg Elems bigBed 9 + Enhancers and promoters from GeneHancer 1 5 0 0 0 127 127 127 0 0 0 http://www.genecards.org/Search/Keyword?queryString=$$ regulation 1 bigDataUrl /gbdb/hg38/geneHancer/geneHancerRegElementsAll.hg38.bb\ longLabel Enhancers and promoters from GeneHancer\ parent ghGeneHancer off\ shortLabel GH Reg Elems\ subGroups set=b_ALL view=a_GH\ track geneHancerRegElements\ hgdp1k gnomAD HGDP+1000G 4k WGS vcfTabix Phased Variants: gnomAD HGDP + 1000 genomes callset - 4094 whole genomes, 80 populations 3 5 0 0 0 127 127 127 0 0 0\ This tracks contains variants of individual genotypes, usually phased, from the projects\ Human Diversity Genome Project, Simons Genome Diversity Project, gnomad's HGDP+1000 Genomes callset,\ and the Mexico Biobank.\ The original release of 1000 Genomes has its own, separate track.\ Projects where the released variants are not phased can be found in the container track "SNV Frequencies".\
\ \\ Available on hg19 and hg38:
\\ Available only on hg38:
\\ Full haplotype display:\ In "pack" mode, this track sorts the haplotypes. This can be\ useful for determining the similarity between the samples and inferring\ inheritance at a particular locus.\ Each sample's phased and/or homozygous genotypes are split into haplotypes,\ clustered by similarity around a central variant (in pink), and sorted for\ display by their position in the clustering tree. Click a variant to center on it.\ The tree (as space allows) is drawn in the label area next to the track image.\ Leaf clusters, in which all haplotypes are identical (at least for the variants\ used in clustering), are colored purple. \
\\ For a full description of how the display works, please see our \ Haplotype Display help page.\ \
\ MXB: Allele frequencies by geographical state and ancestry are available via\ the MexVar platform.\ Raw genotype data are available under controlled access at the\ EGA (Study: EGAS00001005797; Dataset: EGAD00010002361). For the VCFs, email\ andres.moreno@cinvestav.mx.\
\ \\ SGDP: The version used was\ https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/,\ merged with bcftools and lifted to hg38 with CrossMap. \
\ \\ MXB: We thank the Center for Research and Advanced Studies (Cinvestav) of Mexico for\ generating and providing the frequency data, the National Institute of Medical\ Sciences and Nutrition (INCMNSZ) for DNA extraction, and the Ministry of Health\ together with the National Institute of Public Health (INSP) for the design and\ implementation of the National Health Survey 2000 (ENSA 2000). We also thank\ the ENSA-Genomics Consortium for their contributions to sample collection and\ data processing that made possible the construction of the MXB genomic\ resource.\
\\ SGDP: This project was funded by the Simons Foundation. Thanks to David Reich and Swapan \ Mallick for help with importing the data.\
\ \\ Barberena-Jonas C, Medina-Muñoz SG, Cedillo-Castelán V, Sepúlveda-Morales T,\ Gonzaga-Jáuregui C, ENSA Genomics Consortium, García-García L, Ioannidis AG,\ Moreno-Estrada A.\ \ Clinical genetic variation across Hispanic populations in the Mexican Biobank.\ Nat Med. 2026 Jan 21;.\ DOI: 10.1038/s41591-025-04100-z; PMID: 41566040\
\ \\ Sohail M, Moreno-Estrada A.\ \ The Mexican Biobank Project promotes genetic discovery, inclusive science and local capacity\ building.\ Dis Model Mech. 2024 Jan 1;17(1).\ PMID: 38299665; PMC: PMC10855211\
\ \\ Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Corés CD, Barberena-Jonas C, Medina-Muñoz SG,\ Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L et al.\ \ Mexican Biobank advances population and medical genomics of diverse ancestries.\ Nature. 2023 Oct;622(7984):775-783.\ PMID: 37821706; PMC: PMC10600006\
\ \\ Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J\ et al.\ \ Insights into human genetic variation and population history from 929 diverse genomes.\ Science. 2020 Mar 20;367(6484).\ PMID: 32193295; PMC: PMC7115999\
\ \\ Koenig Z, Yohannes MT, Nkambule LL, Zhao X, Goodrich JK, Kim HA, Wilson MW, Tiao G, Hao SP, Sahakian\ N et al.\ \ A harmonized public resource of deeply sequenced diverse human genomes.\ Genome Res. 2024 Jun 25;34(5):796-809.\ PMID: 38749656; PMC: PMC11216312\
\ \\ Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S,\ Tandon A et al.\ \ The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.\ Nature. 2016 Oct 13;538(7624):201-206.\ PMID: 27654912; PMC: PMC5161557\
\ \ varRep 1 bigDataUrl /gbdb/hg38/phasedVars/hgdp1k/gnomad.genomes.v3.1.2.hgdp_tgp.vcf.gz\ dataVersion v 3.1.2\ html phasedVars.html\ longLabel Phased Variants: gnomAD HGDP + 1000 genomes callset - 4094 whole genomes, 80 populations\ parent phasedVars on\ priority 5\ shortLabel gnomAD HGDP+1000G 4k WGS\ track hgdp1k\ type vcfTabix\ visibility pack\ gnomad3Coverage gnomAD v3 Genome Coverage bigWig Genome Aggregation Database (gnomAD) Genome Sample Coverage v3.0.1 2 5 0 0 0 127 127 127 0 0 0\ The Genome Aggregation Database (gnomAD) v3 - Genome Coverage track shows how many\ times regions of the genomes were sequenced. This track includes several subtracks of average\ coverage metrics and sample percentage of coverage.\
\\ There is no gnomAD v4 genome coverage track because the genomes were unchanged from V3. There is no\ gnomAD v3 exomes track because v3 was a genome-only release.
\ \\ The Average Sample Coverage tracks display the mean and median read depth of the\ samples at each base position. The details page shows calculated sample percentages for the range\ of sequence within the browser window.\
\ \\ The nX Coverage Percentage tracks display the percentage of samples whose read\ depth is at least 1X, 5X, 10X, 15X, 20X, 25X, 30X, 50X, and 100X at each base position. The details\ page shows calculated sample percentages for the range of sequence within the browser window.\
\ \\ Coverage was computed using all 71,702 gnomAD v3.01 samples from their gVCFs. The gVCFs were\ produced using a 3-bin blocking scheme:\
\\ The coverage was binned by quality using the thresholds above and the median coverage value for each\ of the resulting coverage blocks was used to compute the coverage metrics presented in the browser.\ Coverage was computed for all callable bases in the genome (all non-N bases, minus telomeres and\ centromeres).\
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that \ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below). Coverage values\ for the genome are in bigWig files in\ the coverage/ subdirectory. Variant VCFs can be found in the vcf/ subdirectory.
\\ The data can also be found directly from the gnomAD downloads page. Please refer to \ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ More information about using and understanding the gnomAD data can be found in the\ gnomAD FAQ site.\
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the ODC Open Database License\ (ODbL) as described here.\
\ \\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ \ Analysis of protein-coding genetic variation in 60,706 humans.\ Nature. 2016 Aug 18;536(7616):285-91.\ PMID: 27535533; PMC: PMC5018207\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM,\ Ganna A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\ \\ Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, Khera AV, Lowther C,\ Gauthier LD, Wang H et al.\ \ A structural variation reference for medical and population genetics.\ Nature. 2020 May;581(7809):444-451.\ PMID: 32461652; PMC: PMC7334194\
\ \\ Cummings BB, Karczewski KJ, Kosmicki JA, Seaby EG, Watts NA, Singer-Berk M, Mudge JM, Karjalainen J,\ Satterstrom FK, O'Donnell-Luria AH et al.\ \ Transcript expression-aware annotation improves rare variant interpretation.\ Nature. 2020 May;581(7809):452-458.\ PMID: 32461655; PMC: PMC7334198\
\ \ varRep 0 compositeTrack on\ dataVersion Release 3.0.1\ group varRep\ longLabel Genome Aggregation Database (gnomAD) Genome Sample Coverage v3.0.1\ maxHeightPixels 100:24:8\ parent gnomadVariants\ priority 5\ shortLabel gnomAD v3 Genome Coverage\ track gnomad3Coverage\ type bigWig\ visibility full\ chainHprcGCA_018467015v1 HG02486.mat chain GCA_018467015.1 HG02486.mat HG02486.pri.mat.f1_v2 (May 2021 GCA_018467015.1_HG02486.pri.mat.f1_v2) HPRC project computed Chained Alignments 3 5 0 0 0 255 255 0 1 0 0 hprc 1 longLabel HG02486.mat HG02486.pri.mat.f1_v2 (May 2021 GCA_018467015.1_HG02486.pri.mat.f1_v2) HPRC project computed Chained Alignments\ otherDb GCA_018467015.1\ parent hprcChainNetViewchain off\ priority 22\ shortLabel HG02486.mat\ subGroups view=chain sample=s022 population=afr subpop=acb hap=mat\ track chainHprcGCA_018467015v1\ type chain GCA_018467015.1\ hr_na10835Vcf HR_NA10835 Variants vcfTabix HR_NA10835 Variants 0 5 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/highRepro/HR_NA10835.sort.vcf.gz\ longLabel HR_NA10835 Variants\ parent highReproVcfs\ shortLabel HR_NA10835 Variants\ subGroups view=vcfs\ track hr_na10835Vcf\ type vcfTabix\ wgEncodeRegTxnCaltechRnaSeqHsmmR2x75Il200SigPooled HSMM bigWig 0 65535 Transcription of HSMM cells from ENCODE 0 5 128 255 242 191 255 248 0 0 0 regulation 1 color 128,255,242\ longLabel Transcription of HSMM cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegTxn\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 5\ shortLabel HSMM\ track wgEncodeRegTxnCaltechRnaSeqHsmmR2x75Il200SigPooled\ type bigWig 0 65535\ snpArrayIllumina850k Illumina 850k bigBed 6 Illumina 850k EPIC Methylation Array 3 5 0 0 0 127 127 127 0 0 0\ The arrays listed in this track are probes from the\ Agilent Catalog Oligonucleotide Microarrays.\
\Please note that more microarray tracks are available on the hg19 genome assembly. \ To view those tracks, please \ click this link for hg19 microarrays.\ Microarrays that are not listed can be added as Custom Tracks with data from the companies.\
\ \\ Agilent's oligonucleotide CGH (Comparative Genomic Hybridization) platform enables the\ study of genome-wide DNA copy number changes at a high resolution. The CGH probes on Agilent\ CGH microarrays are 60-mer oligonucleotides synthesized in situ using Agilent's inkjet\ SurePrint technology. The probes represented on the Agilent CGH microarrays have been\ selected using algorithms developed specifically for the CGH application, assuring optimal\ performance of these probes in detecting DNA copy number changes.\
\ \\ With the Infinium MethylationEPIC BeadChip Kit, researchers can interrogate over 850,000\ methylation sites quantitatively across the genome at single-nucleotide resolution. Multiple\ samples, including FFPE, can be analyzed in parallel to deliver high-throughput power while\ minimizing the cost per sample. These tracks show positions being measured on the Illumina 450k and\ 850k (EPIC) microarray tracks, not the probe locations themselves. Contact us\ or Illumina if you need the probe locations directly. More information about\ the arrays can be found on the\ Infinium MethylationEPIC Kit website.\
\ Note: The 450k track on hg38 contains 128,989 regions representing the target regions, not the probes\ themselves.
\ \\ The Infinium CytoSNP-850K v1.2 BeadChip provides comprehensive coverage of\ cytogenetically relevant genes on a proven platform, helping researchers find valuable information\ that may be missed by other technologies. It contains approximately 850,000 empirically selected\ single nucleotide polymorphisms (SNPs) spanning the entire genome with enriched coverage for 3,262\ genes of known cytogenetics relevance in both constitutional and cancer applications. \
\ \\ The CytoScan HD Array, which is included in the\ CytoScan HD Suite, provides the broadest coverage and highest performance for\ detecting chromosomal aberrations. CytoScan HD Suite has greater than 99% sensitivity and can\ reliably detect 25-50kb copy number changes across the genome at high specificity with\ single-nucleotide polymorphism (SNP) allelic corroboration. With more than 2.6 million copy number\ markers, CytoScan HD Suite covers all OMIM and RefSeq genes.\
\ \\ Bionano Laboratories provides access to Optical Genome Mapping (OGM) data for projects across a variety of\ applications for researchers, clinicians, and pharmaceutical companies.
\This track shows the CTTAAG sites used by the \ Bionano Optical Genome Mapping system,\ an assay to detect structural variants.\
\ \\ Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \ \\ The Agilent arrays were downloaded from their \ Agilent SureDesign website tool on March 2022.
\\ The Illumina 450k and 850k (EPIC) tracks were created using a few columns from the\ Infinium MethylationEPIC v1.0 B5 Manifest File (CSV Format)\ and was then converted into a bigBed.
\\ The Illumina CytoSNP-850K track was created by downloading the\ CytoSNP-850K v1.2 Manifest File (CSV Format) (GRCh38) file and then converted\ into a bigBed file.\
\\ The Affymetrix Cytoscan HD GeneChip Array track was created by converting the \ CytoScanHD_Accel_Array.na36.bed.zip\ into a bigBed file.\
\\ The Bionano track was created by receiving the BED files from\ \ apang@bionano.\ com\ \ and converted to bigBed files using the bedToBigBed tool.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API \ or downloaded from our \ Downloads site. Please refer to our\ \ mailing list archives for questions, or our\ \ Data Access FAQ for more information.\
\ \\ Thanks to the Agilent and Illumina support teams for sharing the data and the UCSC Genome Browser\ engineers for configuring the data.
\\ Thanks to Andy Pang from Bionano Genomics for providing the BED data file.
\ varRep 1 bigDataUrl /gbdb/hg38/bbi/illumina/epic850K.bb\ colorByStrand 255,0,0 0,0,255\ html genotypeArrays\ longLabel Illumina 850k EPIC Methylation Array\ noScoreFilter on\ parent genotypeArrays on\ priority 5\ shortLabel Illumina 850k\ track snpArrayIllumina850k\ type bigBed 6\ urls refGeneAccession="https://www.ncbi.nlm.nih.gov/nuccore/$$" rsID="https://www.ncbi.nlm.nih.gov/snp/?term=$$"\ visibility pack\ wgEncodeRegMarkH3k27acK562 K562 bigWig 0 6249 H3K27Ac Mark (Often Found Near Regulatory Elements) on K562 Cells from ENCODE 2 5 128 128 255 191 191 255 0 0 0 regulation 1 color 128,128,255\ longLabel H3K27Ac Mark (Often Found Near Regulatory Elements) on K562 Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k27ac\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel K562\ table wgEncodeBroadHistoneK562H3k27acStdSig\ track wgEncodeRegMarkH3k27acK562\ type bigWig 0 6249\ wgEncodeRegMarkH3k4me1K562 K562 bigWig 0 5716 H3K4Me1 Mark (Often Found Near Regulatory Elements) on K562 Cells from ENCODE 0 5 128 128 255 191 191 255 0 0 0 regulation 1 color 128,128,255\ longLabel H3K4Me1 Mark (Often Found Near Regulatory Elements) on K562 Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k4me1\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel K562\ table wgEncodeBroadHistoneK562H3k4me1StdSig\ track wgEncodeRegMarkH3k4me1K562\ type bigWig 0 5716\ wgEncodeRegMarkH3k4me3K562 K562 bigWig 0 9918 H3K4Me3 Mark (Often Found Near Promoters) on K562 Cells from ENCODE 0 5 128 128 255 191 191 255 0 0 0 regulation 1 color 128,128,255\ longLabel H3K4Me3 Mark (Often Found Near Promoters) on K562 Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k4me3\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel K562\ table wgEncodeBroadHistoneK562H3k4me3StdSig\ track wgEncodeRegMarkH3k4me3K562\ type bigWig 0 9918\ KAPA_HyperExome_hg38_capture_targets KAPA Hyper P bigBed Roche - KAPA HyperExome Capture Probe Footprint 0 5 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/KAPA_HyperExome_hg38_capture_targets.bb\ color 100,143,255\ longLabel Roche - KAPA HyperExome Capture Probe Footprint\ parent exomeProbesets off\ shortLabel KAPA Hyper P\ track KAPA_HyperExome_hg38_capture_targets\ type bigBed\ wgEncodeReg4AtacLiver Liver bigWig Avg. ATAC level of 7 liver experiments (tissues and primary cells only) 0 5 137 152 82 196 203 168 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLiverATAC.bw\ color 137,152,82\ longLabel Avg. ATAC level of 7 liver experiments (tissues and primary cells only)\ parent wgEncodeReg4Atac\ priority 5\ shortLabel Liver\ track wgEncodeReg4AtacLiver\ type bigWig\ dbSnp153BadCoords Map Err dbSnp(153) bigBed 4 Mappings with Inconsistent Coordinates from dbSNP 153 1 5 100 100 100 177 177 177 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$ varRep 1 bigDataUrl /gbdb/hg38/snp/dbSnp153BadCoords.bb\ color 100,100,100\ longLabel Mappings with Inconsistent Coordinates from dbSNP 153\ parent dbSnp153ViewErrs off\ priority 5\ shortLabel Map Err dbSnp(153)\ subGroups view=errs\ track dbSnp153BadCoords\ type bigBed 4\ dbSnp155BadCoords Map Err dbSnp(155) bigBed 4 Mappings with Inconsistent Coordinates from dbSNP 155 1 5 100 100 100 177 177 177 0 0 0 https://www.ncbi.nlm.nih.gov/snp/$$ varRep 1 bigDataUrl /gbdb/hg38/snp/dbSnp155BadCoords.bb\ color 100,100,100\ longLabel Mappings with Inconsistent Coordinates from dbSNP 155\ parent dbSnp155ViewErrs off\ priority 5\ shortLabel Map Err dbSnp(155)\ subGroups view=errs\ track dbSnp155BadCoords\ type bigBed 4\ multiz470way Multiz 470-way bigMaf Multiz Alignments of 470 mammals 3 5 0 10 100 0 90 10 0 0 0 compGeno 1 altColor 0,90,10\ bigDataUrl https://hgdownload.soe.ucsc.edu/goldenPath/hg38/multiz470way/multiz470way.bigMaf\ color 0, 10, 100\ frames https://hgdownload.soe.ucsc.edu/goldenPath/hg38/multiz470way/multiz470wayFrames.bb\ group compGeno\ irows on\ itemFirstCharCase noChange\ longLabel Multiz Alignments of 470 mammals\ noInherit on\ parent cons470wayViewalign on\ priority 5\ sGroup_Afrotheria triMan1 HLloxAfr4 HLeleMax1 HLdugDug1 oryAfe1 HLproCap3 HLhetBru1 chrAsi1 echTel2 HLhydGig1 eleEdw1 HLmicTal1\ sGroup_Artiodactyla HLbalMus1 HLeubGla1 HLbalEde1 balAcu1 HLeubJap1 HLmegNov1 HLmonMon1 HLphoSin1 HLlagObl1 HLgloMel1 HLpepEle1 HLbalMys1 HLneoAsi1 HLplaMin1 HLbalPhy1 HLmesBid1 HLkogBre1 HLlniGeo1 HLlamGuaCac1 HLvicVicMen1 HLvicPacHua3 HLlamGlaCha1 HLponBla1 HLzipCav1 HLlamGla1 HLcerHanYar1 HLmunMun1 HLbosGau1 HLaxiPor1 HLranTarGra2 HLranTar1 HLoviOri1 HLgirCam1 HLsynCaf1 HLoryDam1 HLcapSib1 HLmunRee1 HLcatWag1 HLhipEqu1 HLhipNig1 HLcapPyg1 HLhydIne1 HLodoVir1 HLprzAlb1 HLgirCam2 HLhemHyl1 HLmosMos1 HLbeaHun1 HLalcAlc1 HLoviNivLyd1 HLconTau2 HLdamLun1 HLodoHem1 HLoryGaz1 HLkobLecLec1 HLantAme1 HLbosFro1 HLeudTho1 HLkobEll1 HLlitWal1 HLoreAme1 HLbosGru1 bisBis1 HLmosBer1 HLcepHar1 HLaepMel1 HLantMar1 HLoviCan1 HLoreOre1 HLmunCri1 HLproPrz1 HLmadKir1 HLsylGri1 HLtraJav1 HLredRed1 HLsaiTat1 HLtraScr1 HLcerEla1 HLneoMos1 HLnanGra1 HLtraImb1 HLneoPyg1 HLphiMax1 HLrapCam1 HLtraKan1 HLmosChr1 HLcapIbe1\ sGroup_Carnivore HLphoVit1 HLzalCal1 HLodoRos1 HLursArc1 HLeumJub1 odoRosDiv1 neoSch1 HLcalUrs1 HLmirLeo1 HLursThi1 HLeriBar1 HLhalGry1 HLmirAng2 ursMar1 HLursAme2 HLailMel2 felCat9 HLneoNeb1 HLaciJub2 HLpanPar1 HLpanLeo1 HLursAme1 HLlynCan1 HLpumYag1 lepWed1 panTig1 HLpanOnc1 HLpanOnc2 HLarcGaz2 HLcryFer2 canFam4 HLlynPar1 HLpriBen1 enhLutKen1 canFam5 HLailFul2 HLcanLupDin1 HLlonCan1 HLlutLut1 HLvulVul1 HLpumCon1 HLmusErm1 HLpotFla1 HLpteBra2 HLlycPic3 HLhyaHya1 enhLutNer1 HLpteBra1 HLfelNig1 HLmelCap1 HLmusFur2 HLmarZib1 HLvulLag1 HLmusPut1 HLcroCro1 HLparHer1 HLgulGul1 HLneoVis1 HLsurSur1 HLmunMug1 HLbasSum1 HLspiGra1 HLhelPar1 HLsurSur2 HLtaxTax1 HLnasNar1 HLproLot1 HLlycPic2\ sGroup_Cetartiodactyla HLescRob1 HLphyCat2 HLhipAmp3 HLdelLeu2 HLturTru4 phyCat1 HLturAdu2 orcOrc1 HLsouChi1 lipVex1 HLturAdu1 HLbalBon1 HLphoPho1 HLphoPho2 HLturTru3 turTru2 HLcamDro2 HLhipAmp1 HLcamFer3 HLcamBac1 vicPac2 susScr11 bosTau9 HLbubBub2 HLodoVir3 HLelaDav1 HLbosInd2 HLoviAri5 HLcapHir2 HLodoVir2 HLbosMut2 HLcapAeg1 HLammLer1 panHod1 HLgirTip1 HLoviCan2 HLokaJoh2 HLtraStr1 HLoviAmm1\ sGroup_Chiroptera HLrhiFer5 HLrhiSin1 HLptePse1 HLpteGig1 pteAle1 HLpteRuf1 HLhipGal1 HLrouAeg4 HLpteVam2 HLrouLes1 HLhipArm1 HLeonSpe1 HLeidDup1 HLmacSob1 HLeidHel2 HLtadBra1 HLrouMad1 HLdesRot2 HLmolMol2 HLphyDis3 HLlepYer1 HLmorBla1 HLmegLyr2 HLtonSau1 HLanoCau1 HLcarPer3 HLartJam2 HLminNat1 ptePar1 HLartJam1 HLminSch1 HLmacCal1 HLstuHon1 HLcynBra1 HLmyoMyo6 HLmicHir1 HLcraTho1 HLmyoSep1 myoBra1 eptFus1 HLmyoLuc1 myoLuc2 HLnocLep1 myoDav1 HLmurAurFea1 HLnycHum2 HLantPal1 HLaeoCin1 HLlasBor1 HLpipKuh2 HLpipPip2 HLpipPip1\ sGroup_Euarchontoglires HLgalVar2 tupChi1 tupBel1\ sGroup_Glires HLsciCar1 HLsciVul1 HLmarMon2 HLmarFla1 HLxerIna1 HLmarMar1 HLcynGun1 HLmarMon1 HLmarHim1 HLspeDau1 HLmarVan1 HLuroPar1 speTri2 HLereDor1 HLaplRuf1 HLpedCap1 HLgliGli1 hetGla2 HLhysCri1 HLcoePre1 chiLan1 HLdasPun1 HLcasCan3 HLoryCunCun4 HLgraMur1 HLdinBra1 HLfukDam2 oryCun2 HLlepAme1 HLhydHyd1 HLsylBac1 HLcavTsc1 HLdolPat1 cavPor3 HLmusAve1 HLoryCun3 HLlepTim1 octDeg1 HLcteGun1 HLcteSoc1 HLpetTyp1 nanGal1 HLthrSwi1 HLmyoCoy1 HLperCal2 HLperCri1 HLperManBai2 HLperPol1 HLperNas1 HLperLeu1 HLperEre1 HLrhiPru1 HLonyTor1 HLallBul1 jacJac1 HLcriGam1 HLcriGri3 HLondZib1 ochPri3 HLgraSur1 HLarvAmp1 HLarvNil1 HLellTal1 mm10 mm39 HLdipSte1 dipOrd2 HLacoRus1 HLmasCou1 HLzapHud1 HLmicAgr2 HLpsaObe1 HLmusSpr1 HLmyoGla2 HLmusCar1 micOch1 HLratNor7 HLellLut1 HLmusPah1 HLrhoOpi1 HLacoCah1 rn6 HLmusSpi1 HLmicFor1 HLmicArv1 HLmicOec1 HLsigHis1 HLratRat7 HLneoLep1 mesAur1 HLmerUng1 HLperLonPac1 cavApe1 HLapoSyl1\ sGroup_Laurasiatheria HLdicBic1 cerSim1 HLrhiUni1 HLdicSum1 HLtapInd1 HLtapTer1 HLtapInd2 equCab3 HLcerSimCot1 HLequAsi1 HLequQuaBoe1 HLequAsiAsi2 equPrz1 HLmanPen2 HLphaTri2 HLmanJav1 HLmanJav2 HLmanTri1 manPen1 HLsolPar1 HLtalOcc1 HLscaAqu1 HLuroGra1 conCri1 sorAra2 eriEur2\ sGroup_Metatheria HLvomUrs1 HLphaCin1 HLtriVul1 HLgymLea1 HLmacGig1 HLphaGym1 HLnotEug3 HLantFla1 HLmacFul1 HLsarHar2 HLpseCup1 monDom5 HLgraAgi1 HLospRuf1 HLdidVir1 HLpseCor1 HLpseOcc1 HLthyCyn1 macEug2\ sGroup_Monotremata HLornAna3 HLtacAcu1\ sGroup_Primates panTro6 panPan3 gorGor6 ponAbe3 HLnomLeu4 HLhylMol2 rheMac10 HLmacFas6 HLtheGel1 HLmacFus1 HLrhiRox2 chlSab2 HLpapAnu5 cerAty1 HLmanSph1 macNem1 HLtraFra1 HLpygNem1 HLpilTep2 HLeryPat1 HLallNig1 rhiBie1 HLcerMon1 manLeu1 HLsemEnt1 colAng1 HLcerNeg1 HLpitPit1 HLateGeo1 HLsapApe1 HLaloPal1 HLpleDon1 cebCap1 HLcalJac4 aotNan1 HLsagImp1 HLcalPym1 HLcebAlb1 nasLar1 HLsaiBol1 saiBol1 HLdauMad1 tarSyr2 HLindInd1 HLmirZaz1 eulMac1 micMur3 HLproSim1 HLeulFla1 HLmirCoq1 HLlemCat1 HLcheMed1 HLmicSpe31 HLeulFul1 eulFla1 HLmicTav1 HLeulMon1 proCoq1 HLnycCou1 otoGar3\ sGroup_Xenarthra HLchoDid2 HLchoHof3 HLchoDid1 HLtamTet1 HLmyrTri1 dasNov3 HLtolMat1\ shortLabel Multiz 470-way\ speciesCodonDefault hg38\ speciesDefaultOff panPan3 gorGor6 ponAbe3 HLnomLeu4 HLhylMol2 macNem1 HLtheGel1 HLmacFas6 HLcerMon1 HLpilTep2 colAng1 manLeu1 cerAty1 HLpapAnu5 HLmanSph1 HLsemEnt1 HLmacFus1 HLtraFra1 rhiBie1 HLrhiRox2 HLpygNem1 HLcerNeg1 nasLar1 HLallNig1 chlSab2 HLeryPat1 HLpitPit1 HLateGeo1 aotNan1 HLpleDon1 HLaloPal1 HLsaiBol1 HLsagImp1 saiBol1 HLcalJac4 HLcalPym1 HLsapApe1 cebCap1 HLcebAlb1 HLdauMad1 proCoq1 HLgalVar2 HLindInd1 HLeulFul1 eulFla1 HLlemCat1 HLproSim1 HLeulMon1 HLeulFla1 HLcheMed1 eulMac1 tarSyr2 micMur3 HLmirZaz1 HLmirCoq1 HLmicSpe31 HLmicTav1 HLnycCou1 otoGar3 HLtapTer1 HLrhiUni1 HLtapInd1 HLtapInd2 HLdicBic1 HLdicSum1 HLcerSimCot1 cerSim1 HLequQuaBoe1 equPrz1 HLequAsi1 HLequAsiAsi2 HLeubGla1 HLsciCar1 HLeubJap1 HLsciVul1 balAcu1 HLbalBon1 HLxerIna1 HLmegNov1 HLbalPhy1 HLbalMys1 tupChi1 HLbalEde1 HLaplRuf1 HLescRob1 HLmarFla1 phyCat1 HLphyCat2 HLmarMar1 HLmesBid1 HLchoHof3 HLchoDid2 HLmarVan1 HLplaMin1 HLmarHim1 HLspeDau1 HLmarMon1 HLmarMon2 HLzipCav1 HLuroPar1 lipVex1 HLdelLeu2 HLlniGeo1 HLcynGun1 HLmonMon1 HLhipAmp3 HLhipAmp1 speTri2 HLneoAsi1 HLpumCon1 panTig1 HLkogBre1 HLneoNeb1 HLpanPar1 HLphoSin1 HLeriBar1 HLmanTri1 HLgliGli1 HLdugDug1 HLphaTri2 HLpanOnc1 HLphoVit1 HLaciJub2 HLhalGry1 HLchoDid1 HLpedCap1 HLphoPho2 HLphoPho1 neoSch1 lepWed1 HLpanOnc2 HLponBla1 HLpriBen1 HLcamFer3 orcOrc1 HLmanPen2 HLcasCan3 HLursThi1 HLrhiSin1 HLlynPar1 HLmirLeo1 HLlynCan1 HLmirAng2 HLpanLeo1 HLcamBac1 HLsouChi1 manPen1 HLodoRos1 HLcalUrs1 odoRosDiv1 HLvicPacHua3 HLhipArm1 HLlamGla1 HLailMel2 HLzalCal1 HLeumJub1 felCat9 HLpumYag1 pteAle1 HLursArc1 ursMar1 HLpepEle1 HLgloMel1 HLrhiFer5 HLarcGaz2 HLcamDro2 HLlagObl1 HLptePse1 HLmanJav1 HLmanJav2 vicPac2 HLvicVicMen1 HLlamGuaCac1 HLlamGlaCha1 HLturTru4 HLturAdu1 HLturAdu2 HLursAme1 HLursAme2 triMan1 HLtadBra1 HLpteVam2 turTru2 HLpteRuf1 HLpteGig1 HLturTru3 HLfelNig1 HLeidDup1 HLeidHel2 HLeleMax1 HLhipGal1 HLloxAfr4 tupBel1 HLcynBra1 HLeonSpe1 HLcryFer2 HLgraMur1 HLmyrTri1 oryAfe1 HLrouLes1 HLrouAeg4 HLrouMad1 HLvulVul1 HLvulLag1 HLlycPic3 HLtamTet1 HLcanLupDin1 canFam5 HLpotFla1 HLhydGig1 HLmegLyr2 HLlycPic2 HLailFul2 HLmolMol2 HLcroCro1 HLmacSob1 HLhyaHya1 HLlepTim1 susScr11 HLminSch1 HLparHer1 HLminNat1 HLlepAme1 HLcraTho1 HLcatWag1 HLmorBla1 HLoryCunCun4 ptePar1 oryCun2 HLoryCun3 HLsylBac1 HLnasNar1 HLhysCri1 HLereDor1 HLcoePre1 HLmarZib1 HLgulGul1 HLproLot1 HLbasSum1 HLspiGra1 HLtaxTax1 HLmelCap1 HLmusAve1 HLsurSur2 HLsurSur1 HLmunMug1 HLhelPar1 HLscaAqu1 HLlonCan1 HLpteBra2 HLpteBra1 enhLutNer1 enhLutKen1 HLlutLut1 HLmyoMyo6 hetGla2 myoBra1 HLdesRot2 HLtalOcc1 HLgirCam1 HLokaJoh2 HLmacCal1 HLmusErm1 HLmyoSep1 HLneoVis1 HLgirCam2 HLgirTip1 HLmyoLuc1 myoLuc2 HLmusPut1 HLmusFur2 HLlepYer1 myoDav1 HLsynCaf1 HLbubBub2 HLmosBer1 HLmosMos1 HLmosChr1 HLcerHanYar1 HLmicHir1 HLbosInd2 chrAsi1 HLbosGau1 HLanoCau1 HLbosFro1 HLbosMut2 HLprzAlb1 HLmurAurFea1 HLnocLep1 HLhipEqu1 HLcepHar1 HLhipNig1 HLbosGru1 HLoryDam1 HLsylGri1 HLphiMax1 HLoryGaz1 HLtraStr1 HLantAme1 HLmunRee1 HLmunCri1 HLcerEla1 HLtraImb1 HLconTau2 HLtraScr1 HLkobEll1 HLmunMun1 HLammLer1 HLdamLun1 HLoviCan1 HLkobLecLec1 HLcapPyg1 HLcapHir2 HLcapAeg1 panHod1 HLalcAlc1 HLbeaHun1 HLaepMel1 HLodoHem1 HLredRed1 HLfukDam2 HLcapSib1 HLodoVir3 HLranTarGra2 HLranTar1 HLoreOre1 HLhydIne1 HLoviCan2 HLoviNivLyd1 HLneoMos1 HLodoVir2 HLodoVir1 HLhemHyl1 HLoviOri1 HLoviAri5 HLneoPyg1 nanGal1 HLnanGra1 HLproPrz1 HLrapCam1 HLeudTho1 HLantMar1 chiLan1 HLdasPun1 HLcteGun1 HLlitWal1 HLmadKir1 HLcarPer3 HLaxiPor1 HLphyDis3 HLtonSau1 HLtraJav1 HLartJam1 HLartJam2 HLhetBru1 HLuroGra1 HLtraKan1 conCri1 HLstuHon1 HLoreAme1 HLallBul1 HLelaDav1 HLsaiTat1 HLaeoCin1 HLdipSte1 dipOrd2 HLtolMat1 HLantPal1 HLrhiPru1 HLnycHum2 HLoviAmm1 HLcapIbe1 HLdinBra1 jacJac1 HLzapHud1 HLdolPat1 HLlasBor1 HLpipKuh2 HLperLonPac1 HLhydHyd1 HLpipPip1 HLpipPip2 ochPri3 eleEdw1 cavApe1 HLpetTyp1 HLcavTsc1 cavPor3 HLthrSwi1 octDeg1 HLcriGam1 HLneoLep1 HLcteSoc1 HLmyoCoy1 echTel2 eriEur2 HLperNas1 HLcriGri3 HLperCri1 HLondZib1 HLperCal2 HLperEre1 HLonyTor1 mesAur1 HLperLeu1 HLellTal1 HLperPol1 HLperManBai2 HLsigHis1 HLellLut1 HLmyoGla2 HLarvAmp1 HLpsaObe1 HLacoRus1 HLgraSur1 HLarvNil1 HLmicOec1 HLmicTal1 HLmicAgr2 HLmicFor1 HLacoCah1 HLmicArv1 micOch1 HLrhoOpi1 HLmasCou1 HLmerUng1 HLratRat7 HLratNor7 rn6 HLmusPah1 HLmusCar1 HLmusSpi1 mm10 HLmusSpr1 HLapoSyl1 sorAra2 HLvomUrs1 HLphaCin1 HLgraAgi1 HLtriVul1 HLdidVir1 HLphaGym1 monDom5 HLgymLea1 HLthyCyn1 HLpseCup1 HLmacGig1 HLpseCor1 HLmacFul1 HLnotEug3 HLospRuf1 HLpseOcc1 macEug2 HLantFla1 HLornAna3 HLtacAcu1\ speciesDefaultOn panTro6 rheMac10 canFam4 equCab3 HLsolPar1 bosTau9 HLbalMus1 bisBis1 dasNov3 eptFus1 mm39 HLproCap3 HLsarHar2 HLtacAcu1\ speciesGroups Primates Euarchontoglires Carnivore Laurasiatheria Cetartiodactyla Artiodactyla Xenarthra Chiroptera Glires Afrotheria Metatheria Monotremata\ speciesLabels HLnomLeu4="northern white-cheeked gibbon" HLhylMol2="silvery gibbon" HLtheGel1=gelada HLmacFas6="crab-eating macaque" HLcerMon1="Mona monkey" HLpilTep2="Ugandan red Colobus" HLpapAnu5="olive baboon" HLmanSph1=mandrill HLsemEnt1="Hanuman langur" HLmacFus1="Japanese macaque" HLtraFra1="Francois's langur" HLrhiRox2="golden snub-nosed monkey" HLpygNem1="Red shanked douc langur" HLcerNeg1="De Brazza's monkey" HLallNig1="Allen's swamp monkey" HLeryPat1="red guenon" HLpitPit1="white-faced saki" HLateGeo1="black-handed spider monkey" HLpleDon1="Bolivian titi" HLaloPal1="mantled howler monkey" HLsaiBol1="Bolivian squirrel monkey" HLsagImp1=tamarin HLcalJac4="white-tufted-ear marmoset" HLcalPym1="pygmy marmoset" HLsapApe1="tufted capuchin" HLcebAlb1="white-fronted capuchin" HLdauMad1=aye-aye HLgalVar2="Sunda flying lemur" HLindInd1=babakoto HLeulFul1="brown lemur" HLlemCat1="Ring-tailed lemur" HLproSim1="greater bamboo lemur" HLeulMon1="mongoose lemur" HLeulFla1="Sclater's lemur" HLcheMed1="Lesser dwarf lemur" HLmirZaz1="Northern giant mouse lemur" HLmirCoq1="Coquerel's mouse lemur" HLmicSpe31="mouse lemur" HLmicTav1="Northern rufous mouse lemur" HLnycCou1="slow loris" HLtapTer1="Brazilian tapir" HLrhiUni1="greater Indian rhinoceros" HLtapInd1="Asiatic tapir" HLtapInd2="Asiatic tapir" HLdicBic1="black rhinoceros" HLdicSum1="Sumatran rhinoceros" HLcerSimCot1="northern white rhinoceros" HLequQuaBoe1="Equus burchelli boehmi" HLequAsi1=ass HLequAsiAsi2=donkey HLeubGla1="North Atlantic right whale" HLsciCar1="gray squirrel" HLeubJap1="North Pacific right whale" HLsciVul1="Eurasian red squirrel" HLbalBon1="Antarctic minke whale" HLxerIna1="South African ground squirrel" HLmegNov1="humpback whale" HLbalPhy1="Fin whale" HLbalMys1="bowhead whale" HLbalMus1="Blue whale" HLbalEde1="pygmy Bryde's whale" HLaplRuf1="mountain beaver" HLescRob1="grey whale" HLmarFla1="yellow-bellied marmot" HLphyCat2="sperm whale" HLmarMar1="Alpine marmot" HLmesBid1="Sowerby's beaked whale" HLchoHof3="Hoffmann's two-fingered sloth" HLchoDid2="southern two-toed sloth" HLmarVan1="Vancouver Island marmot" HLplaMin1="Indus River dolphin" HLmarHim1="Himalayan marmot" HLspeDau1="Daurian ground squirrel" HLmarMon1=woodchuck HLmarMon2=woodchuck HLzipCav1="Cuvier's beaked whale" HLuroPar1="Arctic ground squirrel" HLdelLeu2="beluga whale" HLlniGeo1=boutu HLcynGun1="Gunnison's prairie dog" HLmonMon1=narwhal HLhipAmp3=hippopotamus HLhipAmp1=hippopotamus HLneoAsi1="Yangtze finless porpoise" HLpumCon1=puma HLkogBre1="pygmy sperm whale" HLneoNeb1="Clouded leopard" HLpanPar1=leopard HLphoSin1=vaquita HLeriBar1="bearded seal" HLmanTri1="Tree pangolin" HLgliGli1="Fat dormouse" HLdugDug1=dugong HLphaTri2="Tree pangolin" HLpanOnc1=jaguar HLphoVit1="harbor seal" HLaciJub2=cheetah HLhalGry1="gray seal" HLchoDid1="southern two-toed sloth" HLpedCap1=springhare HLphoPho2="harbor porpoise" HLphoPho1="harbor porpoise" HLpanOnc2=jaguar HLponBla1=franciscana HLpriBen1="Amur leopard cat" HLcamFer3="Wild Bactrian camel" HLmanPen2="Chinese pangolin" HLcasCan3="American beaver" HLursThi1="Asian black bear" HLrhiSin1="Chinese rufous horseshoe bat" HLlynPar1="Spanish lynx" HLmirLeo1="Southern elephant seal" HLlynCan1="Canada lynx" HLmirAng2="Northern elephant seal" HLpanLeo1=lion HLcamBac1="Bactrian camel" HLsouChi1="Indo-pacific humpbacked dolphin" HLodoRos1=walrus HLcalUrs1="northern fur seal" HLvicPacHua3="Lama pacos huacaya" HLhipArm1="great roundleaf bat" HLlamGla1=llama HLailMel2="giant panda" HLzalCal1="California sea lion" HLeumJub1="Steller sea lion" HLpumYag1=jaguarundi HLursArc1="grizzly bear" HLpepEle1="melon-headed whale" HLgloMel1="long-finned pilot whale" HLrhiFer5="greater horseshoe bat" HLarcGaz2="antarctic fur seal" HLcamDro2="Arabian camel" HLlagObl1="Pacific white-sided dolphin" HLptePse1="Bonin flying fox" HLmanJav1="Malayan pangolin" HLmanJav2="Malayan pangolin" HLvicVicMen1="Vicugna mensalis" HLlamGuaCac1=guanaco HLlamGlaCha1=llama HLturTru4="common bottlenose dolphin" HLturAdu1="Indo-pacific bottlenose dolphin" HLturAdu2="Indo-pacific bottlenose dolphin" HLursAme1="American black bear" HLursAme2="American black bear" HLtadBra1="Brazilian free-tailed bat" HLpteVam2="large flying fox" HLpteRuf1="Malagasy flying fox" HLpteGig1="Indian flying fox" HLturTru3="common bottlenose dolphin" HLfelNig1="black-footed cat" HLeidDup1="Malagasy straw-colored fruit bat" HLeidHel2="straw-colored fruit bat" HLeleMax1="Asiatic elephant" HLhipGal1="Cantor's roundleaf bat" HLloxAfr4="African savanna elephant" HLcynBra1="lesser short-nosed fruit bat" HLeonSpe1="lesser dawn bat" HLcryFer2=fossa HLgraMur1="woodland dormouse" HLmyrTri1="giant anteater" HLrouLes1="Leschenault's rousette" HLrouAeg4="Egyptian rousette" HLrouMad1="Madagascan rousette" HLvulVul1="red fox" HLvulLag1="Arctic fox" HLlycPic3="African hunting dog" HLtamTet1="southern tamandua" HLcanLupDin1=dingo HLpotFla1=kinkajou HLhydGig1="Steller's sea cow" HLmegLyr2="Indian false vampire" HLlycPic2="African hunting dog" HLailFul2="lesser panda" HLmolMol2="Pallas's mastiff bat" HLcroCro1="spotted hyena" HLmacSob1="long-tongued fruit bat" HLhyaHya1="striped hyena" HLlepTim1="Mountain hare" HLminSch1="Schreibers' long-fingered bat" HLparHer1="Asian palm civet" HLminNat1="Miniopterus schreibersii natalensis" HLlepAme1="snowshoe hare" HLcraTho1="hog-nosed bat" HLcatWag1="Chacoan peccary" HLmorBla1="Antillean ghost-faced bat" HLoryCunCun4="European rabbit" HLoryCun3=rabbit HLsylBac1="brush rabbit" HLnasNar1="White-nosed coati" HLhysCri1="crested porcupine" HLereDor1="North American porcupine" HLcoePre1="Brazilian porcupine" HLmarZib1=sable HLsolPar1="Hispaniolan solenodon" HLgulGul1=wolverine HLproLot1=raccoon HLbasSum1=Cacomistle HLspiGra1="western spotted skunk" HLtaxTax1="North American badger" HLmelCap1=ratel HLmusAve1="hazel dormouse" HLsurSur2=meerkat HLsurSur1=meerkat HLmunMug1="banded mongoose" HLhelPar1="dwarf mongoose" HLscaAqu1="eastern mole" HLlonCan1="Northern American river otter" HLpteBra2="giant otter" HLpteBra1="giant otter" HLlutLut1="Eurasian river otter" HLmyoMyo6="greater mouse-eared bat" HLdesRot2="common vampire bat" HLtalOcc1="Iberian mole" HLgirCam1=giraffe HLokaJoh2=okapi HLmacCal1="California big-eared bat" HLmusErm1=ermine HLmyoSep1="Northern long-eared myotis" HLneoVis1="American mink" HLgirCam2=giraffe HLgirTip1="Masai giraffe" HLmyoLuc1="little brown bat" HLmusPut1="European polecat" HLmusFur2="domestic ferret" HLlepYer1="Lesser long-nosed bat" HLsynCaf1="African buffalo" HLbubBub2="water buffalo" HLmosBer1="Chinese forest musk deer" HLmosMos1="Siberian musk deer" HLmosChr1="alpine musk deer" HLcerHanYar1="Yarkand deer" HLmicHir1="Schizostoma hirsutum" HLbosInd2="zebu cattle" HLbosGau1=gaur HLanoCau1="tailed tailless bat" HLbosFro1=gayal HLbosMut2="wild yak" HLprzAlb1="white-lipped deer" HLmurAurFea1="Murina feae" HLnocLep1="greater bulldog bat" HLhipEqu1="roan antelope" HLcepHar1="Harvey's duiker" HLhipNig1="sable antelope" HLbosGru1="domestic yak" HLoryDam1="scimitar-horned oryx" HLsylGri1="bush duiker" HLphiMax1="Maxwell's duiker" HLoryGaz1=gemsbok HLtraStr1="greater kudu" HLantAme1=pronghorn HLmunRee1="Reeves' muntjac" HLmunCri1="black muntjac" HLcerEla1="Central European red deer" HLtraImb1="lesser kudu" HLconTau2="brindled gnu" HLtraScr1=bushbuck HLkobEll1=waterbuck HLmunMun1=muntjak HLammLer1=aoudad HLdamLun1=topi HLoviCan1="bighorn sheep" HLkobLecLec1=lechwe HLcapPyg1="Eastern roe deer" HLcapHir2=goat HLcapAeg1="wild goat" HLalcAlc1="Eurasian elk" HLbeaHun1="Cobus hunteri" HLaepMel1=impala HLodoHem1="mule deer" HLredRed1="Bohar reedbuck" HLfukDam2="Damara mole-rat" HLcapSib1="Siberian ibex" HLodoVir3="white-tailed deer" HLranTarGra2="porcupine caribou" HLranTar1=reindeer HLoreOre1=klipspringer HLhydIne1="Chinese water deer" HLoviCan2="bighorn sheep" HLoviNivLyd1="snow sheep" HLneoMos1=suni HLodoVir2="white-tailed deer" HLodoVir1="white-tailed deer" HLhemHyl1="Nilgiri tahr" HLoviOri1="Asiatic mouflon" HLoviAri5=sheep HLneoPyg1="royal antelope" HLnanGra1="Grant's gazelle" HLproPrz1="Przewalski's gazelle" HLrapCam1=steenbok HLeudTho1="Thomson's gazelle" HLantMar1=springbok HLdasPun1="punctate agouti" HLcteGun1="northern gundi" HLlitWal1=gerenuk HLmadKir1="Kirk's dik-dik" HLcarPer3="Seba's short-tailed bat" HLaxiPor1="Hog deer" HLphyDis3="pale spear-nosed bat" HLtonSau1="stripe-headed round-eared bat" HLtraJav1="Java mouse-deer" HLartJam1="Jamaican fruit-eating bat" HLartJam2="Jamaican fruit-eating bat" HLhetBru1="yellow-spotted hyrax" HLproCap3="Cape rock hyrax" HLuroGra1="gracile shrew mole" HLtraKan1="lesser mouse-deer" HLstuHon1="Honduran yellow-shouldered bat" HLoreAme1="mountain goat" HLallBul1="Gobi jerboa" HLelaDav1="Pere David's deer" HLsaiTat1="saiga antelope" HLaeoCin1="hoary bat" HLdipSte1="Stephens's kangaroo rat" HLtolMat1="Southern three-banded armadillo" HLantPal1="pallid bat" HLrhiPru1="hoary bamboo rat" HLnycHum2="evening bat" HLoviAmm1=argali HLcapIbe1="Alpine ibex" HLdinBra1=pacarana HLzapHud1="meadow jumping mouse" HLdolPat1="Patagonian cavy" HLlasBor1="red bat" HLpipKuh2="Kuhl's pipistrelle" HLperLonPac1="Pacific pocket mouse" HLhydHyd1=capybara HLpipPip1="common pipistrelle" HLpipPip2="common pipistrelle" HLpetTyp1=dassie-rat HLcavTsc1="Montane guinea pig" HLthrSwi1="Greater cane rat" HLcriGam1="Gambian giant pouched rat" HLneoLep1="desert woodrat" HLcteSoc1="social tuco-tuco" HLmyoCoy1=nutria HLperNas1="northern rock mouse" HLcriGri3="Chinese hamster" HLperCri1="Hesperomys crinitus" HLondZib1=muskrat HLperCal2="Peromyscus californicus subsp. insignis" HLperEre1="cactus mouse" HLonyTor1="southern grasshopper mouse" HLperLeu1="white-footed mouse" HLellTal1="Northern mole vole" HLperPol1="oldfield mouse" HLperManBai2="prairie deer mouse" HLsigHis1="hispid cotton rat" HLellLut1="Transcaucasian mole vole" HLmyoGla2="Bank vole" HLarvAmp1="Eurasian water vole" HLpsaObe1="fat sand rat" HLacoRus1="golden spiny mouse" HLgraSur1="African woodland thicket rat" HLarvNil1="African grass rat" HLmicOec1="root vole" HLmicTal1="Talazac's shrew tenrec" HLmicAgr2="short-tailed field vole" HLmicFor1="reed vole" HLacoCah1="Egyptian spiny mouse" HLmicArv1="Common vole" HLrhoOpi1="great gerbil" HLmasCou1="southern multimammate mouse" HLmerUng1="Mongolian gerbil" HLratRat7="black rat" HLratNor7="Norway rat" HLmusPah1="shrew mouse" HLmusCar1="Ryukyu mouse" HLmusSpi1="steppe mouse" HLmusSpr1="western wild mouse" HLapoSyl1="European woodmouse" HLvomUrs1="common wombat" HLphaCin1=koala HLgraAgi1="Agile Gracile Mouse Opossum" HLtriVul1="common brushtail" HLdidVir1="North American opossum" HLphaGym1="ground cuscus" HLgymLea1="Leadbeater's possum" HLthyCyn1="Tasmanian wolf" HLpseCup1="coppery ringtail possum" HLmacGig1="eastern gray kangaroo" HLpseCor1="golden ringtail possum" HLmacFul1="western gray kangaroo" HLnotEug3="tammar wallaby" HLospRuf1="red kangaroo" HLpseOcc1="Western ringtail oppossum" HLantFla1="yellow-footed antechinus" HLsarHar2="Tasmanian devil" HLornAna3=platypus HLtacAcu1="Australian echidna"\ subGroups view=align\ summary https://hgdownload.soe.ucsc.edu/goldenPath/hg38/multiz470way/multiz470waySummary.bb\ track multiz470way\ treeImage phylo/hg38_470way.png\ type bigMaf\ viewUi on\ nmdDetectiveB_ptc NMDetective-B PTC bigWig NMDetective-B: Decision tree NMD efficiency for first out-of-frame PTC 0 5 0 153 102 127 204 178 0 0 0\ The NMDetective tracks display genome-wide predictions of nonsense-mediated mRNA\ decay (NMD) efficiency from\ Lindeboom et al. 2016.\ NMDetective scores predict whether a premature termination codon (PTC) at a given position\ will trigger NMD and mRNA degradation, or whether the transcript will escape NMD and\ potentially produce a truncated protein.\
\ \\ Scores range from approximately −1 to +1. Positive values indicate that a PTC at\ that position is predicted to trigger NMD (the mRNA is degraded). Negative values indicate\ that the PTC is predicted to escape NMD (the truncated mRNA may be translated into an\ aberrant protein). Values near zero indicate intermediate or uncertain NMD efficiency.\
\ \| Track | Description |
|---|---|
| NMDetective-A | \Random forest model predicting NMD efficiency for all possible PTCs introduced\ by single-nucleotide variants. Explains ~71% of systematic variance in NMD\ efficiency. |
| NMDetective-B | \Simplified decision tree model for all possible PTCs. Slightly lower accuracy\ (~68% variance explained) but more interpretable, making it suitable for\ clinical applications. |
| NMDetective-A PTC | \Random forest model predicting NMD efficiency specifically for the first\ out-of-frame PTC introduced by frameshifting indel mutations. |
| NMDetective-B PTC | \Decision tree model for the first out-of-frame PTC from frameshifting\ indels. |
\ Each subtrack is displayed as a signal (bigWig) track. By default, the vertical axis\ ranges from −1 to +1. Regions with positive values (predicted NMD-triggering) are\ shown above the baseline; regions with negative values (predicted NMD escape) are shown\ below.\
\\ The NMDetective models were trained on somatic nonsense mutation data from 9,769 cancer\ patients and validated with frameshift mutations and germline variants\ (Lindeboom et al. 2019).\ The models incorporate the following features to predict NMD efficiency:\
\\ NMDetective-A (random forest regression) captures non-linear interactions among\ these features and achieves the highest predictive accuracy.\ NMDetective-B (decision tree) applies a simpler rule-based classification that\ is more transparent, with a modest reduction in accuracy.\
\ \\ The predictions were generated for every possible PTC-introducing single-nucleotide\ variant and for the first out-of-frame PTC from every possible single-nucleotide\ frameshifting indel across all human protein-coding transcripts. The original bedGraph\ custom track files were downloaded from the\ NMDetective Figshare page\ resource and converted to bigWig format at UCSC.\
\ \\ The data underlying these tracks can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Rik Lindeboom for providing custom tracks and the original NMDetective data\ on Figshare.\
\ \\ Lindeboom RG, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Lindeboom RGH, Vermeulen M, Lehner B, Supek F.\ \ The impact of nonsense-mediated mRNA decay on genetic disease, gene editing and cancer\ immunotherapy.\ Nat Genet. 2019 Nov;51(11):1645-1651.\ PMID: 31659324; PMC: PMC6858879\
\ \ genes 0 autoScale off\ bigDataUrl /gbdb/hg38/nmd/nmdDectB-ptc.bw\ color 0,153,102\ html nmdDetective\ longLabel NMDetective-B: Decision tree NMD efficiency for first out-of-frame PTC\ maxHeightPixels 128:32:8\ parent nmd off\ priority 5\ shortLabel NMDetective-B PTC\ track nmdDetectiveB_ptc\ type bigWig\ viewLimits -1:1\ visibility hide\ panelAppAusCNVs PanelApp Australia CNVs bigBed 9 + PanelApp Australia CNV Regions 3 5 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/panelApp/cnvAus.bb\ filter.versionCreated 0\ filterLabel.versionCreated Minimum panel version to display\ filterValues.confidenceLevel 3,2,1,0\ itemRgb on\ labelFields entityName\ longLabel PanelApp Australia CNV Regions\ mouseOver Gene: $entityName\ The recombination rate track represents calculated rates of recombination based\ on the genetic maps from deCODE (Halldorsson et al., 2019) and 1000 Genomes\ (2013 Phase 3 release, lifted from hg19). The deCODE map is more recent, has a higher \ resolution and was natively created on hg38 and therefore recommended. \ For the Recomb. deCODE average track, the recombination rates for chrX represent the female rate.\
\ \This track also includes a subtrack with all the\ individual deCODE recombination events and another subtrack with several thousand\ de-novo mutations found in the deCODE sequencing data. These two tracks are hidden by\ default and have to be switched on explicitly on the configuration page.\
\ \\ This is a super track that contains different subtracks, three with the deCODE\ recombination rates (paternal, maternal and average) and one with the 1000\ Genomes recombination rate (average). These tracks are in \ signal graph\ (wiggle) format. By default, to show most recombination hotspots, their maximum\ value is set to 100 cM, even though many regions have values higher than 100.\ The maximum value can be changed on the configuration pages of the tracks.\
\ \\ There are two more tracks that show additional details provided by deCODE: one\ subtrack with the raw data of all cross-overs tagged with their proband ID and\ another one with around 8000 human de-novo mutation variants that are linked to\ cross-over changes.\
\ \\ The deCODE genetic map was created at \ deCODE Genetics. It is based \ on microarrays assaying 626,828 SNP markers that allowed to identify 1,476,140 crossovers in\ 56,321 paternal meioses and 3,055,395 crossovers in 70,086 maternal meioses.\ In total, the data is based on 4,531,535 crossovers in 126,427 meioses. By\ using WGS data with 9,305,070 SNPs, the boundaries for 761,981 crossovers were\ refined: 247,942 crossovers in 9423 paternal meioses and 514,039 crossovers in\ 11,750 maternal meioses. The average resolution of the genetic map is 682 base\ pairs (bp): 655 and 708 bp for the paternal and maternal maps, respectively.\
\ \The 1000 Genomes genetic map is based on the IMPUTE genetic map based on 1000 Genomes Phase 3, on hg19 coordinates. It\ was converted to hg38 by Po-Ru Loh at the Broad Institute. After a run of \ liftOver, he post-processed the data to deal with situations in which\ consecutive map locations became much closer/farther after lifting. The\ heuristic used is sufficient for statistical phasing but may not be optimal for\ other analyses. For this reason, and because of its higher resolution, the DeCODE\ map is therefore recommended for hg38.\
\ \As with all other tracks, the data conversion commands and pointers to the\ original data files are documented in the \ makeDoc file of this track.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr17 -start=45941345 -end=45942345 http://hgdownload.soe.ucsc.edu/gbdb/hg38/recombRate/recombAvg.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ This track was produced at UCSC using data that are freely available for\ the deCODE\ and 1000 Genomes genetic maps. Thanks to Po-Ru Loh at the\ Broad Institute for providing the code to lift the hg19 1000 Genomes map data to hg38.\
\ \\ 1000 Genomes Project Consortium., Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA,\ Hurles ME, McVean GA.\ \ A map of human genome variation from population-scale sequencing.\ Nature. 2010 Oct 28;467(7319):1061-73.\ PMID: 20981092; PMC: PMC3042601\
\ \\ Halldorsson BV, Palsson G, Stefansson OA, Jonsson H, Hardarson MT, Eggertsson HP, Gunnarsson B,\ Oddsson A, Halldorsson GH, Zink F et al.\ \ Characterizing mutagenic effects of recombination through a sequence-level genetic map.\ Science. 2019 Jan 25;363(6425).\ PMID: 30679340\
\ map 1 bigDataUrl /gbdb/hg38/recombRate/recombDenovo.bb\ html recombRate2.html\ longLabel Recombination rate: De-novo mutations found in deCODE samples\ parent recombRate2\ priority 5\ shortLabel Recomb. deCODE Dmn\ track recombDnm\ type bigBed 4 +\ visibility hide\ ncbiRefSeqPsl RefSeq Alignments psl RefSeq Alignments of RNAs 1 5 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault diffCodons\ baseColorUseCds table ncbiRefSeqCds\ baseColorUseSequence extFile seqNcbiRefSeq extNcbiRefSeq\ color 0,0,0\ idXref ncbiRefSeqLink mrnaAcc name\ indelDoubleInsert on\ indelQueryInsert on\ longLabel RefSeq Alignments of RNAs\ parent refSeqComposite off\ pepTable ncbiRefSeqPepTable\ priority 5\ pslSequence no\ shortLabel RefSeq Alignments\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ track ncbiRefSeqPsl\ type psl\ revelOverlaps REVEL overlaps bigBed 9 + REVEL: Positions with >1 score due to overlapping transcripts (mouseover for details) 1 5 150 80 200 202 167 227 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/revel/overlap.bb\ extraTableFields _jsonTable|Title\ longLabel REVEL: Positions with >1 score due to overlapping transcripts (mouseover for details)\ mouseOverField _mouseOver\ parent revel on\ shortLabel REVEL overlaps\ track revelOverlaps\ type bigBed 9 +\ visibility dense\ gnomad320XPercentage Sample % > 20X bigWig gnomAD Percentage of Genome Samples with at least 20X Coverage v3.0.1 2 5 135 0 120 195 127 187 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_20.bw\ color 135,0,120\ longLabel gnomAD Percentage of Genome Samples with at least 20X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 5\ shortLabel Sample % > 20X\ track gnomad320XPercentage\ viewLimits 0:1\ gnomad4Exome20XPercentage Sample % > 20X bigWig gnomAD Percentage of Exome Samples with at least 20X Coverage v4.0 2 5 135 0 120 195 127 187 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_20.bw\ color 135,0,120\ longLabel gnomAD Percentage of Exome Samples with at least 20X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 5\ shortLabel Sample % > 20X\ track gnomad4Exome20XPercentage\ viewLimits 0:1\ genomicSuperDups Segmental Dups bed 6 + Duplications of >1000 Bases of Non-RepeatMasked Sequence 0 5 0 0 0 127 127 127 0 0 0\ This track shows regions detected as putative genomic duplications within the\ golden path. The following display conventions are used to distinguish\ levels of similarity:\
\ Segmental duplications play an important role in both genomic disease \ and gene evolution. This track displays an analysis of the global \ organization of these long-range segments of identity in genomic sequence.\
\ \Large recent duplications (>= 1 kb and >= 90% identity) were detected\ by identifying high-copy repeats, removing these repeats from the genomic \ sequence ("fuguization") and searching all sequence for similarity. The\ repeats were then reinserted into the pairwise alignments, the ends of \ alignments trimmed, and global alignments were generated.\ For a full description of the "fuguization" detection method, see Bailey\ et al., 2001. This method has become\ known as WGAC (whole-genome assembly comparison); for example, see Bailey \ et al., 2002.\ \
\ These data were provided by Ginger Cheng, Xinwei She,\ Archana Raja,\ Tin Louie and\ Evan Eichler \ at the University of Washington.
\ \\ Bailey JA, Gu Z, Clark RA, Reinert K, Samonte RV, Schwartz S, Adams MD, \ Myers EW, Li PW, Eichler EE.\ Recent segmental duplications in the human genome.\ Science. 2002 Aug 9;297(5583):1003-7.\ PMID: 12169732\
\ \\ Bailey JA, Yavor AM, Massa HF, Trask BJ, Eichler EE.\ Segmental duplications: organization and impact within the \ current human genome project assembly.\ Genome Res. 2001 Jun;11(6):1005-17.\ PMID: 11381028; PMC: PMC311093\
\ rep 1 group rep\ longLabel Duplications of >1000 Bases of Non-RepeatMasked Sequence\ noScoreFilter .\ priority 5\ shortLabel Segmental Dups\ track genomicSuperDups\ type bed 6 +\ visibility hide\ tgpHG00702_SH089_CHS SH089 CHS Trio vcfPhasedTrio 1000 Genomes Southern Han Chinese Trio 2 5 0 0 0 127 127 127 0 0 23 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX, varRep 0 longLabel 1000 Genomes Southern Han Chinese Trio\ parent tgpTrios\ shortLabel SH089 CHS Trio\ track tgpHG00702_SH089_CHS\ type vcfPhasedTrio\ vcfChildSample HG00702|child\ vcfParentSamples HG00657|mother,HG00656|father\ visibility full\ wgEncodeRegDnaseUwT47dPeak T-47D Pk narrowPeak T-47D mammary ductal carcinoma cell line DNaseI Peaks from ENCODE 1 5 255 124 85 255 189 170 1 0 0 regulation 1 color 255,124,85\ longLabel T-47D mammary ductal carcinoma cell line DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel T-47D Pk\ subGroups view=a_Peaks cellType=T-47D treatment=n_a tissue=breast cancer=cancer\ track wgEncodeRegDnaseUwT47dPeak\ wgEncodeRegDnaseUwT47dWig T-47D Sg bigWig 0 34214.8 T-47D mammary ductal carcinoma cell line DNaseI Signal from ENCODE 0 5 255 124 85 255 189 170 0 0 0 regulation 1 color 255,124,85\ longLabel T-47D mammary ductal carcinoma cell line DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.04106\ shortLabel T-47D Sg\ subGroups cellType=T-47D treatment=n_a tissue=breast cancer=cancer\ table wgEncodeRegDnaseUwT47dSignal\ track wgEncodeRegDnaseUwT47dWig\ type bigWig 0 34214.8\ unipLocTransMemb Transmembrane bigBed 12 + UniProt Transmembrane Domains 1 5 0 150 0 127 202 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipLocTransMemb.bb\ color 0,150,0\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ itemRgb off\ longLabel UniProt Transmembrane Domains\ mouseOver UniProt record: $uniProtId\ This track shows allele frequencies for 78.6 million variants from\ 4,480 whole-genome-sequenced Chinese individuals released by the\ Westlake BioBank for Chinese\ (WBBC) pilot project. The WBBC is a population study of about 35,000\ Chinese volunteers across 31 provinces; about 15,000 have been deeply\ phenotyped and a subset have been whole-genome sequenced.\ The frequencies are also broken down into four Han Chinese regional\ groups (North, Central, South, Lingnan) defined by recruitment province\ in the WBBC paper.\
\ \\ The pilot project has been folded into the larger\ China Precision BioBank\ (CPBB) initiative, which is collecting up to 100,000 samples\ nationwide. The variant frequencies on this track are from the original\ WBBC Phase I release (v20210103) and are unchanged by the rebranding.\
\ \\ The track uses the standard UCSC VCF display. Hovering a variant shows\ the cohort allele frequency, the four regional frequencies, sequencing\ depth, GATK VQSR log-odds score, and the per-genotype hom-ref / het /\ hom-alt sample counts as reported by WBBC.\
\ \\ The WBBC pilot whole-genome-sequenced 4,535 individuals at a mean depth\ of 13.9x on Illumina HiSeq X10 platforms, after dropping samples that\ failed standard QC. Reads were aligned to GRCh38 with BWA-MEM, variants\ were jointly called with GATK 4.0 HaplotypeCaller, and the callset was\ hard-filtered with VQSR. The 4,480 unrelated samples released for download\ were stratified into four Han Chinese regional groups (North, Central,\ South and Lingnan, which together cover 27 of the administrative divisions\ the pilot reached). Allele counts and frequencies are reported overall\ and per region. See Cong et al. 2022 (in References below) for\ full sample-selection and pipeline details.\
\\ The per-chromosome WGS sites VCFs (chr1-22) were downloaded from\ https://wbbc.westlake.edu.cn/\ (URL pattern: WBBC.chr<N>.GRCh38.vcf.gz). We concatenated\ the 22 files with bcftools concat, re-headered the result to\ add the standard hg38 contig lines and proper INFO definitions, then\ dropped variants with cohort allele count zero (multi-allelic splits\ that no WBBC sample carries; ~1.9% of rows), and sorted, bgzipped and\ tabix-indexed the result. No coordinate liftover was\ needed: the upstream files are already on GRCh38 with chr-prefixed\ chromosomes. The pipeline is recorded in the\ makeDoc\ file of the track.\
\ \\ Only autosomes (chr1-22) are present; chrX/Y/M are not in the WBBC\ download. Variants reported as AC=0 in the WBBC release (about 1.9 %\ of rows, mostly multi-allelic split sites that no WBBC individual\ carries) have been removed from this track.\
\ \\ The variant frequencies can be explored interactively using the\ Table Browser or the\ Data Integrator, and exported to\ spreadsheet or tab-separated tables. From scripts, the data can be\ accessed via our REST\ API with track=wbbc.\
\\ The VCF file is also available from\ our\ download server as wbbc.vcf.gz. Individual regions can be\ extracted with tabix, for example\ tabix http://hgdownload.soe.ucsc.edu/gbdb/hg38/varFreqs/wbbc/wbbc.vcf.gz chr21:1-100000000.\ The original per-chromosome WBBC release is distributed at\ https://wbbc.westlake.edu.cn/.\
\ \\ Thanks to the WBBC participants and to the Westlake University team\ (Pei-Kuan Cong, Hou-Feng Zheng and colleagues) for making the pilot\ sites-only VCFs publicly available.\
\ \\ Cong PK, Bai WY, Li JC, Yang MY, Khederzadeh S, Gai SR, Li N, Liu YH, Yu SH, Zhao WW et al.\ \ Genomic analyses of 10,376 individuals in the Westlake BioBank for Chinese (WBBC) pilot project.\ Nat Commun. 2022 May 26;13(1):2939.\ PMID: 35618720; PMC: PMC9135724\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/wbbc/wbbc.vcf.gz\ dataVersion Phase I v20210103\ longLabel SNV Frequencies: Westlake BioBank for Chinese - 4,480 WGS, 4 regional Han groups\ parent varFreqs on\ priority 6\ shortLabel China WBBC 4.5k WGS\ track wbbc\ type vcfTabix\ visibility hide\ CHOL CHOL bigLolly 12 + Cholangiocarcinoma 0 6 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/CHOL.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Cholangiocarcinoma\ parent gdcCancer off\ priority 6\ shortLabel CHOL\ track CHOL\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ lincRNAsCTColon Colon bed 5 + lincRNAs from colon 1 6 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from colon\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Colon\ subGroups view=lincRNAsRefseqExp tissueType=colon\ track lincRNAsCTColon\ wgEncodeReg4DnaseConnectiveTissue Connective tissue bigWig DNase level of 1 connective tissue experiment (tissues and primary cells only) 0 6 138 135 169 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpConnectiveTissueDNase.bw\ color 138,135,169\ longLabel DNase level of 1 connective tissue experiment (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 6\ shortLabel Connective tissue\ track wgEncodeReg4DnaseConnectiveTissue\ type bigWig\ cortexNeuron42H Cortex - Neuron - Z0000042H bigWig Methylation Atlas: Cortex - Neuron - Z0000042H 2 6 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/cortexNeuron42H.bw\ color 138,43,226\ longLabel Methylation Atlas: Cortex - Neuron - Z0000042H\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 6\ shortLabel Cortex - Neuron - Z0000042H\ subGroups cellType=Neuron dataType=Replicate\ track cortexNeuron42H\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ unipLocCytopl Cytoplasmic bigBed 12 + UniProt Cytoplasmic Domains 1 6 255 150 0 255 202 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipLocCytopl.bb\ color 255,150,0\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ itemRgb off\ longLabel UniProt Cytoplasmic Domains\ mouseOver UniProt record: $uniProtId\ The arrays listed in this track are probes from the\ Agilent Catalog Oligonucleotide Microarrays.\
\Please note that more microarray tracks are available on the hg19 genome assembly. \ To view those tracks, please \ click this link for hg19 microarrays.\ Microarrays that are not listed can be added as Custom Tracks with data from the companies.\
\ \\ Agilent's oligonucleotide CGH (Comparative Genomic Hybridization) platform enables the\ study of genome-wide DNA copy number changes at a high resolution. The CGH probes on Agilent\ CGH microarrays are 60-mer oligonucleotides synthesized in situ using Agilent's inkjet\ SurePrint technology. The probes represented on the Agilent CGH microarrays have been\ selected using algorithms developed specifically for the CGH application, assuring optimal\ performance of these probes in detecting DNA copy number changes.\
\ \\ With the Infinium MethylationEPIC BeadChip Kit, researchers can interrogate over 850,000\ methylation sites quantitatively across the genome at single-nucleotide resolution. Multiple\ samples, including FFPE, can be analyzed in parallel to deliver high-throughput power while\ minimizing the cost per sample. These tracks show positions being measured on the Illumina 450k and\ 850k (EPIC) microarray tracks, not the probe locations themselves. Contact us\ or Illumina if you need the probe locations directly. More information about\ the arrays can be found on the\ Infinium MethylationEPIC Kit website.\
\ Note: The 450k track on hg38 contains 128,989 regions representing the target regions, not the probes\ themselves.
\ \\ The Infinium CytoSNP-850K v1.2 BeadChip provides comprehensive coverage of\ cytogenetically relevant genes on a proven platform, helping researchers find valuable information\ that may be missed by other technologies. It contains approximately 850,000 empirically selected\ single nucleotide polymorphisms (SNPs) spanning the entire genome with enriched coverage for 3,262\ genes of known cytogenetics relevance in both constitutional and cancer applications. \
\ \\ The CytoScan HD Array, which is included in the\ CytoScan HD Suite, provides the broadest coverage and highest performance for\ detecting chromosomal aberrations. CytoScan HD Suite has greater than 99% sensitivity and can\ reliably detect 25-50kb copy number changes across the genome at high specificity with\ single-nucleotide polymorphism (SNP) allelic corroboration. With more than 2.6 million copy number\ markers, CytoScan HD Suite covers all OMIM and RefSeq genes.\
\ \\ Bionano Laboratories provides access to Optical Genome Mapping (OGM) data for projects across a variety of\ applications for researchers, clinicians, and pharmaceutical companies.
\This track shows the CTTAAG sites used by the \ Bionano Optical Genome Mapping system,\ an assay to detect structural variants.\
\ \\ Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \ \\ The Agilent arrays were downloaded from their \ Agilent SureDesign website tool on March 2022.
\\ The Illumina 450k and 850k (EPIC) tracks were created using a few columns from the\ Infinium MethylationEPIC v1.0 B5 Manifest File (CSV Format)\ and was then converted into a bigBed.
\\ The Illumina CytoSNP-850K track was created by downloading the\ CytoSNP-850K v1.2 Manifest File (CSV Format) (GRCh38) file and then converted\ into a bigBed file.\
\\ The Affymetrix Cytoscan HD GeneChip Array track was created by converting the \ CytoScanHD_Accel_Array.na36.bed.zip\ into a bigBed file.\
\\ The Bionano track was created by receiving the BED files from\ \ apang@bionano.\ com\ \ and converted to bigBed files using the bedToBigBed tool.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API \ or downloaded from our \ Downloads site. Please refer to our\ \ mailing list archives for questions, or our\ \ Data Access FAQ for more information.\
\ \\ Thanks to the Agilent and Illumina support teams for sharing the data and the UCSC Genome Browser\ engineers for configuring the data.
\\ Thanks to Andy Pang from Bionano Genomics for providing the BED data file.
\ varRep 1 bigDataUrl /gbdb/hg38/bbi/cytoSnp/cytoSnp850k.bb\ colorByStrand 255,0,0 0,0,255\ html genotypeArrays\ longLabel Illumina 850k CytoSNP Array\ noScoreFilter on\ parent genotypeArrays on\ priority 6\ shortLabel CytoSNP 850k\ track snpArrayCytoSnp850k\ type bigBed 6 +\ urls rsID="https://www.ncbi.nlm.nih.gov/snp/?term=$$"\ visibility pack\ dbVar_common_byrska_bishop dbVar Curated Byrska-Bishop SVs bigBed 9 + . NCBI dbVar Curated Common SVs: all populations from Byrska-Bishop 3 6 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/variants/$$ varRep 1 bigDataUrl /gbdb/hg38/bbi/dbVar/common_byrska_bishop.bb\ longLabel NCBI dbVar Curated Common SVs: all populations from Byrska-Bishop\ parent dbVar_common off\ priority 6\ shortLabel dbVar Curated Byrska-Bishop SVs\ track dbVar_common_byrska_bishop\ type bigBed 9 + .\ url https://www.ncbi.nlm.nih.gov/dbvar/variants/$$\ urlLabel NCBI Variant Page:\ wgEncodeReg4MarkH3k27acEmbryo Embryo bigWig Avg. H3K27ac level of 4 embryo experiments (tissues and primary cells only) 2 6 118 158 101 186 206 178 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpEmbryoH3K27ac.bw\ color 118,158,101\ longLabel Avg. H3K27ac level of 4 embryo experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 6\ shortLabel Embryo\ track wgEncodeReg4MarkH3k27acEmbryo\ type bigWig\ wgEncodeReg4MarkH3k4me3Embryo Embryo bigWig Avg. H3K4me3 level of 2 embryo experiments (tissues and primary cells only) 0 6 118 158 101 186 206 178 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpEmbryoH3K4me3.bw\ color 118,158,101\ longLabel Avg. H3K4me3 level of 2 embryo experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 6\ shortLabel Embryo\ track wgEncodeReg4MarkH3k4me3Embryo\ type bigWig\ ENCFF414OGC_ENCFF806YEZ_ENCFF849TDM_ENCFF736UDR ENCFF414OGC_ENCFF806YEZ_ENCFF849TDM_ENCFF736UDR bigBed 9 + 5 K562: (1) cCREs 4 6 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF414OGC_ENCFF806YEZ_ENCFF849TDM_ENCFF736UDR.bb\ longLabel K562: (1) cCREs\ mouseOver ID: ${name}\ The GENCODE Genes track (version 39, December 2021) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The following table provides statistics for the v39 release derived from the GTF file that contains\ annotations only on the main chromosomes. More information on how they were generated can be found\ in the GENCODE site.
\ \\
\ \\
\ GENCODE v39 Release Stats \ Genes Observed Transcripts Observed \ Protein-coding genes 19,982 Protein-coding transcripts 87,151 \ Long non-coding RNA genes 18,811 - full length protein-coding 61,516 \ Small non-coding RNA genes 7,567 - partial length protein-coding 25,635 \ Pseudogenes 14,763 Nonsense mediated decay transcripts 19,762 \ Immunoglobulin/T-cell receptor gene segments 409 Long non-coding RNA loci transcripts 53,009
\
\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \\
The GENCODE v39 track was built from the GENCODE downloads file \
gencode.v39.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources \
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney.
\ \\ Harrow J, Frankish A, Gonzalez JM, Tapanari E, Diekhans M, Kokocinski F, Aken BL, Barrell D, Zadissa\ A, Searle S et al.\ \ GENCODE: the reference human genome annotation for The ENCODE Project.\ Genome Res. 2012 Sep;22(9):1760-74.\ PMID: 22955987; PMC: PMC3431492\
\ \\ Harrow J, Denoeud F, Frankish A, Reymond A, Chen CK, Chrast J, Lagarde J, Gilbert JG, Storey R,\ Swarbreck D et al.\ \ GENCODE: producing a reference annotation for ENCODE.\ Genome Biol. 2006;7 Suppl 1:S4.1-9.\ PMID: 16925838; PMC: PMC1810553\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV39.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV39\ group genes\ html knownGeneV39\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V39\ maxItems 50000\ parent knownGeneArchive\ priority 6\ searchIndex name\ shortLabel GENCODE V39\ track knownGeneV39\ type bigGenePred\ visibility hide\ geneHancerGenes GH genes TSS bigBed 9 GH genes TSS 3 6 0 0 0 127 127 127 0 0 0 http://www.genecards.org/cgi-bin/carddisp.pl?gene=$$ regulation 1 bigDataUrl /gbdb/hg38/geneHancer/geneHancerGenesTssAll.hg38.bb\ longLabel GH genes TSS\ parent ghGeneTss off\ shortLabel GH genes TSS\ subGroups set=b_ALL view=b_TSS\ track geneHancerGenes\ type bigBed 9\ urlLabel In GeneCards:\ wgEncodeReg4MarkCtcfHeart Heart bigWig Avg. CTCF level of 24 heart experiments (tissues and primary cells only) 0 6 116 50 165 185 152 210 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpHeartCTCF.bw\ color 116,50,165\ longLabel Avg. CTCF level of 24 heart experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf off\ priority 6\ shortLabel Heart\ track wgEncodeReg4MarkCtcfHeart\ type bigWig\ netHprcGCA_018467015v1 HG02486.mat netAlign GCA_018467015.1 chainHprcGCA_018467015v1 HG02486.mat HG02486.pri.mat.f1_v2 (May 2021 GCA_018467015.1_HG02486.pri.mat.f1_v2) HPRC project computed Chain Nets 1 6 0 0 0 255 255 0 0 0 0 hprc 0 longLabel HG02486.mat HG02486.pri.mat.f1_v2 (May 2021 GCA_018467015.1_HG02486.pri.mat.f1_v2) HPRC project computed Chain Nets\ otherDb GCA_018467015.1\ parent hprcChainNetViewnet off\ priority 22\ shortLabel HG02486.mat\ subGroups view=net sample=s022 population=afr subpop=acb hap=mat\ track netHprcGCA_018467015v1\ type netAlign GCA_018467015.1 chainHprcGCA_018467015v1\ hr_na12248Vcf HR_NA12248 Variants vcfTabix HR_NA12248 Variants 0 6 0 0 0 127 127 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/problematic/highRepro/HR_NA12248.sort.vcf.gz\ longLabel HR_NA12248 Variants\ parent highReproVcfs\ shortLabel HR_NA12248 Variants\ subGroups view=vcfs\ track hr_na12248Vcf\ type vcfTabix\ wgEncodeRegTxnCaltechRnaSeqHuvecR2x75Il200SigPooled HUVEC bigWig 0 65535 Transcription of HUVEC cells from ENCODE 0 6 128 199 255 191 227 255 0 0 0 regulation 1 color 128,199,255\ longLabel Transcription of HUVEC cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegTxn\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 6\ shortLabel HUVEC\ track wgEncodeRegTxnCaltechRnaSeqHuvecR2x75Il200SigPooled\ type bigWig 0 65535\ KAPA_HyperExome_hg38_primary_targets KAPA Hyper T bigBed Roche - KAPA HyperExome Primary Target Regions 0 6 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/KAPA_HyperExome_hg38_primary_targets.bb\ color 100,143,255\ longLabel Roche - KAPA HyperExome Primary Target Regions\ parent exomeProbesets off\ shortLabel KAPA Hyper T\ track KAPA_HyperExome_hg38_primary_targets\ type bigBed\ wgEncodeReg4AtacLung Lung bigWig Avg. ATAC level of 6 lung experiments (tissues and primary cells only) 0 6 130 163 45 192 209 150 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLungATAC.bw\ color 130,163,45\ longLabel Avg. ATAC level of 6 lung experiments (tissues and primary cells only)\ parent wgEncodeReg4Atac off\ priority 6\ shortLabel Lung\ track wgEncodeReg4AtacLung\ type bigWig\ wgEncodeRegMarkH3k27acNhek NHEK bigWig 0 23439 H3K27Ac Mark (Often Found Near Regulatory Elements) on NHEK Cells from ENCODE 2 6 212 128 255 233 191 255 0 0 0 regulation 1 color 212,128,255\ longLabel H3K27Ac Mark (Often Found Near Regulatory Elements) on NHEK Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k27ac\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel NHEK\ table wgEncodeBroadHistoneNhekH3k27acStdSig\ track wgEncodeRegMarkH3k27acNhek\ type bigWig 0 23439\ wgEncodeRegMarkH3k4me1Nhek NHEK bigWig 0 2669 H3K4Me1 Mark (Often Found Near Regulatory Elements) on NHEK Cells from ENCODE 0 6 212 128 255 233 191 255 0 0 0 regulation 1 color 212,128,255\ longLabel H3K4Me1 Mark (Often Found Near Regulatory Elements) on NHEK Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k4me1\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel NHEK\ table wgEncodeBroadHistoneNhekH3k4me1StdSig\ track wgEncodeRegMarkH3k4me1Nhek\ type bigWig 0 2669\ wgEncodeRegMarkH3k4me3Nhek NHEK bigWig 0 8230 H3K4Me3 Mark (Often Found Near Promoters) on NHEK Cells from ENCODE 0 6 212 128 255 233 191 255 0 0 0 regulation 1 color 212,128,255\ longLabel H3K4Me3 Mark (Often Found Near Promoters) on NHEK Cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegMarkH3k4me3\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel NHEK\ table wgEncodeBroadHistoneNhekH3k4me3StdSig\ track wgEncodeRegMarkH3k4me3Nhek\ type bigWig 0 8230\ wgEncodeRegDnaseUwPanc1Peak PANC-1 Pk narrowPeak PANC-1 pancreatic carcinoma cell line DNaseI Peaks from ENCODE 1 6 255 141 85 255 198 170 1 0 0 regulation 1 color 255,141,85\ longLabel PANC-1 pancreatic carcinoma cell line DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel PANC-1 Pk\ subGroups view=a_Peaks cellType=PANC-1 treatment=n_a tissue=pancreas cancer=cancer\ track wgEncodeRegDnaseUwPanc1Peak\ wgEncodeRegDnaseUwPanc1Wig PANC-1 Sg bigWig 0 12279.3 PANC-1 pancreatic carcinoma cell line DNaseI Signal from ENCODE 0 6 255 141 85 255 198 170 0 0 0 regulation 1 color 255,141,85\ longLabel PANC-1 pancreatic carcinoma cell line DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.05908\ shortLabel PANC-1 Sg\ subGroups cellType=PANC-1 treatment=n_a tissue=pancreas cancer=cancer\ table wgEncodeRegDnaseUwPanc1Signal\ track wgEncodeRegDnaseUwPanc1Wig\ type bigWig 0 12279.3\ panelAppAusTandRep PanelApp Australia STRs bigBed 9 + PanelApp Australia Short Tandem Repeats 3 6 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/panelApp/tandRepAus.bb\ filter.version 0\ filterLabel.version Minimum panel version to display\ filterValues.confidenceLevel 3,2,1,0\ itemRgb on\ labelFields hgncSymbol\ longLabel PanelApp Australia Short Tandem Repeats\ mouseOver Gene name: $geneName\ The recombination rate track represents calculated rates of recombination based\ on the genetic maps from deCODE (Halldorsson et al., 2019) and 1000 Genomes\ (2013 Phase 3 release, lifted from hg19). The deCODE map is more recent, has a higher \ resolution and was natively created on hg38 and therefore recommended. \ For the Recomb. deCODE average track, the recombination rates for chrX represent the female rate.\
\ \This track also includes a subtrack with all the\ individual deCODE recombination events and another subtrack with several thousand\ de-novo mutations found in the deCODE sequencing data. These two tracks are hidden by\ default and have to be switched on explicitly on the configuration page.\
\ \\ This is a super track that contains different subtracks, three with the deCODE\ recombination rates (paternal, maternal and average) and one with the 1000\ Genomes recombination rate (average). These tracks are in \ signal graph\ (wiggle) format. By default, to show most recombination hotspots, their maximum\ value is set to 100 cM, even though many regions have values higher than 100.\ The maximum value can be changed on the configuration pages of the tracks.\
\ \\ There are two more tracks that show additional details provided by deCODE: one\ subtrack with the raw data of all cross-overs tagged with their proband ID and\ another one with around 8000 human de-novo mutation variants that are linked to\ cross-over changes.\
\ \\ The deCODE genetic map was created at \ deCODE Genetics. It is based \ on microarrays assaying 626,828 SNP markers that allowed to identify 1,476,140 crossovers in\ 56,321 paternal meioses and 3,055,395 crossovers in 70,086 maternal meioses.\ In total, the data is based on 4,531,535 crossovers in 126,427 meioses. By\ using WGS data with 9,305,070 SNPs, the boundaries for 761,981 crossovers were\ refined: 247,942 crossovers in 9423 paternal meioses and 514,039 crossovers in\ 11,750 maternal meioses. The average resolution of the genetic map is 682 base\ pairs (bp): 655 and 708 bp for the paternal and maternal maps, respectively.\
\ \The 1000 Genomes genetic map is based on the IMPUTE genetic map based on 1000 Genomes Phase 3, on hg19 coordinates. It\ was converted to hg38 by Po-Ru Loh at the Broad Institute. After a run of \ liftOver, he post-processed the data to deal with situations in which\ consecutive map locations became much closer/farther after lifting. The\ heuristic used is sufficient for statistical phasing but may not be optimal for\ other analyses. For this reason, and because of its higher resolution, the DeCODE\ map is therefore recommended for hg38.\
\ \As with all other tracks, the data conversion commands and pointers to the\ original data files are documented in the \ makeDoc file of this track.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr17 -start=45941345 -end=45942345 http://hgdownload.soe.ucsc.edu/gbdb/hg38/recombRate/recombAvg.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ This track was produced at UCSC using data that are freely available for\ the deCODE\ and 1000 Genomes genetic maps. Thanks to Po-Ru Loh at the\ Broad Institute for providing the code to lift the hg19 1000 Genomes map data to hg38.\
\ \\ 1000 Genomes Project Consortium., Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA,\ Hurles ME, McVean GA.\ \ A map of human genome variation from population-scale sequencing.\ Nature. 2010 Oct 28;467(7319):1061-73.\ PMID: 20981092; PMC: PMC3042601\
\ \\ Halldorsson BV, Palsson G, Stefansson OA, Jonsson H, Hardarson MT, Eggertsson HP, Gunnarsson B,\ Oddsson A, Halldorsson GH, Zink F et al.\ \ Characterizing mutagenic effects of recombination through a sequence-level genetic map.\ Science. 2019 Jan 25;363(6425).\ PMID: 30679340\
\ map 0 bigDataUrl /gbdb/hg38/recombRate/recomb1000GAvg.bw\ html recombRate2.html\ longLabel Recombination rate: 1000 Genomes, lifted from hg19 (PR Loh)\ maxHeightPixels 128:60:8\ parent recombRate2\ priority 6\ shortLabel Recomb. 1k Genomes\ track recomb1000GAvg\ type bigWig\ viewLimits 0.0:100\ viewLimitsMax 0:150000\ visibility full\ ncbiRefSeqGenomicDiff RefSeq Diffs bigBed 9 + Differences between NCBI RefSeq Transcripts and the Reference Genome 1 6 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/ncbiRefSeq/ncbiRefSeqGenomicDiff.bb\ itemRgb on\ longLabel Differences between NCBI RefSeq Transcripts and the Reference Genome\ parent refSeqComposite off\ priority 6\ shortLabel RefSeq Diffs\ skipEmptyFields on\ track ncbiRefSeqGenomicDiff\ type bigBed 9 +\ gnomad325XPercentage Sample % > 25X bigWig gnomAD Percentage of Genome Samples with at least 25X Coverage v3.0.1 2 6 105 0 150 180 127 202 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_25.bw\ color 105,0,150\ longLabel gnomAD Percentage of Genome Samples with at least 25X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 6\ shortLabel Sample % > 25X\ track gnomad325XPercentage\ viewLimits 0:1\ gnomad4Exome25XPercentage Sample % > 25X bigWig gnomAD Percentage of Exome Samples with at least 25X Coverage v4.0 2 6 105 0 150 180 127 202 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_25.bw\ color 105,0,150\ longLabel gnomAD Percentage of Exome Samples with at least 25X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 6\ shortLabel Sample % > 25X\ track gnomad4Exome25XPercentage\ viewLimits 0:1\ chainSelf Self Alignment chain hg38 Human Chained Self Alignments 0 6 100 50 0 255 240 200 1 0 0\ This track shows alignments of the human genome with itself, using\ a gap scoring system that allows longer gaps than traditional\ affine gap scoring systems. The system can also tolerate gaps\ in both sets of sequence simultaneously. After filtering out the \ "trivial" alignments produced when identical locations of the \ genome map to one another (e.g. chrN mapping to chrN), \ the remaining alignments point out areas of duplication within the \ human genome. The pseudoautosomal regions of chrX and chrY are an \ exception: in this assembly, these regions have been copied from chrX into \ chrY, resulting in a large amount of self chains aligning in these positions \ on both chromosomes.
\\ The chain track displays boxes joined together by either single or\ double lines. The boxes represent aligning regions. Single lines indicate \ gaps that are largely due to a deletion in the query assembly or an \ insertion in the target assembly. Double lines represent more complex gaps \ that involve substantial sequence in both the query and target assemblies. \ This may result from inversions, overlapping deletions, an abundance of local \ mutation, or an unsequenced gap in one of the assemblies. In cases where \ multiple chains align over a particular region of the human genome, the \ chains with single-lined gaps are often due to processed pseudogenes, while \ chains with double-lined gaps are more often due to paralogs and unprocessed \ pseudogenes.
\\ Chains have both a score and a normalized score. The score is derived by \ comparing sequence similarity, while penalizing both mismatches and gaps\ in a per base fashion. This leads to longer chains having greater scores, \ even if a smaller chain provides a better match. The normalized score divides\ the score by the length of the alignment, providing a more comparable score value\ not dependent on the match length.
\ \By default, the chains are colored by the normalized score. This can be changed\ to color based on which chromosome they map to in the aligning organism. There is also\ an option to color all the chains black.
\\ To display only the chains of one chromosome in the aligning\ organism, enter the name of that chromosome (e.g. chr4) in box next to: \ Filter by chromosome.
\\ By default, chains with a score of 20,000 or more are displayed. This default value provides\ a conservative cutoff, filtering out many false-positive alignments with low sequence \ similarity, or high penalties. It should be noted however, that alignments below this \ threshold may still be indicative of homology.
\\ In the "pack" and "full" display\ modes, the individual feature names indicate the chromosome, strand, and\ location (in thousands) of the match for each matching alignment.
\ \\ The genome was aligned to itself using blastz. Trivial alignments were \ filtered out, and the remaining alignments were converted into axt format\ using the lavToAxt program. The axt alignments were fed into axtChain, which \ organizes all alignments between a single target chromosome and a single\ query chromosome into a group and creates a kd-tree out of the gapless \ subsections (blocks) of the alignments. A dynamic program was then run over \ the kd-trees to find the maximally scoring chains of these blocks. Chains \ scoring below a threshold were discarded; the remaining chains are displayed \ in this track.
\ \\ Blastz was developed at Pennsylvania State University by\ Minmei Hou, Scott Schwartz, Zheng Zhang, and Webb Miller with advice from\ Ross Hardison.
\\ Lineage-specific repeats were identified by Arian Smit and his\ RepeatMasker\ program.
\\ The axtChain program was developed at the University of California\ at Santa Cruz by Jim Kent with advice from Webb Miller and David Haussler.\
\\ The browser display and database storage of the chains were generated\ by Robert Baertsch and Jim Kent.
\ \\ Chiaromonte F, Yap VB, Miller W.\ Scoring pairwise genomic sequence alignments.\ Pac Symp Biocomput 2002, 115-26 (2002).\
\ \\ Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D.\ \ Evolution's cauldron: duplication, deletion, and rearrangement in the mouse and human genomes.\ Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9.\
\ \\ Schwartz S, Kent WJ, Smit A, Zhang Z, Baertsch R, Hardison RC, Haussler D, Miller W.\ \ Human-mouse alignments with BLASTZ.\ Genome Res. 2003 Jan;13(1):103-7.\
\ rep 1 altColor 255,240,200\ baseColorDefault diffBases\ baseColorUseSequence 2bit\ chainColor Normalized Score\ chainNormScoreAvailable yes\ color 100,50,0\ group rep\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Human Chained Self Alignments\ matrix 16 91,-114,-31,-123,-114,100,-125,-31,-31,-125,100,-114,-123,-31,-114,91\ matrixHeader A, C, G, T\ otherDb hg38\ otherTwoBitUrl /gbdb/hg38/hg38.2bit\ priority 6\ scoreFilter 20000\ shortLabel Self Alignment\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 100000.0\ spectrum on\ track chainSelf\ type chain hg38\ visibility hide\ ucneClusters UCNE Clusters bigBed 4 + UCNEBase: 239 Cluster of UCNE elements 0 6 0 0 0 127 127 127 0 0 0 https://epd.expasy.org/ucnebase/view.php?data=cluster&entry=$$ compGeno 1 bigDataUrl /gbdb/hg38/unusualcons/clusters.bb\ longLabel UCNEBase: 239 Cluster of UCNE elements\ parent unusualcons on\ shortLabel UCNE Clusters\ track ucneClusters\ type bigBed 4 +\ url https://epd.expasy.org/ucnebase/view.php?data=cluster&entry=$$\ umap36Quantitative Umap M36 bigWig 0.027778 1.0 Multi-read mappability with 36-mers 0 6 80 70 240 167 162 247 0 0 0 map 0 bigDataUrl /gbdb/hg38/hoffmanMappability/k36.Umap.MultiTrackMappability.bw\ color 80,70,240\ longLabel Multi-read mappability with 36-mers\ parent umapBigWig off\ priority 6\ shortLabel Umap M36\ subGroups view=MR\ track umap36Quantitative\ type bigWig 0.027778 1.0\ visibility hide\ iscaLikelyPathogenic Uncert Path gvf ClinGen CNVs: Uncertain: Likely Pathogenic 3 6 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/?term=$$ phenDis 1 longLabel ClinGen CNVs: Uncertain: Likely Pathogenic\ parent iscaViewDetail off\ shortLabel Uncert Path\ subGroups view=cnv class=likP level=sub\ track iscaLikelyPathogenic\ tgpHG02024_VN049_KHV VN049 KHV Trio vcfPhasedTrio 1000 Genomes Kinh in Ho Chi Minh City, Vietnam Trio 2 6 0 0 0 127 127 127 0 0 23 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX, varRep 0 longLabel 1000 Genomes Kinh in Ho Chi Minh City, Vietnam Trio\ parent tgpTrios\ shortLabel VN049 KHV Trio\ track tgpHG02024_VN049_KHV\ type vcfPhasedTrio\ vcfChildSample HG02024|child\ vcfParentSamples HG02025|mother,HG02026|father\ visibility full\ chainAquChr2 aquChr2 Chain chain aquChr2 Golden eagle (Oct. 2014 (aquChr-1.0.2/aquChr2)) Chained Alignments 3 7 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Golden eagle (Oct. 2014 (aquChr-1.0.2/aquChr2)) Chained Alignments\ otherDb aquChr2\ parent vertebrateChainNetViewchain off\ shortLabel aquChr2 Chain\ subGroups view=chain species=s016 clade=c01\ track chainAquChr2\ type chain aquChr2\ chainPonAbe3 Orangutan Chain chain ponAbe3 Orangutan (Jan. 2018 (Susie_PABv2/ponAbe3)) Chained Alignments 3 7 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Orangutan (Jan. 2018 (Susie_PABv2/ponAbe3)) Chained Alignments\ otherDb ponAbe3\ parent primateChainNetViewchain off\ shortLabel Orangutan Chain\ subGroups view=chain species=s013a clade=c00\ track chainPonAbe3\ type chain ponAbe3\ netMm39 Mouse Net netAlign mm39 chainMm39 Mouse (Jun. 2020 (GRCm39/mm39)) Alignment Net 1 7 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Mouse (Jun. 2020 (GRCm39/mm39)) Alignment Net\ otherDb mm39\ parent placentalChainNetViewnet on\ shortLabel Mouse Net\ subGroups view=net species=s012a clade=c00\ track netMm39\ type netAlign mm39 chainMm39\ encTfChipPkENCFF576PUH A549 CREB1 1 narrowPeak Transcription Factor ChIP-seq Peaks of CREB1 in A549 from ENCODE 3 (ENCFF576PUH) 0 7 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of CREB1 in A549 from ENCODE 3 (ENCFF576PUH)\ parent encTfChipPk off\ shortLabel A549 CREB1 1\ subGroups cellType=A549 factor=CREB1\ track encTfChipPkENCFF576PUH\ cloneEndABC18 ABC18 bed 12 Agencourt fosmid library 18 0 7 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 18\ parent cloneEndSuper off\ priority 7\ shortLabel ABC18\ subGroups source=agencourt\ track cloneEndABC18\ type bed 12\ visibility hide\ affyCytoScanHD Affy CytoScan HD bigBed 12 Affymetrix Cytoscan HD GeneChip Array 3 7 0 0 0 127 127 127 0 0 0\ The arrays listed in this track are probes from the\ Agilent Catalog Oligonucleotide Microarrays.\
\Please note that more microarray tracks are available on the hg19 genome assembly. \ To view those tracks, please \ click this link for hg19 microarrays.\ Microarrays that are not listed can be added as Custom Tracks with data from the companies.\
\ \\ Agilent's oligonucleotide CGH (Comparative Genomic Hybridization) platform enables the\ study of genome-wide DNA copy number changes at a high resolution. The CGH probes on Agilent\ CGH microarrays are 60-mer oligonucleotides synthesized in situ using Agilent's inkjet\ SurePrint technology. The probes represented on the Agilent CGH microarrays have been\ selected using algorithms developed specifically for the CGH application, assuring optimal\ performance of these probes in detecting DNA copy number changes.\
\ \\ With the Infinium MethylationEPIC BeadChip Kit, researchers can interrogate over 850,000\ methylation sites quantitatively across the genome at single-nucleotide resolution. Multiple\ samples, including FFPE, can be analyzed in parallel to deliver high-throughput power while\ minimizing the cost per sample. These tracks show positions being measured on the Illumina 450k and\ 850k (EPIC) microarray tracks, not the probe locations themselves. Contact us\ or Illumina if you need the probe locations directly. More information about\ the arrays can be found on the\ Infinium MethylationEPIC Kit website.\
\ Note: The 450k track on hg38 contains 128,989 regions representing the target regions, not the probes\ themselves.
\ \\ The Infinium CytoSNP-850K v1.2 BeadChip provides comprehensive coverage of\ cytogenetically relevant genes on a proven platform, helping researchers find valuable information\ that may be missed by other technologies. It contains approximately 850,000 empirically selected\ single nucleotide polymorphisms (SNPs) spanning the entire genome with enriched coverage for 3,262\ genes of known cytogenetics relevance in both constitutional and cancer applications. \
\ \\ The CytoScan HD Array, which is included in the\ CytoScan HD Suite, provides the broadest coverage and highest performance for\ detecting chromosomal aberrations. CytoScan HD Suite has greater than 99% sensitivity and can\ reliably detect 25-50kb copy number changes across the genome at high specificity with\ single-nucleotide polymorphism (SNP) allelic corroboration. With more than 2.6 million copy number\ markers, CytoScan HD Suite covers all OMIM and RefSeq genes.\
\ \\ Bionano Laboratories provides access to Optical Genome Mapping (OGM) data for projects across a variety of\ applications for researchers, clinicians, and pharmaceutical companies.
\This track shows the CTTAAG sites used by the \ Bionano Optical Genome Mapping system,\ an assay to detect structural variants.\
\ \\ Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \ \\ The Agilent arrays were downloaded from their \ Agilent SureDesign website tool on March 2022.
\\ The Illumina 450k and 850k (EPIC) tracks were created using a few columns from the\ Infinium MethylationEPIC v1.0 B5 Manifest File (CSV Format)\ and was then converted into a bigBed.
\\ The Illumina CytoSNP-850K track was created by downloading the\ CytoSNP-850K v1.2 Manifest File (CSV Format) (GRCh38) file and then converted\ into a bigBed file.\
\\ The Affymetrix Cytoscan HD GeneChip Array track was created by converting the \ CytoScanHD_Accel_Array.na36.bed.zip\ into a bigBed file.\
\\ The Bionano track was created by receiving the BED files from\ \ apang@bionano.\ com\ \ and converted to bigBed files using the bedToBigBed tool.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API \ or downloaded from our \ Downloads site. Please refer to our\ \ mailing list archives for questions, or our\ \ Data Access FAQ for more information.\
\ \\ Thanks to the Agilent and Illumina support teams for sharing the data and the UCSC Genome Browser\ engineers for configuring the data.
\\ Thanks to Andy Pang from Bionano Genomics for providing the BED data file.
\ varRep 1 bigDataUrl /gbdb/hg38/genotypeArrays/affyCytoScanHD.bb\ html genotypeArrays\ itemRgb on\ longLabel Affymetrix Cytoscan HD GeneChip Array\ mouseOver Probe ID: $name\ This track shows allele frequencies for single-nucleotide variants (SNVs)\ and small indels joint-called across 1,027 population-consented human\ samples sequenced with PacBio HiFi long reads by the Consortium of Long\ Read Sequencing (CoLoRSdb). Sites were joint-genotyped with\ DeepVariant\ and merged across samples with\ GLnexus.
\ \\ Two versions of the callset are displayed, chosen automatically according\ to the browser's reference assembly:
\\ Only population-level allele frequencies (AF), allele counts (AC), and\ total allele numbers (AN) are displayed; per-sample genotypes are not\ included in the released VCF. Multi-allelic sites have been decomposed so\ that each alternate allele has its own VCF row.
\ \\ The track uses the standard UCSC VCF representation. Mouseover\ shows the variant, reference/alternate alleles, AC, AN and AF. At higher\ zoom levels, alleles use base-specific colors. Homozygous ALT positions\ are marked with one letter, heterozygotes with two letters.
\ \\ CoLoRSdb member sites sequenced 1,027 individuals with PacBio HiFi\ long reads. Per-sample variant calls came from DeepVariant. GLnexus\ then joint-genotyped them into a cohort-wide population VCF. Only\ sites that passed CoLoRSdb's standard quality filters are included.
\ \\ The data can be explored interactively with the\ Table Browser or\ Data Integrator, and accessed from\ scripts via our API\ (track=colorsDbSnv).
\ \The VCF files are available on our download server:\ \ GRCh38 VCF and\ \ CHM13 VCF.\ They are hard symlinks to the upstream CoLoRSdb releases, which are\ distributed from\ colorsdb.org and the\ CoLoRSdb GitHub\ repositories.
\ \Thanks to the Consortium of Long Read Sequencing (CoLoRSdb) members\ and PacBio, who produced and released these joint-called long-read\ variant frequencies.
\ \See the main SNV Frequencies\ container track for the general context and a comparison table across\ all included frequency databases.
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/colorsDb/colorsDbSnv.vcf.gz\ dataVersion v1.2.0\ longLabel SNV Frequencies: CoLoRSdb v1.2.0 - 1,027 PacBio HiFi WGS, SNV/indel callset\ parent varFreqs on\ priority 7\ shortLabel CoLoRSdb 1k LR SNV/Ind\ track colorsDbSnv\ type vcfTabix\ visibility hide\ cortexNeuron42J Cortex - Neuron - Z0000042J bigWig Methylation Atlas: Cortex - Neuron - Z0000042J 2 7 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/cortexNeuron42J.bw\ color 138,43,226\ longLabel Methylation Atlas: Cortex - Neuron - Z0000042J\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 7\ shortLabel Cortex - Neuron - Z0000042J\ subGroups cellType=Neuron dataType=Replicate\ track cortexNeuron42J\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ iscaCuratedPathogenic Curated Path gvf ClinGen CNVs: Curated Pathogenic 3 7 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/?term=$$ phenDis 1 longLabel ClinGen CNVs: Curated Pathogenic\ parent iscaViewDetail off\ shortLabel Curated Path\ subGroups view=cnv class=path level=cur\ track iscaCuratedPathogenic\ wgEncodeReg4DnaseEmbryo Embryo bigWig Avg. DNase level of 9 embryo experiments (tissues and primary cells only) 0 7 118 158 101 186 206 178 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpEmbryoDNase.bw\ color 118,158,101\ longLabel Avg. DNase level of 9 embryo experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 7\ shortLabel Embryo\ track wgEncodeReg4DnaseEmbryo\ type bigWig\ ENCFF428XFI_ENCFF280PUF_ENCFF469WVA_ENCFF644EEX ENCFF428XFI_ENCFF280PUF_ENCFF469WVA_ENCFF644EEX bigBed 9 + 5 GM12878: (1) cCREs 4 7 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF428XFI_ENCFF280PUF_ENCFF469WVA_ENCFF644EEX.bb\ longLabel GM12878: (1) cCREs\ mouseOver ID: ${name}\ The GENCODE Genes track (version 38, May 2021) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The following table provides statistics for the v38 release derived from the GTF file that contains\ annotations only on the main chromosomes. More information on how they were generated can be found\ in the GENCODE site.
\ \\
\ \\
\ GENCODE v38 Release Stats \ Genes Observed Transcripts Observed \ Protein-coding genes 19,955 Protein-coding transcripts 86,757 \ Long non-coding RNA genes 17,944 - full length protein-coding 61,015 \ Small non-coding RNA genes 7,567 - partial length protein-coding 25,742 \ Pseudogenes 14,773 Nonsense mediated decay transcripts 18,881 \ Immunoglobulin/T-cell receptor gene segments 409 Long non-coding RNA loci transcripts 48,752
\
\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \\
The GENCODE v38 track was built from the GENCODE downloads file \
gencode.v38.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources \
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney.
\ \\ Harrow J, Frankish A, Gonzalez JM, Tapanari E, Diekhans M, Kokocinski F, Aken BL, Barrell D, Zadissa\ A, Searle S et al.\ \ GENCODE: the reference human genome annotation for The ENCODE Project.\ Genome Res. 2012 Sep;22(9):1760-74.\ PMID: 22955987; PMC: PMC3431492\
\ \\ Harrow J, Denoeud F, Frankish A, Reymond A, Chen CK, Chrast J, Lagarde J, Gilbert JG, Storey R,\ Swarbreck D et al.\ \ GENCODE: producing a reference annotation for ENCODE.\ Genome Biol. 2006;7 Suppl 1:S4.1-9.\ PMID: 16925838; PMC: PMC1810553\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV38.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV38\ group genes\ html knownGeneV38\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V38\ maxItems 50000\ parent knownGeneArchive\ priority 7\ searchIndex name\ shortLabel GENCODE V38\ track knownGeneV38\ type bigGenePred\ visibility hide\ geneHancerInteractions GH Interactions bigInteract Interactions between GeneHancer regulatory elements and genes 2 7 0 0 0 127 127 127 0 0 0 https://www.genecards.org/cgi-bin/carddisp.pl?gene=$\ The chain track shows alignments of human (Dec. 2013 (GRCh38/hg38)) to\ other genomes using a gap scoring system that allows longer gaps \ than traditional affine gap scoring systems. It can also tolerate gaps in both\ human and the other genome simultaneously. These \ "double-sided" gaps can be caused by local inversions and \ overlapping deletions in both species. \
\ The chain track displays boxes joined together by either single or\ double lines. The boxes represent aligning regions.\ Single lines indicate gaps that are largely due to a deletion in the\ other assembly or an insertion in the human assembly.\ Double lines represent more complex gaps that involve substantial\ sequence in both species. This may result from inversions, overlapping\ deletions, an abundance of local mutation, or an unsequenced gap in one\ species. In cases where multiple chains align over a particular region of\ the other genome, the chains with single-lined gaps are often \ due to processed pseudogenes, while chains with double-lined gaps are more \ often due to paralogs and unprocessed pseudogenes.
\\ In the "pack" and "full" display\ modes, the individual feature names indicate the chromosome, strand, and\ location (in thousands) of the match for each matching alignment.
\ \\ The net track shows only the alignments from the highest-scoring chain\ for each region of the human genome assembly. It is useful for finding\ orthologous regions and for studying genome rearrangement. The human\ sequence used in this annotation is from the Dec. 2013 (GRCh38/hg38) assembly.
\ \By default, the chains to chromosome-based assemblies are colored\ based on which chromosome they map to in the aligning organism. To turn\ off the coloring, check the "off" button next to: Color\ track based on chromosome.
\\ To display only the chains of one chromosome in the aligning\ organism, enter the name of that chromosome (e.g. chr4) in box next to: \ Filter by chromosome.
\ \\ In full display mode, the top-level (level 1)\ chains are the largest, highest-scoring chains that\ span this region. In many cases gaps exist in the\ top-level chain. When possible, these are filled in by\ other chains that are displayed at level 2. The gaps in \ level 2 chains may be filled by level 3 chains and so\ forth.
\\ In the graphical display, the boxes represent ungapped \ alignments; the lines represent gaps. Click\ on a box to view detailed information about the chain\ as a whole; click on a line to display information\ about the gap. The detailed information is useful in determining\ the cause of the gap or, for lower level chains, the genomic\ rearrangement.
\\ Individual items in the display are categorized as one of four types\ (other than gap):
\\ The assemblies were examined for any transposons that had been inserted\ since the divergence of the two species. Any such transposons were\ removed before running the alignment. The abbreviated genomes were\ aligned with lastz, and the removed transposons were then added back in.\ The resulting alignments were converted into axt format using the lavToAxt\ program. The axt alignments were fed into axtChain, which organizes all\ alignments between a single human chromosome and a single\ chromosome from the other genome into a group and creates a kd-tree out\ of the gapless subsections (blocks) of the alignments. A dynamic program\ was then run over the kd-trees to find the maximally scoring chains of these\ blocks.\
\ The lastz matrices used for these alignments can be found in our\ download directory\ for the Dec. 2013 (GRCh38/hg38) assembly. See the README.txt file within the relevant\ vsAssembly directory for details (e.g., parameters for the alignment with\ tarSyr2 can be found in the vsTarSyr2/ subdirectory).\
\
For the alignments to Chimp and Rhesus, chains scoring below a minimum\
score of '5000' were discarded; the remaining chains\
are displayed in this track. The linear gap matrix used with axtChain:
\
-linearGap=loose\ \ tablesize 11\ smallSize 111\ position 1 2 3 11 111 2111 12111 32111 72111 152111 252111\ qGap 325 360 400 450 600 1100 3600 7600 15600 31600 56600\ tGap 325 360 400 450 600 1100 3600 7600 15600 31600 56600\ bothGap 625 660 700 750 900 1400 4000 8000 16000 32000 57000\\ \ For the alignments to Tarsier and Bonobo, chains scoring\ below a minimum score of '3000' were discarded; the remaining chains\ are displayed in this track. The same linear gap matrix shown above\ was used with axtChain.\ \
Chains for low-coverage assemblies for which no browser has been built \ are not available as browser tracks, but only from our\ downloads page.\
\ \ See also: lastz parameters and other details (e.g., update time) \ and chain minimum score and gap parameters used in these alignments.\ \\ Chains were derived from lastz alignments, using the methods\ described on the chain tracks description pages, and sorted with the \ highest-scoring chains in the genome ranked first. The program\ chainNet was then used to place the chains one at a time, trimming them as \ necessary to fit into sections not already covered by a higher-scoring chain. \ During this process, a natural hierarchy emerged in which a chain that filled \ a gap in a higher-scoring chain was placed underneath that chain. The program \ netSyntenic was used to fill in information about the relationship between \ higher- and lower-level chains, such as whether a lower-level\ chain was syntenic or inverted relative to the higher-level chain. \ The program netClass was then used to fill in how much of the gaps and chains \ contained Ns (sequencing gaps) in one or both species and how much\ was filled with transposons inserted before and after the two organisms \ diverged.
\ \\ Harris, R.S. (2007) Improved pairwise alignment of genomic DNA. Ph.D. Thesis, \ The Pennsylvania State University.
\\ Lineage-specific repeats were identified by Arian Smit and his \ RepeatMasker\ program.
\\ The axtChain program was developed at the University of California at \ Santa Cruz by Jim Kent with advice from Webb Miller and David Haussler.
\\ The browser display and database storage of the chains and nets were created\ by Robert Baertsch and Jim Kent.
\\ The chainNet, netSyntenic, and netClass programs were\ developed at the University of California\ Santa Cruz by Jim Kent.
\\ \
\ Chiaromonte F, Yap VB, Miller W.\ Scoring pairwise genomic sequence alignments.\ Pac Symp Biocomput. 2002:115-26.\ PMID: 11928468\
\ \\ Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D.\ Evolution's cauldron:\ duplication, deletion, and rearrangement in the mouse and human genomes.\ Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9.\ PMID: 14500911; PMC: PMC208784\
\ compGeno 1 altColor 255,255,0\ chainLinearGap loose\ chainMinScore 5000\ color 0,0,0\ compositeTrack on\ configurable on\ dimensions dimensionX=clade dimensionY=species\ dragAndDrop subTracks\ group compGeno\ html primateChainNet\ longLabel Primate Genomes, Chain and Net Alignments\ noInherit on\ priority 7\ shortLabel Primate Chain/Net\ sortOrder species=+ view=+ clade=+\ subGroup1 view Views chain=Chains net=Nets\ subGroup2 species Species s000a=Human s000b=Hg38P2 s001=Human s002=J._Craig_Venter s002a=HG01243v3 s0025=Chimp s003=Chimp s004=Chimp s005=Chimp s006=Chimp s007a=Bonobo s007b=Bonobo s008=Bonobo s009a=Gorilla s009b=Gorilla s010=Gorilla s011=Gorilla s012=Gorilla s013a=Orangutan s013b=Orangutan s014=Gibbon s015=Gibbon s016=Proboscis_monkey s017=Black_snub-nosed_monkey s018=Golden_snub-nosed_monkey s019=Angolan_colobus s020=Crab-eating_macaque s021=Rhesus s022=Rhesus s023a=Rhesus s023b=Rhesus s024=Baboon s025=Baboon s026=Baboon s027=Pig-tailed_macaque s028=Sooty_mangabey s029=Green_monkey s030=Green_monkey s031=Drill s032=Squirrel_monkey s033=Ma's_night_monkey s034a=Marmoset s034b=Marmoset s035=Marmoset s036=White-faced_sapajou s037=Tarsier s038=Tarsier s039=Sclater's_lemur s040=Black_lemur s041=Coquerel's_sifaka s042=Mouse_lemur s043=Mouse_lemur s044=Mouse_lemur s045=Bushbaby s046=Bushbaby\ subGroup3 clade Clade c00=hominidae c01=cercopithecinae c02=haplorrhini c03=strepsirrhini\ track primateChainNet\ type bed 3\ visibility hide\ gnomad330XPercentage Sample % > 30X bigWig gnomAD Percentage of Genome Samples with at least 30X Coverage v3.0.1 2 7 75 0 180 165 127 217 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_30.bw\ color 75,0,180\ longLabel gnomAD Percentage of Genome Samples with at least 30X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 7\ shortLabel Sample % > 30X\ track gnomad330XPercentage\ viewLimits 0:1\ gnomad4Exome30XPercentage Sample % > 30X bigWig gnomAD Percentage of Exome Samples with at least 30X Coverage v4.0 2 7 75 0 180 165 127 217 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_30.bw\ color 75,0,180\ longLabel gnomAD Percentage of Exome Samples with at least 30X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 7\ shortLabel Sample % > 30X\ track gnomad4Exome30XPercentage\ viewLimits 0:1\ SeqCap-EZ_MedExome_hg38_capture_targets SeqCap EZ Med P bigBed Roche - SeqCap EZ MedExome Capture Probe Footprint 0 7 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/SeqCap_EZ_MedExome_hg38_capture_targets.bb\ color 100,143,255\ longLabel Roche - SeqCap EZ MedExome Capture Probe Footprint\ parent exomeProbesets off\ shortLabel SeqCap EZ Med P\ track SeqCap-EZ_MedExome_hg38_capture_targets\ type bigBed\ simpleRepeat Simple Repeats bed 4 + Simple Tandem Repeats by TRF 0 7 0 0 0 127 127 127 0 0 0\ This track displays simple tandem repeats (possibly imperfect repeats) located\ by Tandem Repeats\ Finder (TRF) which is specialized for this purpose. These repeats can\ occur within coding regions of genes and may be quite\ polymorphic. Repeat expansions are sometimes associated with specific\ diseases.
\ \\ For more information about the TRF program, see Benson (1999).\
\ \\ TRF was written by \ Gary Benson.
\ \\ Benson G.\ \ Tandem repeats finder: a program to analyze DNA sequences.\ Nucleic Acids Res. 1999 Jan 15;27(2):573-80.\ PMID: 9862982; PMC: PMC148217\
\ rep 1 group rep\ longLabel Simple Tandem Repeats by TRF\ priority 7\ shortLabel Simple Repeats\ track simpleRepeat\ type bed 4 +\ visibility hide\ ucneParalogs UCNE Paralogs bigBed 4 + UCNEBase: 987 Paralogous elements 0 7 0 0 0 127 127 127 0 0 0 url="https://epd.expasy.org/ucnebase/view.php?data=ucne&entry=$$" compGeno 1 bigDataUrl /gbdb/hg38/unusualcons/paralogs.bb\ longLabel UCNEBase: 987 Paralogous elements\ parent unusualcons on\ shortLabel UCNE Paralogs\ track ucneParalogs\ type bigBed 4 +\ url url="https://epd.expasy.org/ucnebase/view.php?data=ucne&entry=$$"\ refGene UCSC RefSeq genePred refPep refMrna UCSC annotations of RefSeq RNAs (NM_* and NR_*) 1 7 12 12 120 133 133 187 0 0 0\ The RefSeq Genes track shows known human protein-coding and\ non-protein-coding genes taken from the NCBI RNA reference sequences\ collection (RefSeq). The data underlying this track are updated weekly.
\ \\ Please visit the Feedback for Gene and Reference Sequences (RefSeq) page to\ make suggestions, submit additions and corrections, or ask for help concerning\ RefSeq records.\
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ This track follows the display conventions for\ \ gene prediction tracks.\ The color shading indicates the level of review the RefSeq record has\ undergone: predicted (light), provisional (medium), reviewed (dark).\
\ \\ The item labels and display colors of features within this track can be\ configured through the controls at the top of the track description page.\
\ RefSeq RNAs were aligned against the human genome using BLAT. Those\ with an alignment of less than 15% were discarded. When a single RNA\ aligned in multiple places, the alignment having the highest base identity\ was identified. Only alignments having a base identity level within 0.1% of\ the best and at least 96% base identity with the genomic sequence were kept.\
\ \\ This track was produced at UCSC from RNA sequence data generated by scientists\ worldwide and curated by the NCBI\ RefSeq project.\
\ \\ Kent WJ.\ \ BLAT - the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ \\ Pruitt KD, Brown GR, Hiatt SM, Thibaud-Nissen F, Astashyn A, Ermolaeva O, Farrell CM, Hart J,\ Landrum MJ, McGarvey KM et al.\ \ RefSeq: an update on mammalian reference sequences.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D756-63.\ PMID: 24259432; PMC: PMC3965018\
\ \\ Pruitt KD, Tatusova T, Maglott DR.\ \ NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.\ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4.\ PMID: 15608248; PMC: PMC539979\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 12,12,120\ group genes\ idXref hgFixed.refLink mrnaAcc name\ longLabel UCSC annotations of RefSeq RNAs (NM_* and NR_*)\ parent refSeqComposite off\ priority 7\ shortLabel UCSC RefSeq\ track refGene\ type genePred refPep refMrna\ visibility dense\ umap50Quantitative Umap M50 bigWig 0.02 1.0 Multi-read mappability with 50-mers 0 7 80 120 240 167 187 247 0 0 0 map 0 bigDataUrl /gbdb/hg38/hoffmanMappability/k50.Umap.MultiTrackMappability.bw\ color 80,120,240\ longLabel Multi-read mappability with 50-mers\ parent umapBigWig off\ priority 7\ shortLabel Umap M50\ subGroups view=MR\ track umap50Quantitative\ type bigWig 0.02 1.0\ visibility hide\ tgpNA19240_Y117_YRI Y117 YRI Trio vcfPhasedTrio 1000 Genomes Yoruban in Ibadan, Nigeria Trio 2 7 0 0 0 127 127 127 0 0 23 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX, varRep 0 longLabel 1000 Genomes Yoruban in Ibadan, Nigeria Trio\ parent tgpTrios\ shortLabel Y117 YRI Trio\ track tgpNA19240_Y117_YRI\ type vcfPhasedTrio\ vcfChildSample NA19240|child\ vcfParentSamples NA19238|mother,NA19239|father\ visibility full\ netAquChr2 aquChr2 Net netAlign aquChr2 chainAquChr2 Golden eagle (Oct. 2014 (aquChr-1.0.2/aquChr2)) Alignment Net 1 8 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Golden eagle (Oct. 2014 (aquChr-1.0.2/aquChr2)) Alignment Net\ otherDb aquChr2\ parent vertebrateChainNetViewnet off\ shortLabel aquChr2 Net\ subGroups view=net species=s016 clade=c01\ track netAquChr2\ type netAlign aquChr2 chainAquChr2\ netMm10 Mouse Net netAlign mm10 chainMm10 Mouse (Dec. 2011 (GRCm38/mm10)) Alignment Net 1 8 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Mouse (Dec. 2011 (GRCm38/mm10)) Alignment Net\ otherDb mm10\ parent placentalChainNetViewnet on\ shortLabel Mouse Net\ subGroups view=net species=s012a clade=c00\ track netMm10\ type netAlign mm10 chainMm10\ netPonAbe3 Orangutan Net netAlign ponAbe3 chainPonAbe3 Orangutan (Jan. 2018 (Susie_PABv2/ponAbe3)) Alignment Net 1 8 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Orangutan (Jan. 2018 (Susie_PABv2/ponAbe3)) Alignment Net\ otherDb ponAbe3\ parent primateChainNetViewnet off\ shortLabel Orangutan Net\ subGroups view=net species=s013a clade=c00\ track netPonAbe3\ type netAlign ponAbe3 chainPonAbe3\ encTfChipPkENCFF186ZET A549 CREB1 2 narrowPeak Transcription Factor ChIP-seq Peaks of CREB1 in A549 from ENCODE 3 (ENCFF186ZET) 0 8 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of CREB1 in A549 from ENCODE 3 (ENCFF186ZET)\ parent encTfChipPk off\ shortLabel A549 CREB1 2\ subGroups cellType=A549 factor=CREB1\ track encTfChipPkENCFF186ZET\ cloneEndABC20 ABC20 bed 12 Agencourt fosmid library 20 0 8 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 20\ parent cloneEndSuper off\ priority 8\ shortLabel ABC20\ subGroups source=agencourt\ track cloneEndABC20\ type bed 12\ visibility hide\ AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep2LK5_CNhs13359_ctss_rev AorticSmsToFgf2_00hr15minBr2- bigWig Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep2 (LK5)_CNhs13359_12741-135I5_reverse 0 8 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12741-135I5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr15min%2c%20biol_rep2%20%28LK5%29.CNhs13359.12741-135I5.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep2 (LK5)_CNhs13359_12741-135I5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12741-135I5 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr15minBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep2LK5_CNhs13359_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12741-135I5\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep2LK5_CNhs13359_tpm_rev AorticSmsToFgf2_00hr15minBr2- bigWig Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep2 (LK5)_CNhs13359_12741-135I5_reverse 1 8 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12741-135I5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr15min%2c%20biol_rep2%20%28LK5%29.CNhs13359.12741-135I5.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep2 (LK5)_CNhs13359_12741-135I5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12741-135I5 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr15minBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep2LK5_CNhs13359_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12741-135I5\ urlLabel FANTOM5 Details:\ wgEncodeReg4TxnBloodVesselMinus Blood vessel - bigWig Avg. - strand total RNA-seq level of 21 blood vessel experiments (tissues and primary cells only) 0 8 255 37 41 255 146 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodVesselMinus.bw\ color 255,37,41\ longLabel Avg. - strand total RNA-seq level of 21 blood vessel experiments (tissues and primary cells only)\ negateValues on\ parent wgEncodeReg4Txn off\ priority 8\ shortLabel Blood vessel -\ track wgEncodeReg4TxnBloodVesselMinus\ type bigWig\ gtexCovBrainAmygdala Brain Amygd bigWig Brain Amygdala 0 8 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-T5JC-0011-R4A-SM-32PLT.Brain_Amygdala.RNAseq.bw\ color 238,238,0\ longLabel Brain Amygdala\ parent gtexCov\ shortLabel Brain Amygd\ track gtexCovBrainAmygdala\ placentalChainNetViewchain Chains bed 3 Non-primate Placental Mammal Genomes, Chain and Net Alignments 3 8 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Non-primate Placental Mammal Genomes, Chain and Net Alignments\ parent placentalChainNet\ shortLabel Chains\ spectrum on\ track placentalChainNetViewchain\ view chain\ visibility pack\ phastCons470way Cons 470 Mammals bigWig 0 1 470 mammals conservation by PhastCons 0 8 70 130 70 130 70 70 0 0 0 compGeno 0 altColor 130,70,70\ autoScale off\ bigDataUrl https://hgdownload.soe.ucsc.edu/goldenPath/hg38/phastCons470way/hg38.phastCons470way.bw\ color 70,130,70\ configurable on\ longLabel 470 mammals conservation by PhastCons\ maxHeightPixels 100:40:11\ noInherit on\ parent cons470wayViewphastcons off\ priority 8\ shortLabel Cons 470 Mammals\ spanList 1\ subGroups view=phastcons\ track phastCons470way\ type bigWig 0 1\ windowingFunction mean\ cortexNeuron42K Cortex - Neuron - Z0000042K bigWig Methylation Atlas: Cortex - Neuron - Z0000042K 2 8 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/cortexNeuron42K.bw\ color 138,43,226\ longLabel Methylation Atlas: Cortex - Neuron - Z0000042K\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 8\ shortLabel Cortex - Neuron - Z0000042K\ subGroups cellType=Neuron dataType=Replicate\ track cortexNeuron42K\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ unipDisulfBond Disulf. Bonds bigBed 12 + UniProt Disulfide Bonds 1 8 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipDisulfBond.bb\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ longLabel UniProt Disulfide Bonds\ mouseOver UniProt record: $uniProtId\ FinnGen is a public-private partnership\ that combines genotype data from Finnish biobanks with digital health record data from Finnish\ health registries. The R12 release contains imputed variants from 500,348 biobank samples typed on\ genotyping arrays. The imputation used phased variants from 8,554 high-quality\ whole genome sequences, also from Finland. That is roughly 10% of the Finnish\ population. Phenotype links can be viewed at the\ FinnGen PheWeb.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ TSV data can be requested via the form at\ FinnGen,\ which triggers an automated email containing the download link.\ A script in our GitHub repo converts this file to VCF (see Methods below).\
\ \\ FinnGen participants were genotyped using a custom Axiom FinnGen1 array, supplemented by legacy\ collections genotyped with other arrays. Imputation used a population-specific reference panel of\ high-coverage (25–30x) whole-genome sequences from Finnish individuals. Ancestry outliers were\ removed via PCA against 1000 Genomes reference samples, and 5,780 duplicates and monozygotic twins\ were excluded. Variant quality was assessed using VQSR.\
\\ R12 annotated variants were downloaded from the Google Cloud bucket link received through an email\ and converted to VCF with a\ custom Python script.\ The conversion steps for all source files of the varFreqs track are recorded in the makeDoc file of the track.\ Some tracks also need python scripts, which live on GitHub.\
\ \\ Thanks to the participants and investigators of the FinnGen study.\
\ \\ Kurki MI, Karjalainen J, Palta P, Sipilä TP, Kristiansson K, Donner KM, Reeve MP, Laivuori H,\ Aavikko M, Kaunisto MA et al.\ \ FinnGen provides genetic insights from a well-phenotyped isolated population.\ Nature. 2023 Jan;613(7944):508-518.\ PMID: 36653562; PMC: PMC9849126\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_finngen/finnge_R12_annotated_variants_v1.vcf.gz\ dataVersion R12\ longLabel SNV Frequencies: Finland FinnGen - 500k samples, arrays, imputation used 8.5k WGS\ parent varFreqs on\ priority 8\ shortLabel FinnGen R12 500k imputed\ tableBrowser off\ track finngen\ type vcfTabix\ visibility hide\ knownGeneV36 GENCODE V36 bigGenePred GENCODE V36 0 8 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 36, Oct 2020) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ By default, only the basic gene set is\ displayed, which is a subset of the comprehensive gene set. The basic set represents transcripts\ that GENCODE believes will be useful to the majority of users.
\ \\ The track includes protein-coding genes, non-coding RNA genes, and pseudo-genes, though pseudo-genes\ are not displayed by default. It contains annotations on the reference chromosomes as well as\ assembly patches and alternative loci (haplotypes).
\ \\ The following table provides statistics for the v36 release derived from the GTF file that contains\ annotations only on the main chromosomes. More information on how they were generated can be found\ in the GENCODE site.
\ \\
\ \\
\ GENCODE v36 Release Stats \ Genes Observed Transcripts Observed \ Protein-coding genes 19,965 Protein-coding transcripts 83,986 \ Long non-coding RNA genes 17,910 - full length protein-coding 57,935 \ Small non-coding RNA genes 7,576 - partial length protein-coding 26,051 \ Pseudogenes 14,749 Nonsense mediated decay transcripts 15,811 \ Immunoglobulin/T-cell receptor gene segments 645 Long non-coding RNA loci transcripts 48,351
\
\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ By default, this track displays only the basic GENCODE set, splice variants, and non-coding genes.\ It includes options to display the entire GENCODE set and pseudogenes. To customize these\ options, the respective boxes can be checked or unchecked at the top of this description page. \ \
\ This track also includes a variety of labels which identify the transcripts when visibility is set\ to "full" or "pack". Gene symbols (e.g. NIPA1) are displayed by default, but\ additional options include GENCODE Transcript ID (ENST00000561183.5), UCSC Known Gene ID\ (uc001yve.4), UniProt Display ID (Q7RTP0). Additional information about gene\ and transcript names can be found in our\ FAQ.
\ \\ This track, in general, follows the display conventions for gene prediction tracks. The exons for\ putative non-coding genes and untranslated regions are represented by relatively thin blocks, while\ those for coding open reading frames are thicker. \
Coloring for the gene annotations is based on the annotation type:
\\ This track contains an optional codon coloring feature that allows users to\ quickly validate and compare gene predictions. There is also an option to display the data as\ a density graph, which\ can be helpful for visualizing the distribution of items over a region.
\ \\
The GENCODE v36 track was built from the GENCODE downloads file \
gencode.v36.chr_patch_hapl_scaff.annotation.gff3.gz. Data from other sources \
were correlated with the GENCODE data to build association tables.
\ The GENCODE Genes transcripts are annotated in numerous tables, each of which is also available as a\ downloadable\ file.\ \
\ One can see a full list of the associated tables in the Table Browser by selecting GENCODE Genes from the track menu; this list\ is then available on the table menu.\ \ \
\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator. \ The genePred format files for hg38 are available from our \ \ downloads directory or in our\ \ GTF download directory. \ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\ \\ The GENCODE Genes track was produced at UCSC from the GENCODE comprehensive gene set using a\ computational pipeline developed by Jim Kent and Brian Raney.
\ \\ Harrow J, Frankish A, Gonzalez JM, Tapanari E, Diekhans M, Kokocinski F, Aken BL, Barrell D, Zadissa\ A, Searle S et al.\ \ GENCODE: the reference human genome annotation for The ENCODE Project.\ Genome Res. 2012 Sep;22(9):1760-74.\ PMID: 22955987; PMC: PMC3431492\
\ \\ Harrow J, Denoeud F, Frankish A, Reymond A, Chen CK, Chrast J, Lagarde J, Gilbert JG, Storey R,\ Swarbreck D et al.\ \ GENCODE: producing a reference annotation for ENCODE.\ Genome Biol. 2006;7 Suppl 1:S4.1-9.\ PMID: 16925838; PMC: PMC1810553\
\ \A full list of GENCODE publications is available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/gencode/gencodeV36.bb\ defaultLabelFields geneName\ defaultLinkedTables kgXref\ directUrl /cgi-bin/hgGene?hgg_gene=%s&hgg_chrom=%s&hgg_start=%d&hgg_end=%d&hgg_type=%s&db=%s\ externalDb knownGeneV36\ group genes\ html knownGeneV36\ idXref kgAlias kgID alias\ intronGap 12\ isGencode3 on\ itemRgb on\ labelFields geneName,name,geneName2,name2\ longLabel GENCODE V36\ maxItems 50000\ parent knownGeneArchive\ priority 8\ searchIndex name\ shortLabel GENCODE V36\ track knownGeneV36\ type bigGenePred\ visibility hide\ geneHancerClusteredInteractions GH Clusters bigInteract Clustered interactions of GeneHancer regulatory elements and genes 3 8 0 0 0 127 127 127 0 0 0 https://www.genecards.org/cgi-bin/carddisp.pl?gene=$\ The chain track shows alignments of human (Dec. 2013 (GRCh38/hg38)) to\ other genomes using a gap scoring system that allows longer gaps \ than traditional affine gap scoring systems. It can also tolerate gaps in both\ human and the other genome simultaneously. These \ "double-sided" gaps can be caused by local inversions and \ overlapping deletions in both species. \
\ The chain track displays boxes joined together by either single or\ double lines. The boxes represent aligning regions.\ Single lines indicate gaps that are largely due to a deletion in the\ other assembly or an insertion in the human assembly.\ Double lines represent more complex gaps that involve substantial\ sequence in both species. This may result from inversions, overlapping\ deletions, an abundance of local mutation, or an unsequenced gap in one\ species. In cases where multiple chains align over a particular region of\ the other genome, the chains with single-lined gaps are often \ due to processed pseudogenes, while chains with double-lined gaps are more \ often due to paralogs and unprocessed pseudogenes.
\\ In the "pack" and "full" display\ modes, the individual feature names indicate the chromosome, strand, and\ location (in thousands) of the match for each matching alignment.
\ \\ The net track shows the best human/other chain for \ every part of the other genome. It is useful for\ finding orthologous regions and for studying genome\ rearrangement. The human sequence used in this annotation is from\ the Dec. 2013 (GRCh38/hg38) assembly.
\ \By default, the chains to chromosome-based assemblies are colored\ based on which chromosome they map to in the aligning organism. To turn\ off the coloring, check the "off" button next to: Color\ track based on chromosome.
\\ To display only the chains of one chromosome in the aligning\ organism, enter the name of that chromosome (e.g. chr4) in box next to: \ Filter by chromosome.
\ \\ In full display mode, the top-level (level 1)\ chains are the largest, highest-scoring chains that\ span this region. In many cases gaps exist in the\ top-level chain. When possible, these are filled in by\ other chains that are displayed at level 2. The gaps in \ level 2 chains may be filled by level 3 chains and so\ forth.
\\ In the graphical display, the boxes represent ungapped \ alignments; the lines represent gaps. Click\ on a box to view detailed information about the chain\ as a whole; click on a line to display information\ about the gap. The detailed information is useful in determining\ the cause of the gap or, for lower level chains, the genomic\ rearrangement.
\\ Individual items in the display are categorized as one of four types\ (other than gap):
\\
Transposons that have been inserted since the human/other\
split were removed from the assemblies. The abbreviated genomes were\
aligned with lastz, and the transposons were added back in.\
The resulting alignments were converted into axt format using the lavToAxt\
program. The axt alignments were fed into axtChain, which organizes all\
alignments between a single human chromosome and a single\
chromosome from the other genome into a group and creates a kd-tree out\
of the gapless subsections (blocks) of the alignments. A dynamic program\
was then run over the kd-trees to find the maximally scoring chains of these\
blocks.\
\
\
\
Chains scoring below a minimum score of '5000' were discarded;\
the remaining chains are displayed in this track. The linear gap\
matrix used with axtChain:
\
-linearGap=loose\ \ tablesize 11\ smallSize 111\ position 1 2 3 11 111 2111 12111 32111 72111 152111 252111\ qGap 325 360 400 450 600 1100 3600 7600 15600 31600 56600\ tGap 325 360 400 450 600 1100 3600 7600 15600 31600 56600\ bothGap 625 660 700 750 900 1400 4000 8000 16000 32000 57000\\ \ See also: lastz parameters used in these alignments,\ and chain minimum score and gap parameters used in these alignments.\ \ \
\ Chains were derived from lastz alignments, using the methods\ described on the chain tracks description pages, and sorted with the \ highest-scoring chains in the genome ranked first. The program\ chainNet was then used to place the chains one at a time, trimming them as \ necessary to fit into sections not already covered by a higher-scoring chain. \ During this process, a natural hierarchy emerged in which a chain that filled \ a gap in a higher-scoring chain was placed underneath that chain. The program \ netSyntenic was used to fill in information about the relationship between \ higher- and lower-level chains, such as whether a lower-level\ chain was syntenic or inverted relative to the higher-level chain. \ The program netClass was then used to fill in how much of the gaps and chains \ contained Ns (sequencing gaps) in one or both species and how much\ was filled with transposons inserted before and after the two organisms \ diverged.
\ \\ LASTZ was developed at\ Miller Lab at Pennsylvania State University by \ Bob Harris.\
\\ Lineage-specific repeats were identified by Arian Smit and his \ RepeatMasker\ program.
\\ The axtChain program was developed at the University of California at \ Santa Cruz by Jim Kent with advice from Webb Miller and David Haussler.
\\ The browser display and database storage of the chains and nets were created\ by Robert Baertsch and Jim Kent.
\\ The chainNet, netSyntenic, and netClass programs were\ developed at the University of California\ Santa Cruz by Jim Kent.
\\ \
\ Harris RS.\ Improved pairwise alignment of genomic DNA.\ Ph.D. Thesis. Pennsylvania State University, USA. 2007.\
\ \\ Chiaromonte F, Yap VB, Miller W.\ Scoring pairwise genomic sequence alignments.\ Pac Symp Biocomput. 2002:115-26.\ PMID: 11928468\
\ \\ Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D.\ Evolution's cauldron:\ duplication, deletion, and rearrangement in the mouse and human genomes.\ Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9.\ PMID: 14500911; PMC: PMC208784\
\ \\ Schwartz S, Kent WJ, Smit A, Zhang Z, Baertsch R, Hardison RC,\ Haussler D, Miller W.\ Human-mouse alignments with BLASTZ.\ Genome Res. 2003 Jan;13(1):103-7.\ PMID: 12529312; PMC: PMC430961\
\ compGeno 1 altColor 255,255,0\ chainLinearGap loose\ chainMinScore 5000\ color 0,0,0\ compositeTrack on\ configurable on\ dimensions dimensionX=clade dimensionY=species\ dragAndDrop subTracks\ group compGeno\ html placentalChainNet\ longLabel Non-primate Placental Mammal Genomes, Chain and Net Alignments\ noInherit on\ priority 8\ shortLabel Placental Chain/Net\ sortOrder species=+ view=+ clade=+\ subGroup1 view Views chain=Chains net=Nets\ subGroup2 species Species s000=Guinea_pig s001=Guinea_pig s002=Chinchilla s003=Chinese_hamster s004a=Chinese_hamster_CHOv1 s004b=Chinese_hamster_CHOv2 s004c=RegenCHO1 s005a=Kangaroo_rat s005b=Beaver s006=Malayan_flying_lemur s007=Naked_mole-rat s008=Naked_mole-rat s0081=Damara_mole-rat s009=Lesser_Egyptian_jerboa s010=Golden_hamster s011=Prairie_vole s012a=Mouse s012b=Mouse38B s012c=Mouse s013=Mouse s014=Mouse s015=Mouse s016=Mouse s017=Upper_Galilee_mountains_blind_mole_rat s018=Pika s019=Pika s020=Brush-tailed_rat s021=Rabbit s022=Rabbit s023=Prairie_deer_mouse s024a=Rat s024b=RegenRn1 s024c=RegenRn0 s024d=Rat s025=Rat s026=Rat s027=Rat s028=Rat s029=Squirrel s030=Squirrel s031=Tree_shrew s032=Chinese_tree_shrew s033=Panda s034a=Dog s034b=Dog s034c=Dog s034d=Dog s035=Dog s036=Dog s037a=Dog s037b=Domestic_cat s037c=Cat s038=Cat s039=Cat s040=Cat s041=Cat s042=Weddell_seal s043=Ferret s043a=Southern_sea_otter s044=Hawaiian_monk_seal s045=Pacific_walrus s046=Amur_tiger s047=Polar_bear s048=Minke_whale s049=Bison s050=Wild_yak s051a=Cow s051b=Cow s052=Cow s053=Cow s054=Cow s055=Cow s056=Cow s057=Cow s058=Cow s059=Cow s060=Water_buffalo s061=Bactrian_camel s062=Domestic_goat s063=Yangtze_river_dolphin s064=Killer_whale s065a=Sheep s065b=Sheep s065c=Sheep s065d=Sheep s067=Tibetan_antelope s068=Sperm_whale s069=Pig s070=Pig s071=Pig s072=Pig s073=Dolphin s074=Dolphin s075=Alpaca s076=Alpaca s077=Straw_colored_fruit_bat s078=Big_brown_bat s079=Indian_false_vampire s080=Brandt's_myotis_(bat) s081=David's_myotis_(bat) s082=Microbat s083=Microbat s084=Black_flying-fox s085=Parnell's_mustached_bat s086=Megabat s087=Greater_horseshoe_bat s088=Egyptian_rousette s089=Star-nosed_mole s090=Hedgehog s091=Hedgehog s092=Chinese_pangolin s093=Shrew s094=Shrew s095=White_rhinoceros s096a=Horse s096b=Horse s096c=Horse s098=Przewalski_horse s099=Sloth s100=Armadillo s101=Armadillo s102=Armadillo s103=Cape_golden_mole s104=Tenrec s105=Tenrec s106=Cape_elephant_shrew s107=Elephant s108=Elephant s109a=Asiatic_elephant s109b=Elephant s110=Aardvark s111=Rock_hyrax s112=Manatee\ subGroup3 clade Clade c00=Euarchontoglires c01=Carnivora c02=Cetartiodactyla c03=Chiroptera c04=Laurasiatheria c05=Perissodactyla c06=Xenarthra c07=Afrotheria\ track placentalChainNet\ type bed 3\ visibility hide\ wgEncodeReg4AtacProstate Prostate bigWig ATAC level of 1 prostate experiment (tissues and primary cells only) 0 8 140 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpProstateATAC.bw\ color 140,140,140\ longLabel ATAC level of 1 prostate experiment (tissues and primary cells only)\ parent wgEncodeReg4Atac off\ priority 8\ shortLabel Prostate\ track wgEncodeReg4AtacProstate\ type bigWig\ ncbiRefSeqSelect RefSeq Select and MANE genePred NCBI RefSeq Select and MANE subset: A single representative transcript 1 8 20 20 160 137 137 207 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 20,20,160\ idXref ncbiRefSeqLink mrnaAcc name\ longLabel NCBI RefSeq Select and MANE subset: A single representative transcript\ parent refSeqComposite off\ priority 8\ shortLabel RefSeq Select and MANE\ track ncbiRefSeqSelect\ trackHandler ncbiRefSeq\ type genePred\ gnomad350XPercentage Sample % > 50X bigWig gnomAD Percentage of Genome Samples with at least 50X Coverage v3.0.1 2 8 45 0 210 150 127 232 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_50.bw\ color 45,0,210\ longLabel gnomAD Percentage of Genome Samples with at least 50X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 8\ shortLabel Sample % > 50X\ track gnomad350XPercentage\ viewLimits 0:1\ gnomad4Exome50XPercentage Sample % > 50X bigWig gnomAD Percentage of Exome Samples with at least 50X Coverage v4.0 2 8 45 0 210 150 127 232 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_50.bw\ color 45,0,210\ longLabel gnomAD Percentage of Exome Samples with at least 50X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 8\ shortLabel Sample % > 50X\ track gnomad4Exome50XPercentage\ viewLimits 0:1\ SeqCap-EZ_MedExome_hg19_empirical_targets SeqCap EZ Med T bigBed Roche - SeqCap EZ MedExome Empirical Target Regions 0 8 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/SeqCap_EZ_MedExome_hg38_empirical_targets.bb\ color 100,143,255\ longLabel Roche - SeqCap EZ MedExome Empirical Target Regions\ parent exomeProbesets off\ shortLabel SeqCap EZ Med T\ track SeqCap-EZ_MedExome_hg19_empirical_targets\ type bigBed\ ultras Ultracons bigBed 4 Ultracons: 481 Ultraconserved regions - 100% identical in human, mouse and rat, > 200bp 0 8 0 0 0 127 127 127 0 0 0 compGeno 1 bigDataUrl /gbdb/hg38/unusualcons/hg38.ultraConserved.bb\ longLabel Ultracons: 481 Ultraconserved regions - 100% identical in human, mouse and rat, > 200bp\ parent unusualcons on\ shortLabel Ultracons\ track ultras\ type bigBed 4\ umap100Quantitative Umap M100 bigWig 0.01 1.0 Multi-read mappability with 100-mers 0 8 80 170 240 167 212 247 0 0 0 map 0 bigDataUrl /gbdb/hg38/hoffmanMappability/k100.Umap.MultiTrackMappability.bw\ color 80,170,240\ longLabel Multi-read mappability with 100-mers\ parent umapBigWig off\ priority 8\ shortLabel Umap M100\ subGroups view=MR\ track umap100Quantitative\ type bigWig 0.01 1.0\ visibility hide\ windowmaskerSdust WM + SDust bed 3 Genomic Intervals Masked by WindowMasker + SDust 0 8 0 0 0 127 127 127 0 0 0\ This track depicts masked sequence as determined by\ WindowMasker. The\ WindowMasker tool is included in the NCBI C++ toolkit. The source code\ for the entire toolkit is available from the NCBI\ \ FTP site.\
\ \\ To create this track, WindowMasker was run with the following parameters:\
\ windowmasker -mk_counts true -input hg38.fa -output wm_counts\ windowmasker -ustat wm_counts -sdust true -input hg38.fa -output repeats.bed\\ The repeats.bed (BED3) file was loaded into the "windowmaskerSdust" table for\ this track.\ \ \
\ Morgulis A, Gertz EM, Schäffer AA, Agarwala R.\ WindowMasker: window-based masker for sequenced genomes.\ Bioinformatics. 2006 Jan 15;22(2):134-41.\ PMID: 16287941\
\ rep 1 group rep\ longLabel Genomic Intervals Masked by WindowMasker + SDust\ priority 8\ shortLabel WM + SDust\ track windowmaskerSdust\ type bed 3\ visibility hide\ chainThaSir1 thaSir1 Chain chain thaSir1 Garter snake (Jun. 2015 (Thamnophis_sirtalis-6.0/thaSir1)) Chained Alignments 3 9 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Garter snake (Jun. 2015 (Thamnophis_sirtalis-6.0/thaSir1)) Chained Alignments\ otherDb thaSir1\ parent vertebrateChainNetViewchain off\ shortLabel thaSir1 Chain\ subGroups view=chain species=s028b clade=c02\ track chainThaSir1\ type chain thaSir1\ chainRn7 Rat Chain chain rn7 Rat (Nov. 2020 (mRatBN7.2/rn7)) Chained Alignments 3 9 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Rat (Nov. 2020 (mRatBN7.2/rn7)) Chained Alignments\ otherDb rn7\ parent placentalChainNetViewchain off\ shortLabel Rat Chain\ subGroups view=chain species=s024a clade=c00\ track chainRn7\ type chain rn7\ chainNomLeu3 Gibbon Chain chain nomLeu3 Gibbon (Oct. 2012 (GGSC Nleu3.0/nomLeu3)) Chained Alignments 3 9 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Gibbon (Oct. 2012 (GGSC Nleu3.0/nomLeu3)) Chained Alignments\ otherDb nomLeu3\ parent primateChainNetViewchain off\ shortLabel Gibbon Chain\ subGroups view=chain species=s014 clade=c00\ track chainNomLeu3\ type chain nomLeu3\ phastConsElements470way 470 Mamm. El bigBed 5 . 470 mammals Conserved Elements 0 9 110 10 40 182 132 147 0 0 0 compGeno 1 bigDataUrl https://hgdownload.soe.ucsc.edu/goldenPath/hg38/phastCons470way/hg38.phastConsElements470way.bb\ color 110,10,40\ longLabel 470 mammals Conserved Elements\ noInherit on\ parent cons470wayViewelements off\ priority 9\ shortLabel 470 Mamm. El\ subGroups view=elements\ track phastConsElements470way\ type bigBed 5 .\ encTfChipPkENCFF535MZG A549 CTCF 1 narrowPeak Transcription Factor ChIP-seq Peaks of CTCF in A549 from ENCODE 3 (ENCFF535MZG) 0 9 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of CTCF in A549 from ENCODE 3 (ENCFF535MZG)\ parent encTfChipPk off\ shortLabel A549 CTCF 1\ subGroups cellType=A549 factor=CTCF\ track encTfChipPkENCFF535MZG\ unipModif AA Modifications bigBed 12 + UniProt Amino Acid Modifications 1 9 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipModif.bb\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ longLabel UniProt Amino Acid Modifications\ mouseOver UniProt record: $uniProtId\ This track shows small variants (single-nucleotide variants and short\ insertion/deletion variants) identified by PacBio HiFi long-read sequencing\ of probands and their families enrolled in the Genomic Answers for Kids\ (GA4K) program at Children's Mercy Research Institute. GA4K is a longitudinal\ pediatric genomics initiative that aims to enroll 30,000 children with\ suspected rare genetic disorders, together with their parents, to build a\ large-scale resource of clinical and genomic data.\
\\ The callset contains approximately 36.2 million variants genotyped across\ up to 552 samples (maximum allele number 1104 on the autosomes). Each\ variant is annotated with allele count (AC), total called alleles (AN),\ cohort allele frequency (AF), variant type (substitution, insertion or\ deletion), and the corresponding gnomAD v3.0 allele frequency if one is\ available.\
\ \\ The track uses the standard VCF display. By default, variants appear as\ colored marks along the genome. Click an item to open its detail page,\ which lists the per-site INFO fields AC, AN, AF and the gnomAD v3 allele\ frequency.\
\ \\ Samples were sequenced on PacBio Revio and Sequel II instruments with HiFi\ chemistry. Per-sample variant calls were generated with DeepVariant as gVCFs,\ then merged across the cohort with GLnexus v1.2.7 using the\ DeepVariant_unfiltered configuration. The resulting BCF was converted\ to VCF with bcftools view v1.10.\
\\ To reduce false positives, the merged callset was filtered to variants\ replicated by independent evidence: (1) observed in at least one additional\ unrelated Children's Mercy individual, or (2) matching a variant observed in\ a sample from the Human Pangenome Reference Consortium (HPRC).\
\\ The GA4K release ships as 24 per-chromosome VCF files (chr1-22, chrX,\ chrY). For the Genome Browser, these were concatenated with\ bcftools concat into a single bgzip-compressed, tabix-indexed file.\
\ \\ The VCF file for this track is available from\ our\ download server as ga4kSnv.vcf.gz (with .tbi index).\ Regions can be extracted with tabix, for example:\ tabix http://hgdownload.soe.ucsc.edu/gbdb/hg38/varFreqs/ga4k/ga4kSnv.vcf.gz chr21:1-100000000.\
\\ The original per-chromosome VCFs and full release documentation are\ available from the Children's Mercy Research Institute GA4K data release at\ \ github.com/ChildrensMercyResearchInstitute/GA4K.\
\ \\ Thanks to the Children's Mercy Research Institute and the Genomic Answers\ for Kids participants and their families, who released this dataset to the\ public.\
\ \\ Cohen ASA, Farrow EG, Abdelmoity AT, Alaimo JT, Amudhavalli SM, Anderson JT, Bansal L, Bartik L,\ Baybayan P, Belden B et al.\ \ Genomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes.\ Genet Med. 2022 Jun;24(6):1336-1348.\ PMID: 35305867\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/ga4k/ga4kSnv.vcf.gz\ dataVersion Cohen 2022 release\ longLabel SNV Frequencies: GA4K Children's Mercy - 552 PacBio HiFi WGS, pediatric RD\ parent varFreqs on\ priority 9\ shortLabel GA4K 552 PacBio LR\ track ga4kSnv\ type vcfTabix\ visibility hide\ wgEncodeRegDnaseUwHffPeak HFF Pk narrowPeak HFF foreskin fibroblast DNaseI Peaks from ENCODE 1 9 255 163 85 255 209 170 1 0 0 regulation 1 color 255,163,85\ longLabel HFF foreskin fibroblast DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel HFF Pk\ subGroups view=a_Peaks cellType=HFF treatment=n_a tissue=skin cancer=normal\ track wgEncodeRegDnaseUwHffPeak\ wgEncodeRegDnaseUwHffWig HFF Sg bigWig 0 17635.9 HFF foreskin fibroblast DNaseI Signal from ENCODE 0 9 255 163 85 255 209 170 0 0 0 regulation 1 color 255,163,85\ longLabel HFF foreskin fibroblast DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.0818\ shortLabel HFF Sg\ subGroups cellType=HFF treatment=n_a tissue=skin cancer=normal\ table wgEncodeRegDnaseUwHffSignal\ track wgEncodeRegDnaseUwHffWig\ type bigWig 0 17635.9\ chainHprcGCA_018505825v1 HG02109.mat chain GCA_018505825.1 HG02109.mat HG02109.pri.mat.f1_v2 (May 2021 GCA_018505825.1_HG02109.pri.mat.f1_v2) HPRC project computed Chained Alignments 3 9 0 0 0 255 255 0 1 0 0 hprc 1 longLabel HG02109.mat HG02109.pri.mat.f1_v2 (May 2021 GCA_018505825.1_HG02109.pri.mat.f1_v2) HPRC project computed Chained Alignments\ otherDb GCA_018505825.1\ parent hprcChainNetViewchain off\ priority 25\ shortLabel HG02109.mat\ subGroups view=chain sample=s025 population=afr subpop=acb hap=mat\ track chainHprcGCA_018505825v1\ type chain GCA_018505825.1\ lincRNAsCThLF_r1 hLF_r1 bed 5 + lincRNAs from hlf_r1 1 9 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from hlf_r1\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel hLF_r1\ subGroups view=lincRNAsRefseqExp tissueType=hlf_r1\ track lincRNAsCThLF_r1\ wgEncodeReg4MarkH3k4me3Kidney Kidney bigWig Avg. H3K4me3 level of 5 kidney experiments (tissues and primary cells only) 0 9 92 161 153 173 208 204 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpKidneyH3K4me3.bw\ color 92,161,153\ longLabel Avg. H3K4me3 level of 5 kidney experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3\ priority 9\ shortLabel Kidney\ track wgEncodeReg4MarkH3k4me3Kidney\ type bigWig\ wgEncodeReg4MarkH3k27acLargeIntestine Large intestine bigWig Avg. H3K27ac level of 18 large intestine experiments (tissues and primary cells only) 2 9 86 86 36 170 170 145 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLargeIntestineH3K27ac.bw\ color 86,86,36\ longLabel Avg. H3K27ac level of 18 large intestine experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 9\ shortLabel Large intestine\ track wgEncodeReg4MarkH3k27acLargeIntestine\ type bigWig\ wgEncodeReg4MarkCtcfLiver Liver bigWig Avg. CTCF level of 2 liver experiments (tissues and primary cells only) 0 9 137 152 82 196 203 168 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLiverCTCF.bw\ color 137,152,82\ longLabel Avg. CTCF level of 2 liver experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf\ priority 9\ shortLabel Liver\ track wgEncodeReg4MarkCtcfLiver\ type bigWig\ vertebrateChainNetViewnet Nets bed 3 Non-placental Vertebrate Genomes, Chain and Net Alignments 1 9 0 0 0 255 255 0 0 0 0 compGeno 1 longLabel Non-placental Vertebrate Genomes, Chain and Net Alignments\ parent vertebrateChainNet\ shortLabel Nets\ track vertebrateChainNetViewnet\ view net\ visibility dense\ wgEncodeRegTxnCaltechRnaSeqNhlfR2x75Il200SigPooled NHLF bigWig 0 65535 Transcription of NHLF cells from ENCODE 0 9 255 128 212 255 191 233 0 0 0 regulation 1 color 255,128,212\ longLabel Transcription of NHLF cells from ENCODE\ origAssembly hg19\ parent wgEncodeRegTxn\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ priority 9\ shortLabel NHLF\ track wgEncodeRegTxnCaltechRnaSeqNhlfR2x75Il200SigPooled\ type bigWig 0 65535\ iscaPathGainCum Path Gain bedGraph 4 ClinGen CNVs: Pathogenic Gain Coverage 2 9 0 0 200 127 127 227 0 0 0 phenDis 0 color 0,0,200\ longLabel ClinGen CNVs: Pathogenic Gain Coverage\ parent iscaViewTotal\ shortLabel Path Gain\ subGroups view=cov class=path level=sub\ track iscaPathGainCum\ ncbiRefSeqHgmd RefSeq HGMD genePred NCBI RefSeq HGMD subset: transcripts with clinical variants in HGMD 1 9 20 20 160 137 137 207 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 20,20,160\ idXref ncbiRefSeqLink mrnaAcc name\ longLabel NCBI RefSeq HGMD subset: transcripts with clinical variants in HGMD\ parent refSeqComposite off\ priority 9\ shortLabel RefSeq HGMD\ track ncbiRefSeqHgmd\ trackHandler ncbiRefSeq\ type genePred\ ncbiRefSeqHistorical RefSeq Historical genePred NCBI RefSeq Historical Transcript Versions 1 9 12 12 120 133 133 187 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 12,12,120\ idXref ncbiRefSeqLinkHistorical mrnaAcc name\ longLabel NCBI RefSeq Historical Transcript Versions\ parent refSeqComposite off\ priority 9\ shortLabel RefSeq Historical\ track ncbiRefSeqHistorical\ type genePred\ gnomad3100XPercentage Sample % > 100X bigWig gnomAD Percentage of Genome Samples with at least 100X Coverage v3.0.1 2 9 15 0 240 135 127 247 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v3-genome/gnomad.coverage.over_100.bw\ color 15,0,240\ longLabel gnomAD Percentage of Genome Samples with at least 100X Coverage v3.0.1\ parent gnomad3Coverage off\ priority 9\ shortLabel Sample % > 100X\ track gnomad3100XPercentage\ viewLimits 0:1\ gnomad4Exome100XPercentage Sample % > 100X bigWig gnomAD Percentage of Exome Samples with at least 100X Coverage v4.0 2 9 15 0 240 135 127 247 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/coverage/v4-exome/gnomad.coverage.over_100.bw\ color 15,0,240\ longLabel gnomAD Percentage of Exome Samples with at least 100X Coverage v4.0\ parent gnomad4ExomeCoverage off\ priority 9\ shortLabel Sample % > 100X\ track gnomad4Exome100XPercentage\ viewLimits 0:1\ SeqCap-EZ_MedExomePlusMito_hg19_capture_targets SeqCap EZ Med+Mito P bigBed Roche - SeqCap EZ MedExome + Mito Capture Probe Footprint 0 9 100 143 255 177 199 255 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/SeqCap_EZ_MedExomePlusMito_hg38_capture_targets.bb\ color 100,143,255\ longLabel Roche - SeqCap EZ MedExome + Mito Capture Probe Footprint\ parent exomeProbesets on\ shortLabel SeqCap EZ Med+Mito P\ track SeqCap-EZ_MedExomePlusMito_hg19_capture_targets\ type bigBed\ ultraZoo UltraZoos bigBed 3 UltraZoos: 4552 Ultraconserved regions in Zoonomia alignment - 100% identical in 235 species, >20bp 0 9 0 0 0 127 127 127 0 0 0 compGeno 1 bigDataUrl /gbdb/hg38/unusualcons/zooUCEs.bigBed\ longLabel UltraZoos: 4552 Ultraconserved regions in Zoonomia alignment - 100% identical in 235 species, >20bp\ parent unusualcons on\ shortLabel UltraZoos\ track ultraZoo\ type bigBed 3\ vertebrateChainNet Vertebrate Chain/Net bed 3 Non-placental Vertebrate Genomes, Chain and Net Alignments 0 9 0 0 0 255 255 0 0 0 0\ The chain track shows alignments of human (Dec. 2013 (GRCh38/hg38)) to\ other genomes using a gap scoring system that allows longer gaps \ than traditional affine gap scoring systems. It can also tolerate gaps in both\ human and the other genome simultaneously. These \ "double-sided" gaps can be caused by local inversions and \ overlapping deletions in both species. \
\ The chain track displays boxes joined together by either single or\ double lines. The boxes represent aligning regions.\ Single lines indicate gaps that are largely due to a deletion in the\ other assembly or an insertion in the human assembly.\ Double lines represent more complex gaps that involve substantial\ sequence in both species. This may result from inversions, overlapping\ deletions, an abundance of local mutation, or an unsequenced gap in one\ species. In cases where multiple chains align over a particular region of\ the other genome, the chains with single-lined gaps are often \ due to processed pseudogenes, while chains with double-lined gaps are more \ often due to paralogs and unprocessed pseudogenes.
\\ In the "pack" and "full" display\ modes, the individual feature names indicate the chromosome, strand, and\ location (in thousands) of the match for each matching alignment.
\ \\ The net track shows the best human/other chain for \ every part of the other genome. It is useful for\ finding orthologous regions and for studying genome\ rearrangement. The human sequence used in this annotation is from\ the Dec. 2013 (GRCh38/hg38) assembly.
\ \By default, the chains to chromosome-based assemblies are colored\ based on which chromosome they map to in the aligning organism. To turn\ off the coloring, check the "off" button next to: Color\ track based on chromosome.
\\ To display only the chains of one chromosome in the aligning\ organism, enter the name of that chromosome (e.g. chr4) in box next to: \ Filter by chromosome.
\ \\ In full display mode, the top-level (level 1)\ chains are the largest, highest-scoring chains that\ span this region. In many cases gaps exist in the\ top-level chain. When possible, these are filled in by\ other chains that are displayed at level 2. The gaps in \ level 2 chains may be filled by level 3 chains and so\ forth.
\\ In the graphical display, the boxes represent ungapped \ alignments; the lines represent gaps. Click\ on a box to view detailed information about the chain\ as a whole; click on a line to display information\ about the gap. The detailed information is useful in determining\ the cause of the gap or, for lower level chains, the genomic\ rearrangement.
\\ Individual items in the display are categorized as one of four types\ (other than gap):
\\ Transposons that have been inserted since the human/other\ split were removed from the assemblies. The abbreviated genomes were\ aligned with lastz, and the transposons were added back in.\ The resulting alignments were converted into axt format using the lavToAxt\ program. The axt alignments were fed into axtChain, which organizes all\ alignments between a single human chromosome and a single\ chromosome from the other genome into a group and creates a kd-tree out\ of the gapless subsections (blocks) of the alignments. A dynamic program\ was then run over the kd-trees to find the maximally scoring chains of these\ blocks.\ \
\ \ For the Wallaby alignment, chains scoring below a minimum score\ of '3000' were discarded; the remaining chains are displayed in this track.\ The linear gap matrix used with axtChain:\
\ \
\ The following lastz matrix was used\ \
for the alignments to: Wallaby, Tasmanian Devil\ \\
\ A C G T \ A 91 -114 -31 -123 \ \ C -114 100 -125 -31 \ G -31 -125 100 -114 \ T -123 -31 -114 91 \ \
\ The following lastz matrix was used\
for the alignments to: American Alligator, Medium Ground Finch,
\ Opossum, Platypus, Chicken, Zebra Finch, Lizard, X. tropicalis,
\ Stickleback, Fugu, Zebrafish, Tetraodon, Medaka, Lamprey\ \\
\ A C G T \ A 91 -90 -25 -100 \ C -90 100 -100 -25 \ G -25 -100 100 -90 \ T -100 -25 -90 91
-linearGap=medium\ \ tableSize 11\ smallSize 111\ position 1 2 3 11 111 2111 12111 32111 72111 152111 252111\ qGap 350 425 450 600 900 2900 22900 57900 117900 217900 317900\ tGap 350 425 450 600 900 2900 22900 57900 117900 217900 317900\ bothGap 750 825 850 1000 1300 3300 23300 58300 118300 218300 318300\\ \ For the alignments to: American Alligator, Medium Ground Finch, Tasmanian Devil, Opossum, Platypus, Chicken,\ Zebra Finch, Lizard, X. tropicalis, Stickleback, Fugu, Zebrafish, Tetraodon,\ Medaka and Lamprey, chains scoring below a minimum score\ of '5000' were discarded; the remaining chains are displayed\ in this track. The linear gap matrix used with axtChain:
-linearGap=loose\ \ tablesize 11\ smallSize 111\ position 1 2 3 11 111 2111 12111 32111 72111 152111 252111\ qGap 325 360 400 450 600 1100 3600 7600 15600 31600 56600\ tGap 325 360 400 450 600 1100 3600 7600 15600 31600 56600\ bothGap 625 660 700 750 900 1400 4000 8000 16000 32000 57000\\ \ See also: lastz parameters used in these alignments,\ and chain minimum score and gap parameters used in these alignments.\ \ \
\ Chains were derived from lastz alignments, using the methods\ described on the chain tracks description pages, and sorted with the \ highest-scoring chains in the genome ranked first. The program\ chainNet was then used to place the chains one at a time, trimming them as \ necessary to fit into sections not already covered by a higher-scoring chain. \ During this process, a natural hierarchy emerged in which a chain that filled \ a gap in a higher-scoring chain was placed underneath that chain. The program \ netSyntenic was used to fill in information about the relationship between \ higher- and lower-level chains, such as whether a lower-level\ chain was syntenic or inverted relative to the higher-level chain. \ The program netClass was then used to fill in how much of the gaps and chains \ contained Ns (sequencing gaps) in one or both species and how much\ was filled with transposons inserted before and after the two organisms \ diverged.
\ \\ LASTZ was developed at\ Miller Lab at Pennsylvania State University by \ Bob Harris.\
\\ Lineage-specific repeats were identified by Arian Smit and his \ RepeatMasker\ program.
\\ The axtChain program was developed at the University of California at \ Santa Cruz by Jim Kent with advice from Webb Miller and David Haussler.
\\ The browser display and database storage of the chains and nets were created\ by Robert Baertsch and Jim Kent.
\\ The chainNet, netSyntenic, and netClass programs were\ developed at the University of California\ Santa Cruz by Jim Kent.
\\ \
\ Harris RS.\ Improved pairwise alignment of genomic DNA.\ Ph.D. Thesis. Pennsylvania State University, USA. 2007.\
\ \\ Chiaromonte F, Yap VB, Miller W.\ Scoring pairwise genomic sequence alignments.\ Pac Symp Biocomput. 2002:115-26.\ PMID: 11928468\
\ \\ Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D.\ Evolution's cauldron:\ duplication, deletion, and rearrangement in the mouse and human genomes.\ Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9.\ PMID: 14500911; PMC: PMC208784\
\ \\ Schwartz S, Kent WJ, Smit A, Zhang Z, Baertsch R, Hardison RC,\ Haussler D, Miller W.\ Human-mouse alignments with BLASTZ.\ Genome Res. 2003 Jan;13(1):103-7.\ PMID: 12529312; PMC: PMC430961\
\ compGeno 1 altColor 255,255,0\ chainLinearGap loose\ chainMinScore 5000\ color 0,0,0\ compositeTrack on\ configurable on\ dimensions dimensionX=clade dimensionY=species\ dragAndDrop subTracks\ group compGeno\ html vertebrateChainNet\ longLabel Non-placental Vertebrate Genomes, Chain and Net Alignments\ noInherit on\ priority 9\ shortLabel Vertebrate Chain/Net\ sortOrder species=+ view=+ clade=+\ subGroup1 view Views chain=Chains net=Nets\ subGroup2 species Species s000=Wallaby s001=Wallaby s002=Tasmanian_devil s003=Opossum s004a=Platypus s004b=Platypus s005=Platypus s006=Turkey s007a=Turkey s007b=Japanese_quail s008a=Chicken s008b=Chicken s009=Chicken s010=Chicken s011=Mallard_duck s012=Scarlet_macaw s013=Medium_ground_finch s014=White-throated_sparrow s015=Collared_flycatcher s016=Golden_eagle s017=Peregrine_falcon s018=Saker_falcon s019=Rock_pigeon s020=Parrot s021=Budgerigar s022=Tibetan_ground_jay s023=Zebra_finch s024=Zebra_finch s025=American_alligator s026=Chinese_alligator s027=Lizard s028=Lizard s028b=Garter_snake s029a=Axolotl s029b=X._tropicalis s029c=X._tropicalis s030=X._tropicalis s031=X._tropicalis s032=X._tropicalis s033=X._tropicalis s034=African_clawed_frog s035=Spiny_softshell_turtle s036=Chinese_softshell_turtle s037=Painted_turtle s038=Painted_turtle s039=Green_seaturtle s040=Coelacanth s041=Spotted_gar s042=Mexican_tetra_(cavefish) s043=Zebrafish s044=Zebrafish s045=Zebrafish s046=Zebrafish s047=Zebrafish s048=Atlantic_cod s049=Stickleback s050=Southern_platyfish s051=Medaka s052=Pundamilia_nyererei s053=Zebra_mbuna s054=Princess_of_Burundi s055=Burton's_mouthbreeder s056=Nile_tilapia s057=Nile_tilapia s058=Nile_tilapia s059=Yellowbelly_pufferfish s060=Fugu s061=Fugu s062=Tetraodon s063=Tetraodon s064a=Lamprey s064b=Lamprey s065=Lamprey s066=Arctic_lamprey\ subGroup3 clade Clade c00=mammalia c01=dinosauria c02=lepidosauria c03=amphibia c04=cryptodira c05=coelancanthimorpha c06=neopterygii c07=hyperoartia\ track vertebrateChainNet\ type bed 3\ visibility hide\ netThaSir1 thaSir1 Net netAlign anoCar1 chainAnoCar1 Garter snake (Jun. 2015 (Thamnophis_sirtalis-6.0/thaSir1)) Alignment Net 1 10 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Garter snake (Jun. 2015 (Thamnophis_sirtalis-6.0/thaSir1)) Alignment Net\ otherDb thaSir1\ parent vertebrateChainNetViewnet off\ shortLabel thaSir1 Net\ subGroups view=net species=s028b clade=c02\ track netThaSir1\ type netAlign anoCar1 chainAnoCar1\ netRn7 Rat Net netAlign rn7 chainRn7 Rat (Nov. 2020 (mRatBN7.2/rn7)) Alignment Net 1 10 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Rat (Nov. 2020 (mRatBN7.2/rn7)) Alignment Net\ otherDb rn7\ parent placentalChainNetViewnet on\ shortLabel Rat Net\ subGroups view=net species=s024a clade=c00\ track netRn7\ type netAlign rn7 chainRn7\ netNomLeu3 Gibbon Net netAlign nomLeu3 chainNomLeu3 Gibbon (Oct. 2012 (GGSC Nleu3.0/nomLeu3)) Alignment Net 1 10 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Gibbon (Oct. 2012 (GGSC Nleu3.0/nomLeu3)) Alignment Net\ otherDb nomLeu3\ parent primateChainNetViewnet off\ shortLabel Gibbon Net\ subGroups view=net species=s014 clade=c00\ track netNomLeu3\ type netAlign nomLeu3 chainNomLeu3\ encTfChipPkENCFF615GTV A549 CTCF 2 narrowPeak Transcription Factor ChIP-seq Peaks of CTCF in A549 from ENCODE 3 (ENCFF615GTV) 0 10 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of CTCF in A549 from ENCODE 3 (ENCFF615GTV)\ parent encTfChipPk off\ shortLabel A549 CTCF 2\ subGroups cellType=A549 factor=CTCF\ track encTfChipPkENCFF615GTV\ cloneEndABC22 ABC22 bed 12 Agencourt fosmid library 22 0 10 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 22\ parent cloneEndSuper off\ priority 10\ shortLabel ABC22\ subGroups source=agencourt\ track cloneEndABC22\ type bed 12\ visibility hide\ wgEncodeReg4AtacAllAdrenalGland Adrenal gland (all biosamples) bigWig Avg. ATAC level of 8 adrenal gland experiments (all biosamples) 0 10 90 179 68 172 217 161 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adrenalGlandATAC.bw\ color 90,179,68\ longLabel Avg. ATAC level of 8 adrenal gland experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 10\ shortLabel Adrenal gland (all biosamples)\ track wgEncodeReg4AtacAllAdrenalGland\ type bigWig\ AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep3LK6_CNhs13568_ctss_rev AorticSmsToFgf2_00hr15minBr3- bigWig Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep3 (LK6)_CNhs13568_12839-137B4_reverse 0 10 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12839-137B4 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr15min%2c%20biol_rep3%20%28LK6%29.CNhs13568.12839-137B4.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep3 (LK6)_CNhs13568_12839-137B4_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12839-137B4 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr15minBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep3LK6_CNhs13568_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12839-137B4\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep3LK6_CNhs13568_tpm_rev AorticSmsToFgf2_00hr15minBr3- bigWig Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep3 (LK6)_CNhs13568_12839-137B4_reverse 1 10 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12839-137B4 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr15min%2c%20biol_rep3%20%28LK6%29.CNhs13568.12839-137B4.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr15min, biol_rep3 (LK6)_CNhs13568_12839-137B4_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12839-137B4 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr15minBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr15minBiolRep3LK6_CNhs13568_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12839-137B4\ urlLabel FANTOM5 Details:\ bismap36Quantitative Bismap M36 bigWig 0.027778 1.00 Multi-read mappability with 36-mers after bisulfite conversion 0 10 240 70 80 247 162 167 0 0 0 map 0 bigDataUrl /gbdb/hg38/hoffmanMappability/k36.Bismap.MultiTrackMappability.bw\ color 240,70,80\ longLabel Multi-read mappability with 36-mers after bisulfite conversion\ parent bismapBigWig off\ priority 10\ shortLabel Bismap M36\ subGroups view=MR\ track bismap36Quantitative\ type bigWig 0.027778 1.00\ visibility hide\ wgEncodeReg4TxnBrainMinus Brain - bigWig Avg. - strand total RNA-seq level of 127 brain experiments (tissues and primary cells only) 0 10 155 155 18 205 205 136 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBrainMinus.bw\ color 155,155,18\ longLabel Avg. - strand total RNA-seq level of 127 brain experiments (tissues and primary cells only)\ negateValues on\ parent wgEncodeReg4Txn\ priority 10\ shortLabel Brain -\ track wgEncodeReg4TxnBrainMinus\ type bigWig\ gtexCovBrainCaudatebasalganglia Brain Caud bas gangl bigWig Brain Caudate basal ganglia 0 10 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-1HGF4-0011-R5b-SM-CM2ST.Brain_Caudate_basal_ganglia.RNAseq.bw\ color 238,238,0\ longLabel Brain Caudate basal ganglia\ parent gtexCov\ shortLabel Brain Caud bas gangl\ track gtexCovBrainCaudatebasalganglia\ cortexNeuron42P Cortex - Neuron - Z0000042P bigWig Methylation Atlas: Cortex - Neuron - Z0000042P 2 10 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/cortexNeuron42P.bw\ color 138,43,226\ longLabel Methylation Atlas: Cortex - Neuron - Z0000042P\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 10\ shortLabel Cortex - Neuron - Z0000042P\ subGroups cellType=Neuron dataType=Replicate\ track cortexNeuron42P\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ dbVar_common_global dbVar Curated All Populations bigBed 9 + . NCBI dbVar Curated Common SVs: all populations 3 10 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/variants/$$ varRep 1 bigDataUrl /gbdb/hg38/bbi/dbVar/common_global.bb\ longLabel NCBI dbVar Curated Common SVs: all populations\ parent dbVar_common on\ priority 10\ shortLabel dbVar Curated All Populations\ track dbVar_common_global\ type bigBed 9 + .\ url https://www.ncbi.nlm.nih.gov/dbvar/variants/$$\ urlLabel NCBI Variant Page:\ ENCFF136RNO_ENCFF630BQS_ENCFF611XLA_ENCFF975BGM ENCFF136RNO_ENCFF630BQS_ENCFF611XLA_ENCFF975BGM bigBed 9 + 5 OCI-LY7: (1) cCREs 4 10 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF136RNO_ENCFF630BQS_ENCFF611XLA_ENCFF975BGM.bb\ longLabel OCI-LY7: (1) cCREs\ mouseOver ID: ${name}\ The GenomeAsia 100K project aims\ to sequence 100,000 Asian individuals. This pilot release (GAsP) contains whole-genome sequencing\ data of 1,739 individuals from 219 population groups across Asia. Frequencies are broken down by\ Northeast Asian, Southeast Asian, and South Asian ancestry groups. The data is split into two\ subtracks: substitutions and indels.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is gasp.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original VCFs are also available from the\ GenomeAsia 100K\ website. No license nor login is required.\
\ \\ Samples were sequenced on Illumina HiSeq 2500, HiSeq 4000, and HiSeq X Ten instruments with\ 2×100 bp or 2×150 bp paired-end reads at an average depth of 36x. Reads were aligned to\ GRCh37 using BWA-MEM. Duplicate reads were marked with SAMBLASTER and sorted with Sambamba.\ Per-sample variant calling was performed with GATK HaplotypeCaller in GVCF mode, followed by\ joint genotyping with GenotypeGVCFs. Variant quality score recalibration (VQSR) was applied at\ a 99% sensitivity tranche for both SNPs and indels. Sample-level QC included contamination\ checks with verifyBamID and sex concordance verification. The final callset contains\ ∼65 million variants across 1,739 individuals from 219 populations.\
\\ The upstream callset is on GRCh37. We lifted it to hg38 using\ CrossMap and the UCSC\ hg19ToHg38 chain file. After lifting, variants that landed on alt, random, fix, or\ unplaced contigs were dropped, and the result was sorted and indexed with tabix.\
\\ The makeDoc file documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were needed and are also available from GitHub.\
\ \\ GenomeAsia100K Consortium.\ \ The GenomeAsia 100K Project enables genetic discoveries across Asia.\ Nature. 2019 Dec;576(7785):106-111.\ PMID: 31802016; PMC: PMC7054211\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/ga100k/ga100k.indels.vcf.gz\ dataVersion Pilot 2019 (lifted to hg38, May 2026)\ html gasp\ longLabel SNV Frequencies: GenomeAsia Pilot - Indels\ parent varFreqs on\ priority 10\ shortLabel GenomeAsia 1.7k Indels\ track gaspIndel\ type vcfTabix\ visibility hide\ gnomadConstraint gnomAD Mut Constraint bigWig Gnocchi: Genome Aggregation Database (gnomAD) non-coding constraint of haploinsufficient variation, includes chrX 0 10 150 0 0 0 150 0 0 0 0GnomAD Genome Mutational Constraint, also known as "Genome non-coding constraint of\ haploinsufficient variation (Gnocchi)", is based on v3.1.2 and is available only on hg38.\ It shows the reduced variation caused by purifying\ natural selection. This is similar to negative selection on loss-of-function\ (LoF) for genes, but can be calculated for non-coding regions too.\ Positive values are red and reflect stronger mutation constraint (and less variation), indicating\ higher natural selection pressure in a region. Negative values are green and\ reflect lower mutation constraint\ (and more variation), indicating less selection pressure and less functional effect.\ Briefly, for any 1kbp window in\ the genome, a model based on trinucleotide sequence context, base-level\ methylation, and regional genomic features predicts expected number of mutations,\ and compares this number to the observed number of mutations using a Z-score (see Chen et al 2024\ in the Reference section for details). The chrX scores were added as received from the authors,\ as there are no de novo mutation data available on chrX (for estimating the effects of regional\ genomic features on mutation rates), they are more speculative than the ones on the autosomes.
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below).
\ \The mutational constraints score was updated in October 2022 from a previous,\ now deprecated, pre-publication version. The old version can be found in our\ archive\ directory on the download server. It can be loaded by copying the URL into\ our "Custom tracks" input box.
\ \\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\ varRep 0 altColor 0,150,0\ autoScale on\ bigDataUrl /gbdb/hg38/gnomAD/mutConstraint/mutConstraint.bw\ color 150,0,0\ dataVersion Release 3.1.2 (October 22, 2021)\ html gnomadConstraint\ longLabel Gnocchi: Genome Aggregation Database (gnomAD) non-coding constraint of haploinsufficient variation, includes chrX\ maxHeightPixels 128:40:8\ parent gnomadVariants on\ priority 10\ setColorWith /gbdb/hg38/gnomAD/mutConstraint/mutConstraint.color.bb\ shortLabel gnomAD Mut Constraint\ track gnomadConstraint\ type bigWig\ viewLimitsMax -3:3\ windowingFunction minimum\ wgEncodeReg4DnaseHeart Heart bigWig Avg. DNase level of 53 heart experiments (tissues and primary cells only) 0 10 116 50 165 185 152 210 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpHeartDNase.bw\ color 116,50,165\ longLabel Avg. DNase level of 53 heart experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 10\ shortLabel Heart\ track wgEncodeReg4DnaseHeart\ type bigWig\ wgEncodeRegDnaseUwHffmycPeak HFF-Myc Pk narrowPeak HFF-Myc foreskin fibroblast cell line, cMyc DNaseI Peaks from ENCODE 1 10 255 165 85 255 210 170 1 0 0 regulation 1 color 255,165,85\ longLabel HFF-Myc foreskin fibroblast cell line, cMyc DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel HFF-Myc Pk\ subGroups view=a_Peaks cellType=HFF-Myc treatment=n_a tissue=skin cancer=normal\ track wgEncodeRegDnaseUwHffmycPeak\ wgEncodeRegDnaseUwHffmycWig HFF-Myc Sg bigWig 0 23416.2 HFF-Myc foreskin fibroblast cell line, cMyc DNaseI Signal from ENCODE 0 10 255 165 85 255 210 170 0 0 0 regulation 1 color 255,165,85\ longLabel HFF-Myc foreskin fibroblast cell line, cMyc DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.08427\ shortLabel HFF-Myc Sg\ subGroups cellType=HFF-Myc treatment=n_a tissue=skin cancer=normal\ table wgEncodeRegDnaseUwHffmycSignal\ track wgEncodeRegDnaseUwHffmycWig\ type bigWig 0 23416.2\ netHprcGCA_018505825v1 HG02109.mat netAlign GCA_018505825.1 chainHprcGCA_018505825v1 HG02109.mat HG02109.pri.mat.f1_v2 (May 2021 GCA_018505825.1_HG02109.pri.mat.f1_v2) HPRC project computed Chain Nets 1 10 0 0 0 255 255 0 0 0 0 hprc 0 longLabel HG02109.mat HG02109.pri.mat.f1_v2 (May 2021 GCA_018505825.1_HG02109.pri.mat.f1_v2) HPRC project computed Chain Nets\ otherDb GCA_018505825.1\ parent hprcChainNetViewnet off\ priority 25\ shortLabel HG02109.mat\ subGroups view=net sample=s025 population=afr subpop=acb hap=mat\ track netHprcGCA_018505825v1\ type netAlign GCA_018505825.1 chainHprcGCA_018505825v1\ lincRNAsCThLF_r2 hLF_r2 bed 5 + lincRNAs from hlf_r2 1 10 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from hlf_r2\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel hLF_r2\ subGroups view=lincRNAsRefseqExp tissueType=hlf_r2\ track lincRNAsCThLF_r2\ wgEncodeReg4MarkH3k4me3LargeIntestine Large intestine bigWig Avg. H3K4me3 level of 20 large intestine experiments (tissues and primary cells only) 0 10 86 86 36 170 170 145 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLargeIntestineH3K4me3.bw\ color 86,86,36\ longLabel Avg. H3K4me3 level of 20 large intestine experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 10\ shortLabel Large intestine\ track wgEncodeReg4MarkH3k4me3LargeIntestine\ type bigWig\ wgEncodeReg4MarkH3k27acLiver Liver bigWig Avg. H3K27ac level of 4 liver experiments (tissues and primary cells only) 2 10 137 152 82 196 203 168 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLiverH3K27ac.bw\ color 137,152,82\ longLabel Avg. H3K27ac level of 4 liver experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac\ priority 10\ shortLabel Liver\ track wgEncodeReg4MarkH3k27acLiver\ type bigWig\ wgEncodeReg4MarkCtcfLung Lung bigWig Avg. CTCF level of 13 lung experiments (tissues and primary cells only) 0 10 130 163 45 192 209 150 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLungCTCF.bw\ color 130,163,45\ longLabel Avg. CTCF level of 13 lung experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf off\ priority 10\ shortLabel Lung\ track wgEncodeReg4MarkCtcfLung\ type bigWig\ unipMut Mutations bigBed 12 + UniProt Amino Acid Mutations 1 10 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipMut.bb\ longLabel UniProt Amino Acid Mutations\ mouseOver UniProt record: $uniProtId\ This track contains GENCODE or Ensembl alignments produced by\ the TransMap cross-species alignment algorithm from other vertebrate\ species in the UCSC Genome Browser. GENCODE is Ensembl for human and mouse,\ for other Ensembl sources, only ones with full gene builds are used.\ Projection Ensembl gene annotations will not be used as sources.\ For closer evolutionary distances, the alignments are created using\ syntenically filtered BLASTZ alignment chains, resulting in a prediction of the\ orthologous genes in human.\
\ \ \\ This track follows the display conventions for \ PSL alignment tracks.
\\ This track may also be configured to display codon coloring, a feature that\ allows the user to quickly compare cDNAs against the genomic sequence. For more \ information about this option, click \ here.\ Several types of alignment gap may also be colored; \ for more information, click \ here.\ \
\
\ To ensure unique identifiers for each alignment, cDNA and gene accessions were\ made unique by appending a suffix for each location in the source genome and\ again for each mapped location in the destination genome. The format is:\
\ accession.version-srcUniq.destUniq\\ \ Where srcUniq is a number added to make each source alignment unique, and\ destUniq is added to give the subsequent TransMap alignments unique\ identifiers.\ \
\ For example, in the cow genome, there are two alignments of mRNA BC149621.1.\ These are assigned the identifiers BC149621.1-1 and BC149621.1-2.\ When these are mapped to the human genome, BC149621.1-1 maps to a single\ location and is given the identifier BC149621.1-1.1. However, BC149621.1-2\ maps to two locations, resulting in BC149621.1-2.1 and BC149621.1-2.2. Note\ that multiple TransMap mappings are usually the result of tandem duplications, where both\ chains are identified as syntenic.\
\ \\ The raw data for these tracks can be accessed interactively through the\ Table Browser or the\ Data Integrator.\ For automated analysis, the annotations are stored in\ bigPsl files (containing a\ number of extra columns) and can be downloaded from our\ download server, \ or queried using our API. For more \ information on accessing track data see our \ Track Data Access FAQ.\ The files are associated with these tracks in the following way:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/transMap/V4/hg38.refseq.transMapV4.bigPsl\ -chrom=chr6 -start=0 -end=1000000 stdout\ \ \
\ This track was produced by Mark Diekhans at UCSC from cDNA and EST sequence data\ submitted to the international public sequence databases by \ scientists worldwide and annotations produced by the RefSeq,\ Ensembl, and GENCODE annotations projects.
\ \\ Siepel A, Diekhans M, Brejová B, Langton L, Stevens M, Comstock CL, Davis C, Ewing B, Oommen S,\ Lau C et al.\ \ Targeted discovery of novel human exons by comparative genomics.\ Genome Res. 2007 Dec;17(12):1763-73.\ PMID: 17989246; PMC: PMC2099585\
\ \\ Stanke M, Diekhans M, Baertsch R, Haussler D.\ \ Using native and syntenically mapped cDNA alignments to improve de novo gene finding.\ Bioinformatics. 2008 Mar 1;24(5):637-44.\ PMID: 18218656\
\ \\ Zhu J, Sanborn JZ, Diekhans M, Lowe CB, Pringle TH, Haussler D.\ \ Comparative genomics search for losses of long-established genes on the human lineage.\ PLoS Comput Biol. 2007 Dec;3(12):e247.\ PMID: 18085818; PMC: PMC2134963\
\ \ genes 1 baseColorDefault diffCodons\ baseColorUseCds given\ baseColorUseSequence lfExtra\ bigDataUrl /gbdb/hg38/transMap/V5/hg38.ensembl.transMapV5.bigPsl\ canPack on\ color 0,100,0\ defaultLabelFields orgAbbrev,geneName\ group genes\ html transMapEnsembl\ indelDoubleInsert on\ indelQueryInsert on\ labelFields commonName,orgAbbrev,srcDb,srcTransId,name,geneName,geneId,geneType,transcriptType\ labelSeparator " "\ longLabel TransMap Ensembl and GENCODE Mappings Version 5\ priority 10.001\ searchIndex name,srcTransId,geneName,geneId\ shortLabel TransMap Ensembl\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ superTrack transMapV5 pack\ track transMapEnsemblV5\ transMapSrcSet ensembl\ type bigPsl\ visibility pack\ transMapRefSeqV5 TransMap RefGene bigPsl TransMap RefSeq Gene Mappings Version 5 3 10.003 0 100 0 127 177 127 0 0 0\ This track contains RefSeq Gene alignments produced by\ the TransMap cross-species alignment algorithm\ from other vertebrate species in the UCSC Genome Browser.\ For closer evolutionary distances, the alignments are created using\ syntenically filtered BLASTZ alignment chains, resulting in a prediction of the\ orthologous genes in human.\
\ \ \\ This track follows the display conventions for \ PSL alignment tracks.
\\ This track may also be configured to display codon coloring, a feature that\ allows the user to quickly compare cDNAs against the genomic sequence. For more \ information about this option, click \ here.\ Several types of alignment gap may also be colored; \ for more information, click \ here.\ \
\
\ To ensure unique identifiers for each alignment, cDNA and gene accessions were\ made unique by appending a suffix for each location in the source genome and\ again for each mapped location in the destination genome. The format is:\
\ accession.version-srcUniq.destUniq\\ \ Where srcUniq is a number added to make each source alignment unique, and\ destUniq is added to give the subsequent TransMap alignments unique\ identifiers.\ \
\ For example, in the cow genome, there are two alignments of mRNA BC149621.1.\ These are assigned the identifiers BC149621.1-1 and BC149621.1-2.\ When these are mapped to the human genome, BC149621.1-1 maps to a single\ location and is given the identifier BC149621.1-1.1. However, BC149621.1-2\ maps to two locations, resulting in BC149621.1-2.1 and BC149621.1-2.2. Note\ that multiple TransMap mappings are usually the result of tandem duplications, where both\ chains are identified as syntenic.\
\ \\ The raw data for these tracks can be accessed interactively through the\ Table Browser or the\ Data Integrator.\ For automated analysis, the annotations are stored in\ bigPsl files (containing a\ number of extra columns) and can be downloaded from our\ download server, \ or queried using our API. For more \ information on accessing track data see our \ Track Data Access FAQ.\ The files are associated with these tracks in the following way:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/transMap/V4/hg38.refseq.transMapV4.bigPsl\ -chrom=chr6 -start=0 -end=1000000 stdout\ \ \
\ This track was produced by Mark Diekhans at UCSC from cDNA and EST sequence data\ submitted to the international public sequence databases by \ scientists worldwide and annotations produced by the RefSeq,\ Ensembl, and GENCODE annotations projects.
\ \\ Siepel A, Diekhans M, Brejová B, Langton L, Stevens M, Comstock CL, Davis C, Ewing B, Oommen S,\ Lau C et al.\ \ Targeted discovery of novel human exons by comparative genomics.\ Genome Res. 2007 Dec;17(12):1763-73.\ PMID: 17989246; PMC: PMC2099585\
\ \\ Stanke M, Diekhans M, Baertsch R, Haussler D.\ \ Using native and syntenically mapped cDNA alignments to improve de novo gene finding.\ Bioinformatics. 2008 Mar 1;24(5):637-44.\ PMID: 18218656\
\ \\ Zhu J, Sanborn JZ, Diekhans M, Lowe CB, Pringle TH, Haussler D.\ \ Comparative genomics search for losses of long-established genes on the human lineage.\ PLoS Comput Biol. 2007 Dec;3(12):e247.\ PMID: 18085818; PMC: PMC2134963\
\ \ genes 1 baseColorDefault diffCodons\ baseColorUseCds given\ baseColorUseSequence lfExtra\ bigDataUrl /gbdb/hg38/transMap/V5/hg38.refseq.transMapV5.bigPsl\ canPack on\ color 0,100,0\ defaultLabelFields orgAbbrev,geneName\ group genes\ html transMapRefSeq\ indelDoubleInsert on\ indelQueryInsert on\ labelFields commonName,orgAbbrev,srcDb,srcTransId,name,geneName,geneId\ labelSeparator " "\ longLabel TransMap RefSeq Gene Mappings Version 5\ priority 10.003\ searchIndex name,srcTransId,geneName,geneId\ shortLabel TransMap RefGene\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ superTrack transMapV5 pack\ track transMapRefSeqV5\ transMapSrcSet refseq\ type bigPsl\ visibility pack\ transMapRnaV5 TransMap RNA bigPsl TransMap GenBank RNA Mappings Version 5 0 10.004 0 100 0 127 177 127 0 0 0\ This track contains GenBank mRNA alignments produced by\ the TransMap cross-species alignment algorithm\ from other vertebrate species in the UCSC Genome Browser.\ For closer evolutionary distances, the alignments are created using\ syntenically filtered BLASTZ alignment chains, resulting in a prediction of the\ orthologous genes in human.\
\ \ \\ This track follows the display conventions for \ PSL alignment tracks.
\\ This track may also be configured to display codon coloring, a feature that\ allows the user to quickly compare cDNAs against the genomic sequence. For more \ information about this option, click \ here.\ Several types of alignment gap may also be colored; \ for more information, click \ here.\ \
\
\ To ensure unique identifiers for each alignment, cDNA and gene accessions were\ made unique by appending a suffix for each location in the source genome and\ again for each mapped location in the destination genome. The format is:\
\ accession.version-srcUniq.destUniq\\ \ Where srcUniq is a number added to make each source alignment unique, and\ destUniq is added to give the subsequent TransMap alignments unique\ identifiers.\ \
\ For example, in the cow genome, there are two alignments of mRNA BC149621.1.\ These are assigned the identifiers BC149621.1-1 and BC149621.1-2.\ When these are mapped to the human genome, BC149621.1-1 maps to a single\ location and is given the identifier BC149621.1-1.1. However, BC149621.1-2\ maps to two locations, resulting in BC149621.1-2.1 and BC149621.1-2.2. Note\ that multiple TransMap mappings are usually the result of tandem duplications, where both\ chains are identified as syntenic.\
\ \\ The raw data for these tracks can be accessed interactively through the\ Table Browser or the\ Data Integrator.\ For automated analysis, the annotations are stored in\ bigPsl files (containing a\ number of extra columns) and can be downloaded from our\ download server, \ or queried using our API. For more \ information on accessing track data see our \ Track Data Access FAQ.\ The files are associated with these tracks in the following way:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/transMap/V4/hg38.refseq.transMapV4.bigPsl\ -chrom=chr6 -start=0 -end=1000000 stdout\ \ \
\ This track was produced by Mark Diekhans at UCSC from cDNA and EST sequence data\ submitted to the international public sequence databases by \ scientists worldwide and annotations produced by the RefSeq,\ Ensembl, and GENCODE annotations projects.
\ \\ Siepel A, Diekhans M, Brejová B, Langton L, Stevens M, Comstock CL, Davis C, Ewing B, Oommen S,\ Lau C et al.\ \ Targeted discovery of novel human exons by comparative genomics.\ Genome Res. 2007 Dec;17(12):1763-73.\ PMID: 17989246; PMC: PMC2099585\
\ \\ Stanke M, Diekhans M, Baertsch R, Haussler D.\ \ Using native and syntenically mapped cDNA alignments to improve de novo gene finding.\ Bioinformatics. 2008 Mar 1;24(5):637-44.\ PMID: 18218656\
\ \\ Zhu J, Sanborn JZ, Diekhans M, Lowe CB, Pringle TH, Haussler D.\ \ Comparative genomics search for losses of long-established genes on the human lineage.\ PLoS Comput Biol. 2007 Dec;3(12):e247.\ PMID: 18085818; PMC: PMC2134963\
\ \ genes 1 baseColorDefault diffCodons\ baseColorUseCds given\ baseColorUseSequence lfExtra\ bigDataUrl /gbdb/hg38/transMap/V5/hg38.rna.transMapV5.bigPsl\ canPack on\ color 0,100,0\ defaultLabelFields orgAbbrev,srcTransId\ group genes\ html transMapRna\ indelDoubleInsert on\ indelQueryInsert on\ labelFields commonName,orgAbbrev,srcDb,srcTransId,name,geneName\ labelSeparator " "\ longLabel TransMap GenBank RNA Mappings Version 5\ priority 10.004\ searchIndex name,srcTransId,geneName\ shortLabel TransMap RNA\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ superTrack transMapV5 hide\ track transMapRnaV5\ transMapSrcSet rna\ type bigPsl\ visibility hide\ transMapEstV5 TransMap ESTs bigPsl TransMap EST Mappings Version 5 0 10.005 0 100 0 127 177 127 0 0 0\ This track contains GenBank spliced EST alignments produced by\ the TransMap cross-species alignment algorithm\ from other vertebrate species in the UCSC Genome Browser.\ For closer evolutionary distances, the alignments are created using\ syntenically filtered BLASTZ alignment chains, resulting in a prediction of the\ orthologous genes in human.\
\ \ \\ This track follows the display conventions for \ PSL alignment tracks.
\\ This track may also be configured to display codon coloring, a feature that\ allows the user to quickly compare cDNAs against the genomic sequence. For more \ information about this option, click \ here.\ Several types of alignment gap may also be colored; \ for more information, click \ here.\ \
\
\ To ensure unique identifiers for each alignment, cDNA and gene accessions were\ made unique by appending a suffix for each location in the source genome and\ again for each mapped location in the destination genome. The format is:\
\ accession.version-srcUniq.destUniq\\ \ Where srcUniq is a number added to make each source alignment unique, and\ destUniq is added to give the subsequent TransMap alignments unique\ identifiers.\ \
\ For example, in the cow genome, there are two alignments of mRNA BC149621.1.\ These are assigned the identifiers BC149621.1-1 and BC149621.1-2.\ When these are mapped to the human genome, BC149621.1-1 maps to a single\ location and is given the identifier BC149621.1-1.1. However, BC149621.1-2\ maps to two locations, resulting in BC149621.1-2.1 and BC149621.1-2.2. Note\ that multiple TransMap mappings are usually the result of tandem duplications, where both\ chains are identified as syntenic.\
\ \\ The raw data for these tracks can be accessed interactively through the\ Table Browser or the\ Data Integrator.\ For automated analysis, the annotations are stored in\ bigPsl files (containing a\ number of extra columns) and can be downloaded from our\ download server, \ or queried using our API. For more \ information on accessing track data see our \ Track Data Access FAQ.\ The files are associated with these tracks in the following way:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/transMap/V4/hg38.refseq.transMapV4.bigPsl\ -chrom=chr6 -start=0 -end=1000000 stdout\ \ \
\ This track was produced by Mark Diekhans at UCSC from cDNA and EST sequence data\ submitted to the international public sequence databases by \ scientists worldwide and annotations produced by the RefSeq,\ Ensembl, and GENCODE annotations projects.
\ \\ Siepel A, Diekhans M, Brejová B, Langton L, Stevens M, Comstock CL, Davis C, Ewing B, Oommen S,\ Lau C et al.\ \ Targeted discovery of novel human exons by comparative genomics.\ Genome Res. 2007 Dec;17(12):1763-73.\ PMID: 17989246; PMC: PMC2099585\
\ \\ Stanke M, Diekhans M, Baertsch R, Haussler D.\ \ Using native and syntenically mapped cDNA alignments to improve de novo gene finding.\ Bioinformatics. 2008 Mar 1;24(5):637-44.\ PMID: 18218656\
\ \\ Zhu J, Sanborn JZ, Diekhans M, Lowe CB, Pringle TH, Haussler D.\ \ Comparative genomics search for losses of long-established genes on the human lineage.\ PLoS Comput Biol. 2007 Dec;3(12):e247.\ PMID: 18085818; PMC: PMC2134963\
\ \ genes 1 baseColorDefault none\ baseColorUseSequence lfExtra\ bigDataUrl /gbdb/hg38/transMap/V5/hg38.est.transMapV5.bigPsl\ canPack on\ color 0,100,0\ defaultLabelFields orgAbbrev,srcTransId\ group genes\ html transMapEst\ indelDoubleInsert on\ indelQueryInsert on\ labelFields commonName,orgAbbrev,srcDb,srcTransId,name\ labelSeparator " "\ longLabel TransMap EST Mappings Version 5\ priority 10.005\ searchIndex name,srcTransId\ shortLabel TransMap ESTs\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ superTrack transMapV5 hide\ track transMapEstV5\ transMapSrcSet est\ type bigPsl\ visibility hide\ gtexCov GTEx RNA-Seq Coverage bigWig GTEx V8 RNA-Seq Read Coverage by Tissue 0 10.2 0 0 0 127 127 127 0 0 0\ The\ \ NIH Genotype-Tissue Expression (GTEx) project\ determined genetic variation and gene expression in 52 tissues and 2 cell lines\ using RNA-seq data (V8, August 2019), on 17,382 samples from 948 adults.\ This track focuses on the gene expression part. It shows read coverage, from one\ single sample per tissue, selected for high-quality and high read depth.\ The data is summarized to one number per base pair, the number of sequencing\ reads that cover this position. The plot allows finding out if a given exon is\ transcribed primarily in certain tissues and also whether transcription is\ uniform over the length of a single exon.\
\ \\ This track follows the display conventions for composite \ "wiggle" tracks. The subtracks, one per tissue, of this track \ may be configured in a variety of ways to highlight different aspects of the \ displayed data. The graphical configuration options are shown at the top of \ the track description page, followed by a list of subtracks. To display only \ selected subtracks, uncheck the boxes next to the tracks you wish to hide. \ For more information about the graphical configuration options, click the \ Graph\ configuration help link.
\ Tissue colors were assigned to conform to the GTEx Consortium publication conventions.\ \ \ In Dense mode, the darkness of the grayscale rectangle displayed for the gene reflects the absolute\ read count.\ \ \For background information about GTEx sample selection, see our \ GTEx gene expression\ track. In short, samples were sequenced with the Illumina TrueSeq protocol\ on unstranded polyA+ librarires to obtain 76-bp paired end reads with\ HiSeq 2000 and 2500 machines.
\ \\ Sequence reads were aligned to the hg38/GRCh38 human genome using STAR v2.5.3a\ and the GENCODE 26 transcriptome. \ The alignment pipeline is available\ here.\ For further method details, see the \ \ GTEx Portal Documentation page.\
\ \\ To obtain read coverage, the GTEx Laboratory, Data Analysis and Coordinating\ Center (LDACC) at the Broad Institute decided to select a single, high-quality\ representative sample for each tissue type, since aggregated tracks may\ obscure certain features or even introduce some artifacts (e.g. intronic\ coverage). For each tissue, the selected sample has the highest RIN value with\ a high coverage (>80M reads) and exonic rate (>85%). \ The alignment-to-coverage pipeline is available from Github:\ Python script,\ Docker file and \ Pipeline WDL description. \
\To show the exact GTEx sample that was used for each tissue,\ click the "Schema" link on the track configuration page (above), the filename\ under "bigDataUrl" includes the identifier.
\ \\ The scientific goal of the GTEx project required that the donors and their biospecimen \ present with no evidence of disease. \ The tissue types collected were chosen based on their clinical significance, logistical \ feasibility and their relevance to the scientific goal of the project and the \ research community. \ Summary plots of GTEx sample characteristics are available at the \ \ GTEx Portal Tissue Summary page.
\ \\ The raw data for the GTEx Read Coverage track can be accessed interactively through the \ Table Browser.\
\ \ For automated analysis and downloads, the track data files can be downloaded from \ our downloads server\ or the JSON API.\ Individual regions or the whole genome annotation can be accessed as text using our utility\bigBedToBed. Instructions for downloading the utility can be found \
here. \
That utility can also be used to obtain features within a given range, e.g. \
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gtex/gtexGeneV8.bb -chrom=chr21\
-start=0 -end=100000000 stdout\
\ Data can also be obtained directly from GTEx at the following link:\ \ https://gtexportal.org/home/datasets
\ \\ Statistical analysis and data interpretation was performed by The GTEx Consortium Analysis \ Working Group. \ Data was provided by the GTEx LDACC at The Broad Institute of MIT and Harvard.
\ \\ GTEx Consortium.\ \ The GTEx Consortium atlas of genetic regulatory effects across human tissues.\ Science. 2020 Sep 11;369(6509):1318-1330.\ PMID: 32913098;\ PMC: PMC7737656
\ \ \\ GTEx Consortium.\ \ The Genotype-Tissue Expression (GTEx) project.\ Nat Genet. 2013 Jun;45(6):580-5.\ PMID: 23715323; \ PMC: PMC4010069
\ \\ Carithers LJ, Ardlie K, Barcus M, Branton PA, Britton A, Buia SA, Compton CC, DeLuca DS, \ Peter-Demchok J, Gelfand ET et al.\ \ A Novel Approach to High-Quality Postmortem Tissue Procurement: The GTEx Project.\ Biopreserv Biobank. 2015 Oct;13(5):311-9.\ PMID: 26484571; \ PMC: PMC4675181
\ \ Melé M, Ferreira PG, Reverter F, DeLuca DS, Monlong J, Sammeth M, Young TR, Goldmann JM,\ Pervouchine DD, Sullivan TJ et al.\ \ Human genomics. The human transcriptome across tissues and individuals.\ Science. 2015 May 8;348(6235):660-5.\ PMID: 25954002; PMC: PMC4547472\ \\ DeLuca DS, Levin JZ, Sivachenko A, Fennell T, Nazaire MD, Williams C, Reich M, Winckler W, Getz G.\ \ RNA-SeQC: RNA-seq metrics for quality control and process optimization.\ Bioinformatics. 2012 Jun 1;28(11):1530-2.\ PMID: 22539670; PMC: PMC3356847
\ expression 0 autoScale group\ compositeTrack on\ group expression\ longLabel GTEx V8 RNA-Seq Read Coverage by Tissue\ maxHeightPixels 100:50:8\ priority 10.20\ shortLabel GTEx RNA-Seq Coverage\ track gtexCov\ type bigWig\ ukbDepletion UKB Depl. Rank Score bigWig 0.0 1.0 UK Biobank / deCODE Genetics Depletion Rank Score 1 10.5 0 0 0 127 127 127 0 0 0\ The "Constraint scores" container track includes several subtracks showing the results of\ constraint prediction algorithms. These try to find regions of negative\ selection, where variations likely have functional impact. The algorithms do\ not use multi-species alignments to derive evolutionary constraint, but use\ primarily human variation, usually from variants collected by gnomAD (see the\ gnomAD V2 or V3 tracks on hg19 and hg38) or TOPMED (contained in our dbSNP\ tracks and available as a filter). One of the subtracks is based on UK Biobank\ variants, which are not available publicly, so we have no track with the raw data.\ The number of human genomes that are used as the input for these scores are\ 76k, 53k and 110k for gnomAD, TOPMED and UK Biobank, respectively.\
\ \Note that another important constraint score, gnomAD\ constraint, is not part of this container track but can be found in the hg38 gnomAD\ track.\
\ \ The algorithms included in this track are:\\ JARVIS scores are shown as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The scores were downloaded and converted to a single bigWig file.\ Move the mouse over the bars to display the exact values. A horizontal line is shown at the 0.733\ value which signifies the 90th percentile.
\ See hg19 makeDoc and\ hg38 makeDoc.\\ Interpretation: The authors offer a suggested guideline of > 0.9998 for identifying\ higher confidence calls and minimizing false positives. In addition to that strict threshold, the \ following two more relaxed cutoffs can be used to explore additional hits. Note that these\ thresholds are offered as guidelines and are not necessarily representative of pathogenicity.
\ \\
| Percentile | JARVIS score threshold |
|---|---|
| 99th | 0.9998 |
| 95th | 0.9826 |
| 90th | 0.7338 |
\ HMC scores are displayed as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The highly-constrained cutoff\ of 0.8 is indicated with a line.
\\ Interpretation: \ A protein residue with HMC score <1 indicates that missense variants affecting\ the homologous residues are significantly under negative selection (P-value <\ 0.05) and likely to be deleterious. A more stringent score threshold of HMC<0.8\ is recommended to prioritize predicted disease-associated variants.\
\ \\ Interpretation: The authors suggest the following guidelines for evaluating\ intolerance. By default, the MetaDome track displays a horizontal line at 0.7 which \ signifies the first intolerant bin. For more information see the MetaDome publication.
\ \\
| Classification | MetaDome Tolerance Score |
|---|---|
| Highly intolerant | ≤ 0.175 |
| Intolerant | ≤ 0.525 |
| Slightly intolerant | ≤ 0.7 |
\ MTR data can be found on two tracks, MTR All data and MTR Scores. In the\ MTR Scores track the data has been converted into 4 separate signal tracks\ representing each base pair mutation, with the lowest possible score shown when\ multiple transcripts overlap at a position. Overlaps can happen since this score\ is derived from transcripts and multiple transcripts can overlap. \ A horizontal line is drawn on the 0.8 score line\ to roughly represent the 25th percentile, meaning the items below may be of particular\ interest. It is recommended that the data be explored using\ this version of the track, as it condenses the information substantially while\ retaining the magnitude of the data.
\ \Any specific point mutations of interest can then be researched in the \ MTR All data track. This track contains all of the information from\ \ MTRV2 including more than 3 possible scores per base when transcripts overlap.\ A mouse-over on this track shows the ref and alt allele, as well as the MTR score\ and the MTR score percentile. Filters are available for MTR score, False Discovery Rate\ (FDR), MTR percentile, and variant consequence. By default, only items in the bottom\ 25 percentile are shown. Items in the track are colored according\ to their MTR percentile:
\\ Interpretation: Regions with low MTR scores were seen to be enriched with\ pathogenic variants. For example, ClinVar pathogenic variants were seen to\ have an average score of 0.77 whereas ClinVar benign variants had an average score\ of 0.92. Further validation using the FATHMM cancer-associated training dataset saw\ that scores less than 0.5 contained 8.6% of the pathogenic variants while only containing\ 0.9% of neutral variants. In summary, lower scores are more likely to represent\ pathogenic variants whereas higher scores could be pathogenic, but have a higher chance\ to be a false positive. For more information see the MTR-Viewer publication.
\ \\ Scores were downloaded and converted to a single bigWig file. See the\ hg19 makeDoc and the\ hg38 makeDoc for more info.\
\ \\ Scores were downloaded and converted to .bedGraph files with a custom Python \ script. The bedGraph files were then converted to bigWig files, as documented in our \ makeDoc hg19 build log.
\ \\
The authors provided a bed file containing codon coordinates along with the scores. \
This file was parsed with a python script to create the two tracks. For the first track\
the scores were aggregated for each coordinate, then the lowest score chosen for any\
overlaps and the result written out to bedGraph format. The file was then converted\
to bigWig with the bedGraphToBigWig utility. For the second track the file\
was reorganized into a bed 4+3 and conveted to bigBed with the bedToBigBed\
utility.
\ See the hg19 makeDoc for details including the build script.
\\ The raw MetaDome data can also be accessed via their Zenodo handle.
\ \\ V2\ file was downloaded and columns were reshuffled as well as itemRgb added for the\ MTR All data track. For the MTR Scores track the file was parsed with a python\ script to pull out the highest possible MTR score for each of the 3 possible mutations\ at each base pair and 4 tracks built out of these values representing each mutation.
\\ See the hg19 makeDoc entry on MTR for more info.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/hmc/hmc.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \ \\ Thanks to Jean-Madeleine Desainteagathe (APHP Paris, France) for suggesting the JARVIS, MTR, HMC tracks. Thanks to Xialei Zhang for providing the HMC data file and to Dimitrios Vitsios and Slave Petrovski for helping clean up the hg38 JARVIS files for providing guidance on interpretation. Additional\ thanks to Laurens van de Wiel for providing the MetaDome data as well as guidance on the track development and interpretation. \
\ \ \\ Vitsios D, Dhindsa RS, Middleton L, Gussow AB, Petrovski S.\ \ Prioritizing non-coding regions based on human genomic constraint and sequence context with deep\ learning.\ Nat Commun. 2021 Mar 8;12(1):1504.\ PMID: 33686085; PMC: PMC7940646\
\ \\ Xiaolei Zhang, Pantazis I. Theotokis, Nicholas Li, the SHaRe Investigators, Caroline F. Wright, Kaitlin E. Samocha, Nicola Whiffin, James S. Ware\ \ Genetic constraint at single amino acid resolution improves missense variant prioritisation and gene discovery.\ Medrxiv 2022.02.16.22271023\
\ \\ Wiel L, Baakman C, Gilissen D, Veltman JA, Vriend G, Gilissen C.\ \ MetaDome: Pathogenicity analysis of genetic variants through aggregation of homologous human protein\ domains.\ Hum Mutat. 2019 Aug;40(8):1030-1038.\ PMID: 31116477; PMC: PMC6772141\
\ \\ Silk M, Petrovski S, Ascher DB.\ \ MTR-Viewer: identifying regions within genes under purifying selection.\ Nucleic Acids Res. 2019 Jul 2;47(W1):W121-W126.\ PMID: 31170280; PMC: PMC6602522\
\ \\ Halldorsson BV, Eggertsson HP, Moore KHS, Hauswedell H, Eiriksson O, Ulfarsson MO, Palsson G,\ Hardarson MT, Oddsson A, Jensson BO et al.\ \ The sequences of 150,119 genomes in the UK Biobank.\ Nature. 2022 Jul;607(7920):732-740.\ PMID: 35859178; PMC: PMC9329122\
\ \ \\ Huang YF, Gulko B, Siepel A.\ \ Fast, scalable prediction of deleterious noncoding variants from functional and population genomic\ data.\ Nat Genet. 2017 Apr;49(4):618-624.\ PMID: 28288115; PMC: PMC5395419\
\ \ phenDis 0 bigDataUrl /gbdb/hg38/ukbDepletion/ukbDepletion.bw\ html constraintSuper\ longLabel UK Biobank / deCODE Genetics Depletion Rank Score\ maxHeightPixels 128:40:8\ parent constraintSuper\ priority 10.5\ shortLabel UKB Depl. Rank Score\ track ukbDepletion\ type bigWig 0.0 1.0\ viewLimits 0.0:1.0\ viewLimitsMax 0:1.0\ visibility dense\ chainXenTro10 xenTro10 Chain chain xenTro10 X. tropicalis (Nov. 2019 (UCB_Xtro_10.0/xenTro10)) Chained Alignments 3 11 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel X. tropicalis (Nov. 2019 (UCB_Xtro_10.0/xenTro10)) Chained Alignments\ otherDb xenTro10\ parent vertebrateChainNetViewchain off\ shortLabel xenTro10 Chain\ subGroups view=chain species=s029b clade=c03\ track chainXenTro10\ type chain xenTro10\ chainRn6 Rat Chain chain rn6 Rat (Jul. 2014 (RGSC 6.0/rn6)) Chained Alignments 3 11 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Rat (Jul. 2014 (RGSC 6.0/rn6)) Chained Alignments\ otherDb rn6\ parent placentalChainNetViewchain off\ shortLabel Rat Chain\ subGroups view=chain species=s024d clade=c00\ track chainRn6\ type chain rn6\ chainNasLar1 Proboscis monkey Chain chain nasLar1 Proboscis monkey (Nov. 2014 (Charlie1.0/nasLar1)) Chained Alignments 3 11 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Proboscis monkey (Nov. 2014 (Charlie1.0/nasLar1)) Chained Alignments\ otherDb nasLar1\ parent primateChainNetViewchain off\ shortLabel Proboscis monkey Chain\ subGroups view=chain species=s016 clade=c01\ track chainNasLar1\ type chain nasLar1\ encTfChipPkENCFF646TUX A549 CTCF 3 narrowPeak Transcription Factor ChIP-seq Peaks of CTCF in A549 from ENCODE 3 (ENCFF646TUX) 0 11 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of CTCF in A549 from ENCODE 3 (ENCFF646TUX)\ parent encTfChipPk off\ shortLabel A549 CTCF 3\ subGroups cellType=A549 factor=CTCF\ track encTfChipPkENCFF646TUX\ cloneEndABC23 ABC23 bed 12 Agencourt fosmid library 23 0 11 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 23\ parent cloneEndSuper off\ priority 11\ shortLabel ABC23\ subGroups source=agencourt\ track cloneEndABC23\ type bed 12\ visibility hide\ AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep1LK7_CNhs13341_ctss_fwd AorticSmsToFgf2_00hr30minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep1 (LK7)_CNhs13341_12644-134G7_forward 0 11 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12644-134G7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr30min%2c%20biol_rep1%20%28LK7%29.CNhs13341.12644-134G7.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep1 (LK7)_CNhs13341_12644-134G7_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12644-134G7 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr30minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep1LK7_CNhs13341_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12644-134G7\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep1LK7_CNhs13341_tpm_fwd AorticSmsToFgf2_00hr30minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep1 (LK7)_CNhs13341_12644-134G7_forward 1 11 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12644-134G7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr30min%2c%20biol_rep1%20%28LK7%29.CNhs13341.12644-134G7.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep1 (LK7)_CNhs13341_12644-134G7_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12644-134G7 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr30minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep1LK7_CNhs13341_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12644-134G7\ urlLabel FANTOM5 Details:\ bismap50Quantitative Bismap M50 bigWig 0.02 1.00 Multi-read mappability with 50-mers after bisulfite conversion 0 11 240 120 80 247 187 167 0 0 0 map 0 bigDataUrl /gbdb/hg38/hoffmanMappability/k50.Bismap.MultiTrackMappability.bw\ color 240,120,80\ longLabel Multi-read mappability with 50-mers after bisulfite conversion\ parent bismapBigWig off\ priority 11\ shortLabel Bismap M50\ subGroups view=MR\ track bismap50Quantitative\ type bigWig 0.02 1.00\ visibility hide\ wgEncodeReg4AtacAllBloodVessel Blood vessel (all biosamples) bigWig Avg. ATAC level of 4 blood vessel experiments (all biosamples) 0 11 255 37 41 255 146 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodVesselATAC.bw\ color 255,37,41\ longLabel Avg. ATAC level of 4 blood vessel experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 11\ shortLabel Blood vessel (all biosamples)\ track wgEncodeReg4AtacAllBloodVessel\ type bigWig\ gtexCovBrainCerebellum Brain Cereb bigWig Brain Cerebellum 0 11 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-145MH-2926-SM-5Q5D2.Brain_Cerebellum.RNAseq.bw\ color 238,238,0\ longLabel Brain Cerebellum\ parent gtexCov\ shortLabel Brain Cereb\ track gtexCovBrainCerebellum\ wgEncodeReg4TxnBreastPlus Breast + bigWig Avg. + strand total RNA-seq level of 4 breast experiments (tissues and primary cells only) 0 11 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpBreastPlus.bw\ color 65,171,173\ longLabel Avg. + strand total RNA-seq level of 4 breast experiments (tissues and primary cells only)\ parent wgEncodeReg4Txn off\ priority 11\ shortLabel Breast +\ track wgEncodeReg4TxnBreastPlus\ type bigWig\ dbVar_common_african dbVar Curated African SVs bigBed 9 + . NCBI dbVar Curated Common SVs: African 3 11 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/variants/$$ varRep 1 bigDataUrl /gbdb/hg38/bbi/dbVar/common_african.bb\ longLabel NCBI dbVar Curated Common SVs: African\ parent dbVar_common on\ priority 11\ shortLabel dbVar Curated African SVs\ track dbVar_common_african\ type bigBed 9 + .\ url https://www.ncbi.nlm.nih.gov/dbvar/variants/$$\ urlLabel NCBI Variant Page:\ ENCFF735XLO_ENCFF970LMB_ENCFF481LLD_ENCFF838OJW ENCFF735XLO_ENCFF970LMB_ENCFF481LLD_ENCFF838OJW bigBed 9 + 5 MM.1S: (1) cCREs 4 11 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF735XLO_ENCFF970LMB_ENCFF481LLD_ENCFF838OJW.bb\ longLabel MM.1S: (1) cCREs\ mouseOver ID: ${name}\ The GenomeAsia 100K project aims\ to sequence 100,000 Asian individuals. This pilot release (GAsP) contains whole-genome sequencing\ data of 1,739 individuals from 219 population groups across Asia. Frequencies are broken down by\ Northeast Asian, Southeast Asian, and South Asian ancestry groups. The data is split into two\ subtracks: substitutions and indels.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is gasp.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original VCFs are also available from the\ GenomeAsia 100K\ website. No license nor login is required.\
\ \\ Samples were sequenced on Illumina HiSeq 2500, HiSeq 4000, and HiSeq X Ten instruments with\ 2×100 bp or 2×150 bp paired-end reads at an average depth of 36x. Reads were aligned to\ GRCh37 using BWA-MEM. Duplicate reads were marked with SAMBLASTER and sorted with Sambamba.\ Per-sample variant calling was performed with GATK HaplotypeCaller in GVCF mode, followed by\ joint genotyping with GenotypeGVCFs. Variant quality score recalibration (VQSR) was applied at\ a 99% sensitivity tranche for both SNPs and indels. Sample-level QC included contamination\ checks with verifyBamID and sex concordance verification. The final callset contains\ ∼65 million variants across 1,739 individuals from 219 populations.\
\\ The upstream callset is on GRCh37. We lifted it to hg38 using\ CrossMap and the UCSC\ hg19ToHg38 chain file. After lifting, variants that landed on alt, random, fix, or\ unplaced contigs were dropped, and the result was sorted and indexed with tabix.\
\\ The makeDoc file documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were needed and are also available from GitHub.\
\ \\ GenomeAsia100K Consortium.\ \ The GenomeAsia 100K Project enables genetic discoveries across Asia.\ Nature. 2019 Dec;576(7785):106-111.\ PMID: 31802016; PMC: PMC7054211\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/ga100k/ga100k.subst.vcf.gz\ dataVersion Pilot 2019 (lifted to hg38, May 2026)\ longLabel SNV Frequencies: GenomeAsia Pilot - Substitutions\ parent varFreqs on\ priority 11\ shortLabel GenomeAsia 1.7k SNVs\ track gasp\ type vcfTabix\ visibility hide\ chainHprcGCA_018506125v1 HG02055.mat chain GCA_018506125.1 HG02055.mat HG02055.pri.mat.f1_v2 (May 2021 GCA_018506125.1_HG02055.pri.mat.f1_v2) HPRC project computed Chained Alignments 3 11 0 0 0 255 255 0 1 0 0 hprc 1 longLabel HG02055.mat HG02055.pri.mat.f1_v2 (May 2021 GCA_018506125.1_HG02055.pri.mat.f1_v2) HPRC project computed Chained Alignments\ otherDb GCA_018506125.1\ parent hprcChainNetViewchain off\ priority 28\ shortLabel HG02055.mat\ subGroups view=chain sample=s028 population=afr subpop=acb hap=mat\ track chainHprcGCA_018506125v1\ type chain GCA_018506125.1\ HNSC HNSC bigLolly 12 + Head and Neck squamous cell carcinoma 0 11 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/HNSC.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Head and Neck squamous cell carcinoma\ parent gdcCancer off\ priority 11\ shortLabel HNSC\ track HNSC\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4DnaseKidney Kidney bigWig Avg. DNase level of 78 kidney experiments (tissues and primary cells only) 0 11 92 161 153 173 208 204 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpKidneyDNase.bw\ color 92,161,153\ longLabel Avg. DNase level of 78 kidney experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase\ priority 11\ shortLabel Kidney\ track wgEncodeReg4DnaseKidney\ type bigWig\ lincRNAsCTKidney Kidney bed 5 + lincRNAs from kidney 1 11 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from kidney\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Kidney\ subGroups view=lincRNAsRefseqExp tissueType=kidney\ track lincRNAsCTKidney\ wgEncodeReg4MarkH3k4me3Liver Liver bigWig Avg. H3K4me3 level of 5 liver experiments (tissues and primary cells only) 0 11 137 152 82 196 203 168 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLiverH3K4me3.bw\ color 137,152,82\ longLabel Avg. H3K4me3 level of 5 liver experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3\ priority 11\ shortLabel Liver\ track wgEncodeReg4MarkH3k4me3Liver\ type bigWig\ wgEncodeReg4MarkH3k27acLung Lung bigWig Avg. H3K27ac level of 11 lung experiments (tissues and primary cells only) 2 11 130 163 45 192 209 150 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLungH3K27ac.bw\ color 130,163,45\ longLabel Avg. H3K27ac level of 11 lung experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 11\ shortLabel Lung\ track wgEncodeReg4MarkH3k27acLung\ type bigWig\ wgEncodeReg4MarkCtcfMuscle Muscle bigWig Avg. CTCF level of 14 muscle experiments (tissues and primary cells only) 0 11 137 135 170 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpMuscleCTCF.bw\ color 137,135,170\ longLabel Avg. CTCF level of 14 muscle experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf\ priority 11\ shortLabel Muscle\ track wgEncodeReg4MarkCtcfMuscle\ type bigWig\ neuron0TH Neuron - Z000000TH bigWig Methylation Atlas: Neuron - Z000000TH 2 11 138 43 226 196 149 240 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/neuron0TH.bw\ color 138,43,226\ longLabel Methylation Atlas: Neuron - Z000000TH\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 11\ shortLabel Neuron - Z000000TH\ subGroups cellType=Neuron dataType=Replicate\ track neuron0TH\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ wgEncodeRegDnaseUwNt2d1Peak NT2-D1 Pk narrowPeak NT2-D1 embryonal carcinoma (NTera2) cell line DNaseI Peaks from ENCODE 1 11 255 173 85 255 214 170 1 0 0 regulation 1 color 255,173,85\ longLabel NT2-D1 embryonal carcinoma (NTera2) cell line DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel NT2-D1 Pk\ subGroups view=a_Peaks cellType=NT2-D1 treatment=n_a tissue=testis cancer=cancer\ track wgEncodeRegDnaseUwNt2d1Peak\ wgEncodeRegDnaseUwNt2d1Wig NT2-D1 Sg bigWig 0 8351.64 NT2-D1 embryonal carcinoma (NTera2) cell line DNaseI Signal from ENCODE 0 11 255 173 85 255 214 170 0 0 0 regulation 1 color 255,173,85\ longLabel NT2-D1 embryonal carcinoma (NTera2) cell line DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.09484\ shortLabel NT2-D1 Sg\ subGroups cellType=NT2-D1 treatment=n_a tissue=testis cancer=cancer\ table wgEncodeRegDnaseUwNt2d1Signal\ track wgEncodeRegDnaseUwNt2d1Wig\ type bigWig 0 8351.64\ unipOther Other Annot. bigBed 12 + UniProt Other Annotations 1 11 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipOther.bb\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ longLabel UniProt Other Annotations\ mouseOver UniProt record: $uniProtId\ A reprocessed callset by the gnomAD project combining the 1000 Genomes and Human Genome Diversity Project\ (HGDP) data, with 4,094 whole genomes from 80 populations. The dataset includes per-population\ allele frequencies for all 80 populations as well as broad continental groupings from gnomAD\ (African, Admixed American, East Asian, European, Middle Eastern, South Asian, and others).\
\ \\ This track shows allele frequencies only. The full phased genotype data with haplotype\ clustering display is available in the\ gnomAD HGDP+1000G track under Phased Variants.\ The track here does not include the full variant frequencies for all subpopulations, instead, \ it aggregates frequencies to the main groups, AFR, AMI, AMR, ASJ, EAS, FIN, MID, NFE, OTH, SAS. \ To access the full frequency information, use the track under "Phased Variants".\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is hgdp1kFreq.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original VCFs with full genotypes can also be downloaded from\ gnomAD Downloads.\
\ \\ The gnomAD project reprocessed 4,094 whole genomes from the 1000 Genomes Project and the Human\ Genome Diversity Project (HGDP) through a unified pipeline. Sequencing was performed on Illumina\ platforms at a mean coverage of 32–34x. Reads were aligned to GRCh38 (hs38DH reference with\ decoy and HLA sequences) using BWA-MEM 0.7.15. Variant calling followed GATK best practices:\ per-sample calls with GATK 3.5 HaplotypeCaller, then joint genotyping with GATK4 through\ the Hail VCF combiner, which scales the merge step. Allele-specific variant quality score recalibration\ (AS-VQSR) was applied for both SNPs and indels. Sample QC included contamination estimates\ (verifyBamID), sex concordance, relatedness filters (PC-Relate), and population assignment\ with PCA against gnomAD reference panels. Per-population allele frequencies were computed for\ 80 fine-grained populations and for broad continental groups.\
\\ The makeDoc file documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were also needed and are available from GitHub.\
\ \\ Thanks to the gnomAD team at the Broad Institute for harmonizing and making this dataset\ publicly available, and to all participants of the 1000 Genomes Project and the Human Genome\ Diversity Project.\
\ \\ Koenig Z, Yohannes MT, Nkambule LL, Zhao X, Goodrich JK, Kim HA, Wilson MW, Tiao G, Hao SP, Sahakian\ N et al.\ \ A harmonized public resource of deeply sequenced diverse human genomes.\ Genome Res. 2024 Jun 25;34(5):796-809.\ PMID: 38749656; PMC: PMC11216312\
\ \\ Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J\ et al.\ \ Insights into human genetic variation and population history from 929 diverse genomes.\ Science. 2020 Mar 20;367(6484).\ PMID: 32193295; PMC: PMC7115999\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/hgdp1kFreq/hgdp1k.freq.vcf.gz\ dataVersion v3.1.2\ longLabel SNV Frequencies: gnomAD HGDP + 1000 Genomes - 4,094 WGS, 80 populations\ parent varFreqs on\ priority 12\ shortLabel gnomAD HGDP+1kG 4k WGS\ track hgdp1kFreq\ type vcfTabix\ visibility hide\ netHprcGCA_018506125v1 HG02055.mat netAlign GCA_018506125.1 chainHprcGCA_018506125v1 HG02055.mat HG02055.pri.mat.f1_v2 (May 2021 GCA_018506125.1_HG02055.pri.mat.f1_v2) HPRC project computed Chain Nets 1 12 0 0 0 255 255 0 0 0 0 hprc 0 longLabel HG02055.mat HG02055.pri.mat.f1_v2 (May 2021 GCA_018506125.1_HG02055.pri.mat.f1_v2) HPRC project computed Chain Nets\ otherDb GCA_018506125.1\ parent hprcChainNetViewnet off\ priority 28\ shortLabel HG02055.mat\ subGroups view=net sample=s028 population=afr subpop=acb hap=mat\ track netHprcGCA_018506125v1\ type netAlign GCA_018506125.1 chainHprcGCA_018506125v1\ wgEncodeReg4DnaseLargeIntestine Large intestine bigWig Avg. DNase level of 28 large intestine experiments (tissues and primary cells only) 0 12 86 86 36 170 170 145 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLargeIntestineDNase.bw\ color 86,86,36\ longLabel Avg. DNase level of 28 large intestine experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 12\ shortLabel Large intestine\ track wgEncodeReg4DnaseLargeIntestine\ type bigWig\ lincRNAsCTLiver Liver bed 5 + lincRNAs from liver 1 12 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from liver\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Liver\ subGroups view=lincRNAsRefseqExp tissueType=liver\ track lincRNAsCTLiver\ wgEncodeReg4MarkH3k4me3Lung Lung bigWig Avg. H3K4me3 level of 17 lung experiments (tissues and primary cells only) 0 12 130 163 45 192 209 150 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLungH3K4me3.bw\ color 130,163,45\ longLabel Avg. H3K4me3 level of 17 lung experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 12\ shortLabel Lung\ track wgEncodeReg4MarkH3k4me3Lung\ type bigWig\ wgEncodeReg4MarkH3k27acMuscle Muscle bigWig Avg. H3K27ac level of 20 muscle experiments (tissues and primary cells only) 2 12 137 135 170 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpMuscleH3K27ac.bw\ color 137,135,170\ longLabel Avg. H3K27ac level of 20 muscle experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac\ priority 12\ shortLabel Muscle\ track wgEncodeReg4MarkH3k27acMuscle\ type bigWig\ oligodendMerged Oligodendrocytes Merged bigWig Methylation Atlas: Oligodendrocytes Merged Samples 2 12 148 103 189 201 179 222 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/oligodendMerged.bw\ color 148,103,189\ longLabel Methylation Atlas: Oligodendrocytes Merged Samples\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 12\ shortLabel Oligodendrocytes Merged\ subGroups cellType=Oligodend dataType=Merged\ track oligodendMerged\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ wgEncodeReg4MarkCtcfPancreas Pancreas bigWig Avg. CTCF level of 9 pancreas experiments (tissues and primary cells only) 0 12 175 100 41 215 177 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPancreasCTCF.bw\ color 175,100,41\ longLabel Avg. CTCF level of 9 pancreas experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf off\ priority 12\ shortLabel Pancreas\ track wgEncodeReg4MarkCtcfPancreas\ type bigWig\ unipRepeat Repeats bigBed 12 + UniProt Repeats 1 12 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipRepeat.bb\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ longLabel UniProt Repeats\ mouseOver UniProt record: $uniProtId\ The GREGoR Consortium\ (Genomics Research to Elucidate the Genetics of Rare diseases) is a\ National Human Genome Research Institute (NHGRI)-funded research consortium\ that works to identify the genetic basis of currently unexplained rare diseases.\ GREGoR is a collaboration between multiple research centers and a data coordinating center\ that apply genomic technologies to rare disease cohorts.\
\ \\ This track shows allele frequencies from the GREGoR Release 4 (R04, October 2025)\ joint variant callset of a subset of the 10,683 participants across\ 4,366 families. The joint callset includes only the 8,161 short-read whole-genome sequencing (WGS)\ samples, or a subset of these. The GREGoR site does not specify how many samples exactly are part\ of the joint callset. The callset does not include any of the 2,629 whole-exome sequencing (WES) samples.\ GREGoR also provides some long-read WGS, RNA-seq, and ATAC-seq, but these were not used for the joint callset either.\ The VCF shown here contains variant calls with VEP consequence annotations.\ The INFO fields include allele count (AC), allele frequency (AF), allele number (AN),\ and counts broken down by affected status (AC_AFFECTED, AC_UNAFFECTED, AC_UNKNOWN).\
\ \\ This is a VCF track. When zoomed in, variants are displayed with base-specific coloring.\ Mouseover shows the variant position, alleles, and allele frequency.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is gregor.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The full controlled-access GREGoR data is available through the\ AnVIL platform\ via controlled dbGaP access. More information on data access is available at\ the GREGoR data page.\
\ \The sample processing methods of the GREGoR project depend on the sequencing center, see the Methods document below for details.\ The GREGoR R04 joint callset for short-read whole genome sequencing (srWGS) was generated\ by the GREGoR Data Coordinating Center (DCC) through a two-stage harmonization and joint\ genotyping pipeline. Raw srWGS data from GREGoR Consortium Research Centers were uniformly\ reprocessed with the Whole Genome Germline Single Sample WARP pipeline in DRAGEN-GATK\ mode (v3.1.6). This pipeline aligns reads to the GRCh38 reference genome\ (GCA_000001405.15_GRCh38_no_alt_analysis_set) with the DRAGMAP aligner, marks duplicates\ with Picard v2.26.10, and performs single-sample variant calling with GATK HaplotypeCaller using the\ DragSTR model with hard filtering. The output is per-sample gVCFs. Joint variant calling across all\ harmonized samples was then performed with the Genomic Variant Store (GVS), a\ scalable cloud-native joint genotyping pipeline developed for large cohort analysis in which\ variants are ingested into a query-optimized store and rendered to a multi-sample variant file\ format. The resulting joint callset was functionally annotated with Ensembl Variant Effect Predictor (VEP) v112.\ Full methods, with per-site library preparation and bioinformatics pipelines for independently\ processed samples, are available in the\ GREGoR R04 Methods document.\
\\ At UCSC, site VCF files were downloaded from GREGoR's Google Drive. \ The VCFs were merged with bcftools.\ We provide documentation that indicates how all source files of the varFreqs track were converted in the makeDoc file of the track. \ For some tracks, python scripts were necessary and are also available from GitHub\
\ \\ The GREGoR Consortium is supported by the National Human Genome Research Institute (NHGRI).\ We thank the participants, their families, and the consortium for making this data available.\ For more information, see the\ GREGoR About page.\
\ \\ The GREGoR Consortium does not yet have a peer-reviewed flagship publication describing the\ R04 release. For methods and data access details, see the\ GREGoR R04 Methods document and the\ GREGoR Consortium website.\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/gregor/gregor.vcf.gz\ dataVersion R04 (Oct 2025)\ longLabel SNV Frequencies: GREGoR Consortium - Release 4, 3,624 WGS samples, rare disease families\ parent varFreqs on\ priority 13\ shortLabel GREGoR R4 3.6k WGS\ track gregor\ type vcfTabix\ visibility hide\ chainHprcGCA_018852585v1 HG02145.mat chain GCA_018852585.1 HG02145.mat HG02145.pri.mat.f1_v2 (Jun. 2021 GCA_018852585.1_HG02145.pri.mat.f1_v2) HPRC project computed Chained Alignments 3 13 0 0 0 255 255 0 1 0 0 hprc 1 longLabel HG02145.mat HG02145.pri.mat.f1_v2 (Jun. 2021 GCA_018852585.1_HG02145.pri.mat.f1_v2) HPRC project computed Chained Alignments\ otherDb GCA_018852585.1\ parent hprcChainNetViewchain off\ priority 29\ shortLabel HG02145.mat\ subGroups view=chain sample=s029 population=afr subpop=acb hap=mat\ track chainHprcGCA_018852585v1\ type chain GCA_018852585.1\ KIRC KIRC bigLolly 12 + Kidney renal clear cell carcinoma 0 13 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/KIRC.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Kidney renal clear cell carcinoma\ parent gdcCancer off\ priority 13\ shortLabel KIRC\ track KIRC\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4DnaseLiver Liver bigWig Avg. DNase level of 10 liver experiments (tissues and primary cells only) 0 13 137 152 82 196 203 168 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLiverDNase.bw\ color 137,152,82\ longLabel Avg. DNase level of 10 liver experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase\ priority 13\ shortLabel Liver\ track wgEncodeReg4DnaseLiver\ type bigWig\ lincRNAsCTLung Lung bed 5 + lincRNAs from lung 1 13 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from lung\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Lung\ subGroups view=lincRNAsRefseqExp tissueType=lung\ track lincRNAsCTLung\ wgEncodeReg4MarkH3k4me3Muscle Muscle bigWig Avg. H3K4me3 level of 24 muscle experiments (tissues and primary cells only) 0 13 137 135 170 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpMuscleH3K4me3.bw\ color 137,135,170\ longLabel Avg. H3K4me3 level of 24 muscle experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3\ priority 13\ shortLabel Muscle\ track wgEncodeReg4MarkH3k4me3Muscle\ type bigWig\ oligo0TK Oligodendrocytes - Z000000TK bigWig Methylation Atlas: Oligodendrocytes - Z000000TK 2 13 148 103 189 201 179 222 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/oligo0TK.bw\ color 148,103,189\ longLabel Methylation Atlas: Oligodendrocytes - Z000000TK\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 13\ shortLabel Oligodendrocytes - Z000000TK\ subGroups cellType=Oligodend dataType=Replicate\ track oligo0TK\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ wgEncodeReg4MarkH3k27acPancreas Pancreas bigWig Avg. H3K27ac level of 12 pancreas experiments (tissues and primary cells only) 2 13 175 100 41 215 177 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPancreasH3K27ac.bw\ color 175,100,41\ longLabel Avg. H3K27ac level of 12 pancreas experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 13\ shortLabel Pancreas\ track wgEncodeReg4MarkH3k27acPancreas\ type bigWig\ wgEncodeReg4MarkCtcfProstate Prostate bigWig Avg. CTCF level of 3 prostate experiments (tissues and primary cells only) 0 13 140 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpProstateCTCF.bw\ color 140,140,140\ longLabel Avg. CTCF level of 3 prostate experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf off\ priority 13\ shortLabel Prostate\ track wgEncodeReg4MarkCtcfProstate\ type bigWig\ unipConflict Seq. Conflicts bigBed 12 + UniProt Sequence Conflicts 1 13 0 0 0 127 127 127 0 0 0 genes 1 bigDataUrl /gbdb/hg38/uniprot/unipConflict.bb\ filterValues.status Manually reviewed (Swiss-Prot),Unreviewed (TrEMBL)\ longLabel UniProt Sequence Conflicts\ mouseOver UniProt record: $uniProtId\ The Haplotype Reference Consortium (HRC) is a collaboration among several\ large sequencing projects to create a reference panel for genotype imputation.\ Release 1.1 contains 64,976 haplotypes from 32,488 whole-genome sequenced samples at\ low coverage (average 7x), with 40 million variant sites (minimum allele count of 5).\
\\ The contributing studies include the 1000 Genomes Project, UK10K, and many other cohorts.\ Since 1000 Genomes data is already available as a separate track, this track shows only\ the frequencies from the non-1000 Genomes samples (~30,000 individuals). After the lift\ from GRCh37 to GRCh38, 38.3 million variants remain.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is hrc.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original site list file can also be downloaded from the\ HRC website.\ Our GitHub repo contains a\ script that converts the tab-separated file to VCF and lifts it to hg38.\
\ \\ The HRC r1.1 site list was downloaded from the\ HRC website\ as a tab-separated file on GRCh37, converted to VCF and lifted to GRCh38 with UCSC liftOver.\ Only frequencies from the non-1000 Genomes samples (~30,000 of the 32,488 total) are included,\ since 1000 Genomes data is available separately. Of 40.4M input variants, 8,052 were unmapped\ by liftOver and 2.1M were present only in 1000 Genomes samples and were dropped, leaving\ 38.3M variants.\ The conversion steps for all source files of the varFreqs track are documented in the makeDoc file of the track.\ Some tracks required python scripts, which are also available from GitHub.\
\ \\ Thanks to the Haplotype Reference Consortium and all contributing studies for making this\ reference panel publicly available.\
\ \\ McCarthy S, Das S, Kretzschmar W, Delaneau O, Wood AR, Teumer A, Kang HM, Fuchsberger C,\ Danecek P, Sharp K et al.\ \ A reference panel of 64,976 haplotypes for genotype imputation.\ Nat Genet. 2016 Oct;48(10):1279-83.\ PMID: 27548312; PMC: PMC5388176\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/hrc/hrc.vcf.gz\ dataVersion r1.1\ longLabel SNV Frequencies: Haplotype Reference Consortium - 30k WGS (excl. 1000 Genomes)\ parent varFreqs on\ priority 14\ shortLabel HRC 30k WGS\ track hrc\ type vcfTabix\ visibility hide\ KIRP KIRP bigLolly 12 + Kidney renal papillary cell carcinoma 0 14 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/KIRP.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Kidney renal papillary cell carcinoma\ parent gdcCancer off\ priority 14\ shortLabel KIRP\ track KIRP\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4DnaseLung Lung bigWig Avg. DNase level of 55 lung experiments (tissues and primary cells only) 0 14 130 163 45 192 209 150 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLungDNase.bw\ color 130,163,45\ longLabel Avg. DNase level of 55 lung experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 14\ shortLabel Lung\ track wgEncodeReg4DnaseLung\ type bigWig\ lincRNAsCTLymphNode LymphNode bed 5 + lincRNAs from lymphnode 1 14 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from lymphnode\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel LymphNode\ subGroups view=lincRNAsRefseqExp tissueType=lymphnode\ track lincRNAsCTLymphNode\ oligo42E Oligodendrocytes - Z0000042E bigWig Methylation Atlas: Oligodendrocytes - Z0000042E 2 14 148 103 189 201 179 222 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/oligo42E.bw\ color 148,103,189\ longLabel Methylation Atlas: Oligodendrocytes - Z0000042E\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 14\ shortLabel Oligodendrocytes - Z0000042E\ subGroups cellType=Oligodend dataType=Replicate\ track oligo42E\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ wgEncodeReg4MarkH3k4me3Pancreas Pancreas bigWig Avg. H3K4me3 level of 12 pancreas experiments (tissues and primary cells only) 0 14 175 100 41 215 177 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPancreasH3K4me3.bw\ color 175,100,41\ longLabel Avg. H3K4me3 level of 12 pancreas experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 14\ shortLabel Pancreas\ track wgEncodeReg4MarkH3k4me3Pancreas\ type bigWig\ wgEncodeReg4MarkH3k27acPenis Penis bigWig Avg. H3K27ac level of 2 penis experiments (tissues and primary cells only) 2 14 20 74 159 137 164 207 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPenisH3K27ac.bw\ color 20,74,159\ longLabel Avg. H3K27ac level of 2 penis experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 14\ shortLabel Penis\ track wgEncodeReg4MarkH3k27acPenis\ type bigWig\ wgEncodeReg4MarkCtcfSkin Skin bigWig Avg. CTCF level of 12 skin experiments (tissues and primary cells only) 0 14 127 133 209 191 194 232 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpSkinCTCF.bw\ color 127,133,209\ longLabel Avg. CTCF level of 12 skin experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf off\ priority 14\ shortLabel Skin\ track wgEncodeReg4MarkCtcfSkin\ type bigWig\ Agilent_Human_Exon_Focused_Regions SureSel. Focused T bigBed Agilent - SureSelect Focused Exome Target Regions 0 14 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S07084713_Regions.bb\ color 255,36,36\ longLabel Agilent - SureSelect Focused Exome Target Regions\ parent exomeProbesets off\ shortLabel SureSel. Focused T\ track Agilent_Human_Exon_Focused_Regions\ type bigBed\ chainCanFam4 Dog Chain chain canFam4 Dog (Mar. 2020 (UU_Cfam_GSD_1.0/canFam4)) Chained Alignments 3 15 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Dog (Mar. 2020 (UU_Cfam_GSD_1.0/canFam4)) Chained Alignments\ otherDb canFam4\ parent placentalChainNetViewchain off\ shortLabel Dog Chain\ subGroups view=chain species=s034c clade=c01\ track chainCanFam4\ type chain canFam4\ chainDanRer11 Zebrafish Chain chain danRer11 Zebrafish (May 2017 (GRCz11/danRer11)) Chained Alignments 3 15 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Zebrafish (May 2017 (GRCz11/danRer11)) Chained Alignments\ otherDb danRer11\ parent vertebrateChainNetViewchain off\ shortLabel Zebrafish Chain\ subGroups view=chain species=s043 clade=c06\ track chainDanRer11\ type chain danRer11\ chainMacFas5 Crab-eating macaque Chain chain macFas5 Crab-eating macaque (Jun. 2013 (Macaca_fascicularis_5.0/macFas5)) Chained Alignments 3 15 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Crab-eating macaque (Jun. 2013 (Macaca_fascicularis_5.0/macFas5)) Chained Alignments\ otherDb macFas5\ parent primateChainNetViewchain off\ shortLabel Crab-eating macaque Chain\ subGroups view=chain species=s020 clade=c01\ track chainMacFas5\ type chain macFas5\ encTfChipPkENCFF558UWY A549 ESRRA narrowPeak Transcription Factor ChIP-seq Peaks of ESRRA in A549 from ENCODE 3 (ENCFF558UWY) 0 15 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of ESRRA in A549 from ENCODE 3 (ENCFF558UWY)\ parent encTfChipPk off\ shortLabel A549 ESRRA\ subGroups cellType=A549 factor=ESRRA\ track encTfChipPkENCFF558UWY\ cloneEndABC8 ABC8 bed 12 Agencourt fosmid library 8 0 15 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 8\ parent cloneEndSuper off\ priority 15\ shortLabel ABC8\ subGroups source=agencourt\ track cloneEndABC8\ type bed 12\ visibility hide\ AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_ctss_fwd AorticSmsToFgf2_00hr30minBr3+ bigWig Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_forward 0 15 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr30min%2c%20biol_rep3%20%28LK9%29.CNhs13569.12840-137B5.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12840-137B5 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr30minBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_tpm_fwd AorticSmsToFgf2_00hr30minBr3+ bigWig Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_forward 1 15 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr30min%2c%20biol_rep3%20%28LK9%29.CNhs13569.12840-137B5.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12840-137B5 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr30minBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5\ urlLabel FANTOM5 Details:\ gtexCovBrainHippocampus Brain Hippocamp bigWig Brain Hippocampus 0 15 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-1HSKV-0011-R1b-SM-CMKH7.Brain_Hippocampus.RNAseq.bw\ color 238,238,0\ longLabel Brain Hippocampus\ parent gtexCov\ shortLabel Brain Hippocamp\ track gtexCovBrainHippocampus\ dbVar_common_south_asian dbVar Curated South Asian SVs bigBed 9 + . NCBI dbVar Curated Common SVs: South Asian 3 15 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/variants/$$ varRep 1 bigDataUrl /gbdb/hg38/bbi/dbVar/common_south_asian.bb\ longLabel NCBI dbVar Curated Common SVs: South Asian\ parent dbVar_common off\ priority 15\ shortLabel dbVar Curated South Asian SVs\ track dbVar_common_south_asian\ type bigBed 9 + .\ url https://www.ncbi.nlm.nih.gov/dbvar/variants/$$\ urlLabel NCBI Variant Page:\ wgEncodeReg4TxnEmbryoPlus Embryo + bigWig Avg. + strand total RNA-seq level of 1 embryo experiments (tissues and primary cells only) 0 15 118 158 101 186 206 178 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpEmbryoPlus.bw\ color 118,158,101\ longLabel Avg. + strand total RNA-seq level of 1 embryo experiments (tissues and primary cells only)\ parent wgEncodeReg4Txn off\ priority 15\ shortLabel Embryo +\ track wgEncodeReg4TxnEmbryoPlus\ type bigWig\ ENCFF022SDS_ENCFF132YWJ_ENCFF118EKX_ENCFF857NIC ENCFF022SDS_ENCFF132YWJ_ENCFF118EKX_ENCFF857NIC bigBed 9 + 5 Ascending aorta, female adult (53 years): (1) cCREs 4 15 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF022SDS_ENCFF132YWJ_ENCFF118EKX_ENCFF857NIC.bb\ longLabel Ascending aorta, female adult (53 years): (1) cCREs\ mouseOver ID: ${name}\ The GenomeIndia\ project is a national initiative that coordinates academic and medical institutions\ across India to characterize the genetic diversity of the Indian subcontinent. The\ release used by this track is whole-genome sequencing of 9,768 healthy adults\ sampled from 83 anthropologically defined endogamous populations across India's\ ethnolinguistic and biogeographic range (Indo-European, Dravidian, Austroasiatic,\ and Tibeto-Burman language families, plus a continentally admixed outgroup). After\ joint genotyping and quality filtering, 129,938,889 high-confidence biallelic\ variants (~121M SNVs and ~8M indels) were reported, of which roughly one third are\ absent from gnomAD, 1000 Genomes, and GenomeAsia. This track shows the alternate\ allele frequency in that 9,768-sample autosomal call set.\
\\ Indian populations are underrepresented in global variant\ databases, so many globally rare alleles are at much higher frequencies in specific\ endogamous groups. The release ships only the cohort-wide alternate allele\ frequency (no per-population breakdown), so this track shows the overall\ GenomeIndia AF; AC is derived from AF (see Methods).\
\ \\ Variants are shown as a VCF dense track. Each row reports the genomic position,\ ref/alt alleles, the GenomeIndia alternate allele frequency, and a synthesized\ allele count. The track only includes autosomal variants (chr1–chr22); chrX,\ chrY, and chrM are not in the current release.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API\ can be used; the track name is genomeindia.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original per-chromosome TSV summary statistics can be downloaded directly from\ the GenomeIndia Data Centre at ibdc.dbtindia.gov.in\ (the 9768GI_SummaryStats.tar.gz bundle). Use of the data is subject to\ the GenomeIndia data-access policy listed on that page.\
\ \\ PCR-free whole-genome sequencing libraries were prepared from blood-derived DNA and\ sequenced on Illumina NovaSeq 6000 to a per-sample average depth of at least 23×.\ Reads were processed with the Illumina DRAGEN v4.0.3 germline pipeline against\ GRCh38. The resulting per-sample gVCFs were then joint-genotyped with the Illumina\ gVCF genotyper. Site-level filters retained only PASS variants with\ QUAL ≥ 30, posterior genotype probability ≥ 99.9%, GQ > 20\ at every site (GQ > 40 for singletons and doubletons), heterozygous allele\ balance ≥ 0.2, call rate ≥ 98%, and Hardy–Weinberg equilibrium\ p > 1×10-11; sites with an inbreeding coefficient of 1\ were also excluded as technical artefacts. Variants were annotated for protein\ impact with Ensembl VEP v113 plus LOFTEE; details are in the published methods\ (Bhattacharyya et al. 2025, see References).\
\\ The release was downloaded from\ ibdc.dbtindia.gov.in as 9768GI_SummaryStats.tar.gz, which\ contains 22 per-chromosome TSV files of CHROM, POS, ID, REF, ALT, AF (no header).\ The TSV files were converted to a single sorted, bgzipped, tabix-indexed VCF by the\ script genomeindiaToVcf.py. The release ships only AF; AC and AN are\ synthesized as AN = 2 × 9768 = 19536 and\ AC = round(AF × AN). Variants were kept only\ when called in ≥98% of samples, so AN slightly overstates the true called allele\ count for some sites (worst case ~2%); the AC field is a\ close approximation, not the exact observed count. The processing\ steps are documented in the makeDoc file.\
\ \\ We thank the GenomeIndia consortium for making the 9,768-sample summary statistics\ publicly available. The track was built at UCSC by Max Haeussler.\
\ \\ Bhattacharyya C, Subramanian K, Uppili B, Biswas NK, Ramdas S, Tallapaka KB, Arvind P, Rupanagudi\ KV, Maitra A, Nagabandi T et al.\ \ Mapping genetic diversity with the GenomeIndia project.\ Nat Genet. 2025 Apr;57(4):767-773.\ PMID: 40200122\
\ \\ Subramanian K, Bhattacharyya C, Machha P, Mukherjee A, Tripathi D, Chakraborty S, Majumdar SS,\ Sengupta S, Singh P, More V et al; GenomeIndia Consortium.\ \ An Atlas of Indian Genetic Diversity.\ medRxiv. 2026 Mar 20;2026.03.20.26348801 (preprint).\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_genomeindia/genomeindia.vcf.gz\ dataVersion 9768GI_SummaryStats (Apr 2025)\ longLabel SNV Frequencies: GenomeIndia - 9,768 WGS, 83 populations (Bhattacharyya 2025)\ parent varFreqs on\ priority 15\ shortLabel India GenomeIndia 9.7k WGS\ tableBrowser off\ track genomeindia\ type vcfTabix\ visibility hide\ LAML LAML bigLolly 12 + Acute Myeloid Leukemia 0 15 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/LAML.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Acute Myeloid Leukemia\ parent gdcCancer off\ priority 15\ shortLabel LAML\ track LAML\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4DnaseLymphoidTissue Lymphoid tissue bigWig DNase level of 1 lymphoid tissue experiment (tissues and primary cells only) 0 15 130 141 158 192 198 206 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpLymphoidTissueDNase.bw\ color 130,141,158\ longLabel DNase level of 1 lymphoid tissue experiment (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 15\ shortLabel Lymphoid tissue\ track wgEncodeReg4DnaseLymphoidTissue\ type bigWig\ wgEncodeReg4AtacAllMuscle Muscle (all biosamples) bigWig Avg. ATAC level of 12 muscle experiments (all biosamples) 0 15 137 135 170 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/muscleATAC.bw\ color 137,135,170\ longLabel Avg. ATAC level of 12 muscle experiments (all biosamples)\ parent wgEncodeReg4Atac\ priority 15\ shortLabel Muscle (all biosamples)\ track wgEncodeReg4AtacAllMuscle\ type bigWig\ oligo42L Oligodendrocytes - Z0000042L bigWig Methylation Atlas: Oligodendrocytes - Z0000042L 2 15 148 103 189 201 179 222 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/oligo42L.bw\ color 148,103,189\ longLabel Methylation Atlas: Oligodendrocytes - Z0000042L\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 15\ shortLabel Oligodendrocytes - Z0000042L\ subGroups cellType=Oligodend dataType=Replicate\ track oligo42L\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ lincRNAsCTOvary Ovary bed 5 + lincRNAs from ovary 1 15 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from ovary\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Ovary\ subGroups view=lincRNAsRefseqExp tissueType=ovary\ track lincRNAsCTOvary\ wgEncodeReg4MarkH3k4me3Penis Penis bigWig Avg. H3K4me3 level of 3 penis experiments (tissues and primary cells only) 0 15 20 74 159 137 164 207 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPenisH3K4me3.bw\ color 20,74,159\ longLabel Avg. H3K4me3 level of 3 penis experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 15\ shortLabel Penis\ track wgEncodeReg4MarkH3k4me3Penis\ type bigWig\ wgEncodeReg4MarkH3k27acProstate Prostate bigWig Avg. H3K27ac level of 3 prostate experiments (tissues and primary cells only) 2 15 140 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpProstateH3K27ac.bw\ color 140,140,140\ longLabel Avg. H3K27ac level of 3 prostate experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 15\ shortLabel Prostate\ track wgEncodeReg4MarkH3k27acProstate\ type bigWig\ Agilent_Human_Exon_V4_Covered SureSel. V4+UTR P bigBed Agilent - SureSelect All Exon V4 + UTRs Covered by Probes 0 15 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S04380110_Covered.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V4 + UTRs Covered by Probes\ parent exomeProbesets off\ shortLabel SureSel. V4+UTR P\ track Agilent_Human_Exon_V4_Covered\ type bigBed\ wgEncodeReg4MarkCtcfUterus Uterus bigWig Avg. CTCF level of 2 uterus experiments (tissues and primary cells only) 0 15 186 111 165 220 183 210 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpUterusCTCF.bw\ color 186,111,165\ longLabel Avg. CTCF level of 2 uterus experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkCtcf off\ priority 15\ shortLabel Uterus\ track wgEncodeReg4MarkCtcfUterus\ type bigWig\ wgEncodeRegDnaseUwWi38Peak WI-38 Pk narrowPeak WI-38 embryonic lung fibroblast cell line DNaseI Peaks from ENCODE 1 15 255 192 85 255 223 170 1 0 0 regulation 1 color 255,192,85\ longLabel WI-38 embryonic lung fibroblast cell line DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel WI-38 Pk\ subGroups view=a_Peaks cellType=WI-38 treatment=n_a tissue=lung cancer=normal\ track wgEncodeRegDnaseUwWi38Peak\ wgEncodeRegDnaseUwWi38Wig WI-38 Sg bigWig 0 21133.7 WI-38 embryonic lung fibroblast cell line DNaseI Signal from ENCODE 0 15 255 192 85 255 223 170 0 0 0 regulation 1 color 255,192,85\ longLabel WI-38 embryonic lung fibroblast cell line DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.13075\ shortLabel WI-38 Sg\ subGroups cellType=WI-38 treatment=n_a tissue=lung cancer=normal\ table wgEncodeRegDnaseUwWi38Signal\ track wgEncodeRegDnaseUwWi38Wig\ type bigWig 0 21133.7\ netCanFam4 Dog Net netAlign canFam4 chainCanFam4 Dog (Mar. 2020 (UU_Cfam_GSD_1.0/canFam4)) Alignment Net 1 16 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Dog (Mar. 2020 (UU_Cfam_GSD_1.0/canFam4)) Alignment Net\ otherDb canFam4\ parent placentalChainNetViewnet on\ shortLabel Dog Net\ subGroups view=net species=s034c clade=c01\ track netCanFam4\ type netAlign canFam4 chainCanFam4\ netDanRer11 Zebrafish Net netAlign danRer11 chainDanRer11 Zebrafish (May 2017 (GRCz11/danRer11)) Alignment Net 1 16 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Zebrafish (May 2017 (GRCz11/danRer11)) Alignment Net\ otherDb danRer11\ parent vertebrateChainNetViewnet on\ shortLabel Zebrafish Net\ subGroups view=net species=s043 clade=c06\ track netDanRer11\ type netAlign danRer11 chainDanRer11\ netMacFas5 Crab-eating macaque Net netAlign macFas5 chainMacFas5 Crab-eating macaque (Jun. 2013 (Macaca_fascicularis_5.0/macFas5)) Alignment Net 1 16 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Crab-eating macaque (Jun. 2013 (Macaca_fascicularis_5.0/macFas5)) Alignment Net\ otherDb macFas5\ parent primateChainNetViewnet off\ shortLabel Crab-eating macaque Net\ subGroups view=net species=s020 clade=c01\ track netMacFas5\ type netAlign macFas5 chainMacFas5\ encTfChipPkENCFF896WFR A549 ETS1 narrowPeak Transcription Factor ChIP-seq Peaks of ETS1 in A549 from ENCODE 3 (ENCFF896WFR) 0 16 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of ETS1 in A549 from ENCODE 3 (ENCFF896WFR)\ parent encTfChipPk off\ shortLabel A549 ETS1\ subGroups cellType=A549 factor=ETS1\ track encTfChipPkENCFF896WFR\ cloneEndABC9 ABC9 bed 12 Agencourt fosmid library 9 0 16 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Agencourt fosmid library 9\ parent cloneEndSuper off\ priority 16\ shortLabel ABC9\ subGroups source=agencourt\ track cloneEndABC9\ type bed 12\ visibility hide\ wgEncodeReg4MarkCtcfAllAdipose Adipose (all biosamples) bigWig Avg. CTCF level of 4 adipose experiments (all biosamples) 0 16 255 119 39 255 187 147 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adiposeCTCF.bw\ color 255,119,39\ longLabel Avg. CTCF level of 4 adipose experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 16\ shortLabel Adipose (all biosamples)\ track wgEncodeReg4MarkCtcfAllAdipose\ type bigWig\ AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_ctss_rev AorticSmsToFgf2_00hr30minBr3- bigWig Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_reverse 0 16 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr30min%2c%20biol_rep3%20%28LK9%29.CNhs13569.12840-137B5.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12840-137B5 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr30minBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_tpm_rev AorticSmsToFgf2_00hr30minBr3- bigWig Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_reverse 1 16 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr30min%2c%20biol_rep3%20%28LK9%29.CNhs13569.12840-137B5.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr30min, biol_rep3 (LK9)_CNhs13569_12840-137B5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12840-137B5 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr30minBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr30minBiolRep3LK9_CNhs13569_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12840-137B5\ urlLabel FANTOM5 Details:\ gtexCovBrainHypothalamus Brain Hypothal bigWig Brain Hypothalamus 0 16 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-T5JC-0011-R8A-SM-32PLM.Brain_Hypothalamus.RNAseq.bw\ color 238,238,0\ longLabel Brain Hypothalamus\ parent gtexCov\ shortLabel Brain Hypothal\ track gtexCovBrainHypothalamus\ dbVar_common_other dbVar Curated Other Pop SVs bigBed 9 + . NCBI dbVar Curated Common SVs: Other 3 16 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/dbvar/variants/$$ varRep 1 bigDataUrl /gbdb/hg38/bbi/dbVar/common_other.bb\ longLabel NCBI dbVar Curated Common SVs: Other\ parent dbVar_common off\ priority 16\ shortLabel dbVar Curated Other Pop SVs\ track dbVar_common_other\ type bigBed 9 + .\ url https://www.ncbi.nlm.nih.gov/dbvar/variants/$$\ urlLabel NCBI Variant Page:\ wgEncodeReg4TxnEmbryoMinus Embryo - bigWig Avg. - strand total RNA-seq level of 1 embryo experiments (tissues and primary cells only) 0 16 118 158 101 186 206 178 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpEmbryoMinus.bw\ color 118,158,101\ longLabel Avg. - strand total RNA-seq level of 1 embryo experiments (tissues and primary cells only)\ negateValues on\ parent wgEncodeReg4Txn off\ priority 16\ shortLabel Embryo -\ track wgEncodeReg4TxnEmbryoMinus\ type bigWig\ ENCFF383WYK_ENCFF811RQX_ENCFF130NUG_ENCFF341RAH ENCFF383WYK_ENCFF811RQX_ENCFF130NUG_ENCFF341RAH bigBed 9 + 5 Coronary artery, female adult (53 years): (1) cCREs 4 16 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF383WYK_ENCFF811RQX_ENCFF130NUG_ENCFF341RAH.bb\ longLabel Coronary artery, female adult (53 years): (1) cCREs\ mouseOver ID: ${name}\ IndiGenomes provides\ whole genome sequencing data of 1,029 healthy Indian individuals under the pilot phase of the\ "IndiGen" program. The IndiGenomes website also provides SV call and Alu insertion VCFs.\
\\ The deployed VCF shown in this track is the public release subset distributed by the\ IndiGenomes project (18,016,257 records). The full Jain 2021 callset reports 55.8 million\ variants from the 1,029-genome cohort; the public release is a curated subset of those\ sites. The deployed VCF is sites-only and carries a per-variant VRT (variant type)\ INFO field. Per-variant allele counts and allele frequencies are not distributed with the\ public release and are not shown in this track.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is indigenomes.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original data can also be downloaded from the IndiGen website.\
\ \\ Genomic DNA was extracted from 5 ml of peripheral blood collected via venipuncture from\ 1,029 self-declared healthy Indian individuals representing diverse geographic, ethnic, and\ linguistic groups, using the salting-out method. Whole-genome libraries were prepared using\ the TruSeq DNA PCR-free library preparation kit (Illumina). Sequencing was performed on the\ Illumina NovaSeq 6000 platform with 150×2 bp paired-end reads targeting ≥30×\ mean coverage. Alignment to the GRCh38 reference genome, post-processing, and\ default quality-filtered variant calling were performed end-to-end on the Illumina DRAGEN\ v3.4 Bio-IT platform, which uses field-programmable gate array (FPGA) logic for\ high-throughput processing. The full Jain 2021 callset comprises 55,898,122 single-allelic\ genetic variants (SNVs and indels), of which 32.23% were unique to the Indian samples\ and absent from global reference databases. Variants were annotated using ANNOVAR with\ RefGene, and allele frequencies were cross-referenced against gnomAD v3, 1000 Genomes,\ ExAC, ESP6500, and the Greater Middle East Variome Project. The\ IndiGenomes database\ distributes a public-release subset of these variants (18,016,257 records); that subset is\ the file used in this track.\ (Jain, Bhoyar, Scaria, Sivasubbu & the IndiGen Consortium,\ Nucleic Acids Research 2021).\
\\ The makeDoc file of the track documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were also used and are available from GitHub.\
\ \\ Jain A, Bhoyar RC, Pandhare K, Mishra A, Sharma D, Imran M, Senthivel V, Divakar MK, Rophina M,\ Jolly B et al.\ \ IndiGenomes: a comprehensive resource of genetic variants from over 1000 Indian genomes.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D1225-D1232.\ PMID: 33095885; PMC: PMC7778947\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/indigenomes/IndiGenomes_Variants.vcf.gz\ dataVersion IndiGen pilot (Jain 2021)\ longLabel SNV Frequencies: IndiGenomes India - 1,029 samples\ parent varFreqs on\ priority 16\ shortLabel India IndiGenomes 1k WGS\ track indigenomes\ type vcfTabix\ visibility hide\ LGG LGG bigLolly 12 + Brain Lower Grade Glioma 0 16 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/LGG.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Brain Lower Grade Glioma\ parent gdcCancer off\ priority 16\ shortLabel LGG\ track LGG\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4DnaseMouth Mouth bigWig Avg. DNase level of 4 mouth experiments (tissues and primary cells only) 0 16 130 141 158 192 198 206 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpMouthDNase.bw\ color 130,141,158\ longLabel Avg. DNase level of 4 mouth experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 16\ shortLabel Mouth\ track wgEncodeReg4DnaseMouth\ type bigWig\ wgEncodeReg4AtacAllNerve Nerve (all biosamples) bigWig Avg. ATAC level of 3 nerve experiments (all biosamples) 0 16 160 156 0 207 205 127 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/nerveATAC.bw\ color 160,156,0\ longLabel Avg. ATAC level of 3 nerve experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 16\ shortLabel Nerve (all biosamples)\ track wgEncodeReg4AtacAllNerve\ type bigWig\ oligo42N Oligodendrocytes - Z0000042N bigWig Methylation Atlas: Oligodendrocytes - Z0000042N 2 16 148 103 189 201 179 222 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/oligo42N.bw\ color 148,103,189\ longLabel Methylation Atlas: Oligodendrocytes - Z0000042N\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 16\ shortLabel Oligodendrocytes - Z0000042N\ subGroups cellType=Oligodend dataType=Replicate\ track oligo42N\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ lincRNAsCTPlacenta_R Placenta_R bed 5 + lincRNAs from placenta_r 1 16 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from placenta_r\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Placenta_R\ subGroups view=lincRNAsRefseqExp tissueType=placenta_r\ track lincRNAsCTPlacenta_R\ wgEncodeReg4MarkH3k4me3Prostate Prostate bigWig Avg. H3K4me3 level of 2 prostate experiments (tissues and primary cells only) 0 16 140 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpProstateH3K4me3.bw\ color 140,140,140\ longLabel Avg. H3K4me3 level of 2 prostate experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 16\ shortLabel Prostate\ track wgEncodeReg4MarkH3k4me3Prostate\ type bigWig\ wgEncodeReg4MarkH3k27acSkin Skin bigWig Avg. H3K27ac level of 29 skin experiments (tissues and primary cells only) 2 16 127 133 209 191 194 232 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpSkinH3K27ac.bw\ color 127,133,209\ longLabel Avg. H3K27ac level of 29 skin experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 16\ shortLabel Skin\ track wgEncodeReg4MarkH3k27acSkin\ type bigWig\ Agilent_Human_Exon_V4_Regions SureSel. V4+UTR T bigBed Agilent - SureSelect All Exon V4 + UTRs Target Regions 0 16 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S04380110_Regions.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V4 + UTRs Target Regions\ parent exomeProbesets off\ shortLabel SureSel. V4+UTR T\ track Agilent_Human_Exon_V4_Regions\ type bigBed\ chainFelCat9 Cat Chain chain felCat9 Cat (Nov. 2017 (Felis_catus_9.0/felCat9)) Chained Alignments 3 17 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Cat (Nov. 2017 (Felis_catus_9.0/felCat9)) Chained Alignments\ otherDb felCat9\ parent placentalChainNetViewchain off\ shortLabel Cat Chain\ subGroups view=chain species=s037c clade=c01\ track chainFelCat9\ type chain felCat9\ chainPetMar3 Lamprey Chain chain petMar3 Lamprey (Dec. 2017 (Pmar_germline 1.0/petMar3)) Chained Alignments 3 17 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Lamprey (Dec. 2017 (Pmar_germline 1.0/petMar3)) Chained Alignments\ otherDb petMar3\ parent vertebrateChainNetViewchain off\ shortLabel Lamprey Chain\ subGroups view=chain species=s064a clade=c07\ track chainPetMar3\ type chain petMar3\ chainRheMac10 Rhesus Chain chain rheMac10 Rhesus (Feb. 2019 (Mmul_10/rheMac10)) Chained Alignments 3 17 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Rhesus (Feb. 2019 (Mmul_10/rheMac10)) Chained Alignments\ otherDb rheMac10\ parent primateChainNetViewchain off\ shortLabel Rhesus Chain\ subGroups view=chain species=s021 clade=c01\ track chainRheMac10\ type chain rheMac10\ encTfChipPkENCFF808RWZ A549 FOSL2 narrowPeak Transcription Factor ChIP-seq Peaks of FOSL2 in A549 from ENCODE 3 (ENCFF808RWZ) 0 17 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of FOSL2 in A549 from ENCODE 3 (ENCFF808RWZ)\ parent encTfChipPk off\ shortLabel A549 FOSL2\ subGroups cellType=A549 factor=FOSL2\ track encTfChipPkENCFF808RWZ\ wgEncodeReg4MarkCtcfAllAdrenalGland Adrenal gland (all biosamples) bigWig Avg. CTCF level of 5 adrenal gland experiments (all biosamples) 0 17 90 179 68 172 217 161 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adrenalGlandCTCF.bw\ color 90,179,68\ longLabel Avg. CTCF level of 5 adrenal gland experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 17\ shortLabel Adrenal gland (all biosamples)\ track wgEncodeReg4MarkCtcfAllAdrenalGland\ type bigWig\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_ctss_fwd AorticSmsToFgf2_00hr45minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_forward 0 17 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep1%20%28LK10%29.CNhs13343.12645-134G8.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12645-134G8 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr45minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_tpm_fwd AorticSmsToFgf2_00hr45minBr1+ bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_forward 1 17 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep1%20%28LK10%29.CNhs13343.12645-134G8.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12645-134G8 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr45minBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8\ urlLabel FANTOM5 Details:\ gtexCovBrainNucleusaccumbensbasalganglia Brain Nucl acc bas gang bigWig Brain Nucleus accumbens basal ganglia 0 17 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-14BIN-0011-R6a-SM-5S2RH.Brain_Nucleus_accumbens_basal_ganglia.RNAseq.bw\ color 238,238,0\ longLabel Brain Nucleus accumbens basal ganglia\ parent gtexCov\ shortLabel Brain Nucl acc bas gang\ track gtexCovBrainNucleusaccumbensbasalganglia\ cloneEndCTD CTD bed 12 CalTech BAC library D 0 17 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel CalTech BAC library D\ parent cloneEndSuper on\ priority 20\ shortLabel CTD\ subGroups source=caltech\ track cloneEndCTD\ type bed 12\ visibility hide\ ENCFF013UBZ_ENCFF901QWB_ENCFF972ZHA_ENCFF500RDL ENCFF013UBZ_ENCFF901QWB_ENCFF972ZHA_ENCFF500RDL bigBed 9 + 5 Tibial artery, male adult (37 years): (1) cCREs 4 17 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF013UBZ_ENCFF901QWB_ENCFF972ZHA_ENCFF500RDL.bb\ longLabel Tibial artery, male adult (37 years): (1) cCREs\ mouseOver ID: ${name}\ An allele frequency panel based on short-read whole-genome sequencing analysis of 61,000 Japanese\ individuals, produced by the\ Tohoku Medical Megabank\ Organization (ToMMo) at Tohoku University. The project includes other datatypes such as STRs,\ long-read SVs and short-read CNVs.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is tommo60kjpn.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original data can also be downloaded from the jMorp website, specifically the\ Downloads section.\
\ \\ Genomic DNA was obtained from peripheral blood, saliva, or cord blood samples. Sequencing was\ performed on Illumina HiSeq 2500, HiSeq X Five, NovaSeq 6000, and MGI DNBSeq G400/T7 instruments.\ Reads were aligned to the GRCh38 reference using BWA 0.7.15 or BWA-mem2 2.1. Alignments underwent\ base quality score recalibration (BQSR) with the GATK BaseRecalibrator tool. SNV/indel calling was\ performed using GATK HaplotypeCaller, followed by multisample joint genotyping with Sentieon\ Genomics tools and variant quality score recalibration (VQSR) filtering. Related samples were\ identified and removed with KING 2.3.1 to produce the final allele frequency panel.\
\\ The makeDoc file for this track describes how all source files of the varFreqs track were converted.\ For some tracks, python scripts were needed and are available from GitHub.\
\ \\ Tadaka S, Kawashima J, Hishinuma E, Saito S, Okamura Y, Otsuki A, Kojima K, Komaki S, Aoki Y, Kanno\ T et al.\ \ jMorp: Japanese Multi-Omics Reference Panel update report 2023.\ Nucleic Acids Res. 2024 Jan 5;52(D1):D622-D632.\ PMID: 37930845; PMC: PMC10767895\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/tommo61kjpn/tommo-61kjpn-20250616-GRCh38-snvindel-af-autosome.vcf.gz\ dataVersion 2025-06-16\ longLabel SNV Frequencies: Japan 61k - ToMMo SNV+Indels\ parent varFreqs on\ priority 17\ shortLabel Japan ToMMo 61k WGS\ track tommo60kjpn\ type vcfTabix\ visibility hide\ LIHC LIHC bigLolly 12 + Liver hepatocellular carcinoma 0 17 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/LIHC.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Liver hepatocellular carcinoma\ parent gdcCancer off\ priority 17\ shortLabel LIHC\ track LIHC\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4DnaseMuscle Muscle bigWig Avg. DNase level of 63 muscle experiments (tissues and primary cells only) 0 17 137 135 170 196 195 212 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpMuscleDNase.bw\ color 137,135,170\ longLabel Avg. DNase level of 63 muscle experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase\ priority 17\ shortLabel Muscle\ track wgEncodeReg4DnaseMuscle\ type bigWig\ wgEncodeReg4AtacAllOvary Ovary (all biosamples) bigWig Avg. ATAC level of 5 ovary experiments (all biosamples) 0 17 161 126 151 208 190 203 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/ovaryATAC.bw\ color 161,126,151\ longLabel Avg. ATAC level of 5 ovary experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 17\ shortLabel Ovary (all biosamples)\ track wgEncodeReg4AtacAllOvary\ type bigWig\ lincRNAsCTProstate Prostate bed 5 + lincRNAs from prostate 1 17 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from prostate\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Prostate\ subGroups view=lincRNAsRefseqExp tissueType=prostate\ track lincRNAsCTProstate\ wgEncodeReg4MarkH3k4me3Skin Skin bigWig Avg. H3K4me3 level of 19 skin experiments (tissues and primary cells only) 0 17 127 133 209 191 194 232 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpSkinH3K4me3.bw\ color 127,133,209\ longLabel Avg. H3K4me3 level of 19 skin experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 17\ shortLabel Skin\ track wgEncodeReg4MarkH3k4me3Skin\ type bigWig\ Agilent_Human_Exon_V5_UTRs_Covered SureSel. V5+UTR P bigBed Agilent - SureSelect All Exon V5 + UTRs Covered by Probes 0 17 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S04380219_Covered.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V5 + UTRs Covered by Probes\ parent exomeProbesets off\ shortLabel SureSel. V5+UTR P\ track Agilent_Human_Exon_V5_UTRs_Covered\ type bigBed\ wgEncodeReg4MarkH3k27acUterus Uterus bigWig Avg. H3K27ac level of 2 uterus experiments (tissues and primary cells only) 2 17 186 111 165 220 183 210 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpUterusH3K27ac.bw\ color 186,111,165\ longLabel Avg. H3K27ac level of 2 uterus experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k27ac off\ priority 17\ shortLabel Uterus\ track wgEncodeReg4MarkH3k27acUterus\ type bigWig\ netFelCat9 Cat Net netAlign felCat9 chainFelCat9 Cat (Nov. 2017 (Felis_catus_9.0/felCat9)) Alignment Net 1 18 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Cat (Nov. 2017 (Felis_catus_9.0/felCat9)) Alignment Net\ otherDb felCat9\ parent placentalChainNetViewnet off\ shortLabel Cat Net\ subGroups view=net species=s037c clade=c01\ track netFelCat9\ type netAlign felCat9 chainFelCat9\ netPetMar3 Lamprey Net netAlign petMar3 chainPetMar3 Lamprey (Dec. 2017 (Pmar_germline 1.0/petMar3)) Alignment Net 1 18 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Lamprey (Dec. 2017 (Pmar_germline 1.0/petMar3)) Alignment Net\ otherDb petMar3\ parent vertebrateChainNetViewnet off\ shortLabel Lamprey Net\ subGroups view=net species=s064a clade=c07\ track netPetMar3\ type netAlign petMar3 chainPetMar3\ netRheMac10 Rhesus Net netAlign rheMac10 chainRheMac10 Rhesus (Feb. 2019 (Mmul_10/rheMac10)) Alignment Net 1 18 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Rhesus (Feb. 2019 (Mmul_10/rheMac10)) Alignment Net\ otherDb rheMac10\ parent primateChainNetViewnet off\ shortLabel Rhesus Net\ subGroups view=net species=s021 clade=c01\ track netRheMac10\ type netAlign rheMac10 chainRheMac10\ encTfChipPkENCFF297HAX A549 FOXA1 1 narrowPeak Transcription Factor ChIP-seq Peaks of FOXA1 in A549 from ENCODE 3 (ENCFF297HAX) 0 18 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of FOXA1 in A549 from ENCODE 3 (ENCFF297HAX)\ parent encTfChipPk off\ shortLabel A549 FOXA1 1\ subGroups cellType=A549 factor=FOXA1\ track encTfChipPkENCFF297HAX\ wgEncodeReg4MarkH3k27acAllAdipose Adipose (all biosamples) bigWig Avg. H3K27ac level of 2 adipose experiments (all biosamples) 2 18 255 119 39 255 187 147 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adiposeH3K27ac.bw\ color 255,119,39\ longLabel Avg. H3K27ac level of 2 adipose experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 18\ shortLabel Adipose (all biosamples)\ track wgEncodeReg4MarkH3k27acAllAdipose\ type bigWig\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_ctss_rev AorticSmsToFgf2_00hr45minBr1- bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_reverse 0 18 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep1%20%28LK10%29.CNhs13343.12645-134G8.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12645-134G8 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr45minBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_tpm_rev AorticSmsToFgf2_00hr45minBr1- bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_reverse 1 18 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep1%20%28LK10%29.CNhs13343.12645-134G8.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep1 (LK10)_CNhs13343_12645-134G8_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12645-134G8 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr45minBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep1LK10_CNhs13343_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12645-134G8\ urlLabel FANTOM5 Details:\ bacRearray32k BCGSC 32k rearray bigBed 5 BCGSC Human BAC Rearray (32k minimal tiling set, mostly RP11 and CTD) 3 18 0 60 120 127 157 187 0 0 0\ This track shows the placements of the BCGSC Human BAC Re-Array,\ a curated "32k set" of bacterial artificial chromosome (BAC) clones that\ represents a minimal-overlap tiling path across the human genome. The clones\ are available for ordering from\ BACPAC\ Genomics at the Children's Hospital Oakland Research Institute (CHORI),\ where additional information on the collection, chromosome-specific sub-plates\ and quality control can also be found.\
\ \\ The set was assembled by the BC Cancer Genome Sciences Centre (Marco Marra\ lab) together with the BACPAC Genomics group at CHORI to provide a compact,\ redundant-free clone collection for physical mapping, functional genomics and\ array-based assays such as comparative genomic hybridization (CGH) and FISH.\
\ \\ The 32k set was selected from the human physical fingerprint map so that every\ region of the map is represented at least once. Additional clones were added\ to fill regions that initially lacked coverage, resulting in an average\ resolution of approximately 46 kb between overlapping BAC segments. The set\ provides coverage for more than 99% of both the fingerprint map and the\ reference genome assembly.\
\ \\ Most clones in the set (about 30,388) are drawn from the RPCI-11 and RPCI-13\ human BAC libraries; a smaller contribution (about 2,062 clones) comes from\ the CalTech CIT-D library (CTD). Unlike the other tracks in this container,\ the 32k set does not rely on BAC end sequences: clone identity and assembly\ placement have been verified repeatedly by HindIII fingerprinting at the\ Genome Sciences Centre.\
\ \\ Each item represents the genomic placement of a single BAC clone in the 32k\ re-array, labeled with its clone name (e.g. RP11-…,\ CTD-…). Clicking a clone opens a detail page with a link\ to the corresponding BACPAC Genomics clone record, where ordering and\ additional clone metadata are available.\
\ \\ The original placements were generated by CHORI/BACPAC by lifting an earlier\ clone set onto GRCh38/hg38 from older assemblies. The pre-lifted BED file\ provided by BACPAC was downloaded from\ bacpacresources.org, the custom-track header lines were\ removed, the data were sorted and the file was converted to bigBed (BED5)\ format with bedToBigBed.\
\ \\ Note: earlier supporting data such as the original HindIII fingerprints and\ BAC end sequences for this minimal set are no longer available from BACPAC\ Genomics.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator, and exported from there\ to spreadsheet or tab-separated tables. From scripts, the data can be accessed\ through our REST API,\ using track=bacRearray32k.\
\ \\ For automated download and analysis, the genome annotation is stored in a\ bigBed file that can be downloaded from\ our\ download server as bacRearray32k.bb. Individual regions or the\ whole annotation can be obtained using our tool bigBedToBed, which\ can be compiled from the source code or downloaded as a precompiled binary for\ your system. Instructions for downloading source code and binaries can be\ found here. The tool can also be used to obtain features within\ a given range, e.g.\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/cloneEnd/bacRearray32k.bb\ -chrom=chr21 -start=0 -end=100000000 stdout.\
\ \\ The original annotation source data can be downloaded from\ BACPAC Genomics. Additional information about the clone\ collection, including ordering, chromosome-specific sub-plates, and quality\ control, is available on the BACPAC\ Human BAC Minimal Tiling Set page.\
\ \\ The 32k Human BAC Re-Array was generated in collaboration between BACPAC\ Genomics at CHORI (Pieter J. de Jong and colleagues) and the BC Cancer Genome\ Sciences Centre (Marco Marra lab), Vancouver, BC, Canada. We thank Pieter J.\ de Jong for providing the pre-lifted hg38 placements. The RPCI-11 and RPCI-13\ source libraries were constructed at the Roswell Park Cancer Institute. For\ background on de Jong's role in building these clone libraries, see this\ Undark profile.\
\ \\ Krzywinski M, Bosdet I, Smailus D, Chiu R, Mathewson C, Wye N, Barber S, Brown-John M, Chan S, Chand\ S et al.\ \ A set of BAC clones spanning the human genome.\ Nucleic Acids Res. 2004;32(12):3651-60.\ PMID: 15247347; PMC: PMC484185\
\ \\ Osoegawa K, Mammoser AG, Wu C, Frengen E, Zeng C, Catanese JJ, de Jong PJ.\ \ A bacterial artificial chromosome library for sequencing the complete human genome.\ Genome Res. 2001 Mar;11(3):483-96.\ PMID: 11230172; PMC: PMC311044\
\ map 1 bigDataUrl /gbdb/hg38/bbi/cloneEnd/bacRearray32k.bb\ color 0,60,120\ longLabel BCGSC Human BAC Rearray (32k minimal tiling set, mostly RP11 and CTD)\ parent cloneEndSuper on\ priority 17.5\ shortLabel BCGSC 32k rearray\ subGroups source=chori\ track bacRearray32k\ type bigBed 5\ visibility pack\ wgEncodeReg4MarkCtcfAllBloodVessel Blood vessel (all biosamples) bigWig Avg. CTCF level of 12 blood vessel experiments (all biosamples) 0 18 255 37 41 255 146 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodVesselCTCF.bw\ color 255,37,41\ longLabel Avg. CTCF level of 12 blood vessel experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 18\ shortLabel Blood vessel (all biosamples)\ track wgEncodeReg4MarkCtcfAllBloodVessel\ type bigWig\ gtexCovBrainPutamenbasalganglia Brain Put bas gang bigWig Brain Putamen basal ganglia 0 18 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-1HFI6-0011-R7b-SM-CM2SS.Brain_Putamen_basal_ganglia.RNAseq.bw\ color 238,238,0\ longLabel Brain Putamen basal ganglia\ parent gtexCov\ shortLabel Brain Put bas gang\ track gtexCovBrainPutamenbasalganglia\ ENCFF156LUX_ENCFF696UEY_ENCFF762YWL_ENCFF429ZQN ENCFF156LUX_ENCFF696UEY_ENCFF762YWL_ENCFF429ZQN bigBed 9 + 5 Thoracic aorta, male adult (37 years): (1) cCREs 4 18 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF156LUX_ENCFF696UEY_ENCFF762YWL_ENCFF429ZQN.bb\ longLabel Thoracic aorta, male adult (37 years): (1) cCREs\ mouseOver ID: ${name}\ The Korean Variant Archive (KOVA)\ contains 1,896 whole genome sequencing and 3,409 whole exome sequencing data from healthy\ individuals of Korean ethnicity. Most of the samples originated from normal tissue of cancer\ patients (40.16%), healthy parents of rare disease patients (28.4%), or healthy volunteers\ (31.44%). Korean ancestry is not broken down further in the INFO field. Coverage 100x for WES, 30x for WGS.\ SVs called with Manta are also available.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ TSV data can be requested on the KOVA Downloads website. Our GitHub repo contains a script that\ converts this format to VCF.\
\ \\ Raw reads were aligned to the GRCh38+decoy reference with BWA-MEM v0.7.17 with default\ parameters. Duplicates were marked and coordinates sorted with MarkDuplicatesSpark, then base\ quality scores were recalibrated with BQSRPipelineSpark in GATK v4.1.3.0. Mapping quality control\ metrics were generated with Qualimap v2.2.1. Single-nucleotide variants and small\ insertions/deletions were called per sample with GATK HaplotypeCaller in GVCF mode (-ERC GVCF), and\ joint genotyping was performed by creating a GenomicsDB with GenomicsDBImport and followed GATK\ Best Practices. Variant quality score recalibration (VQSR) retained 99.7% of true SNVs\ and 99.0% of true indels based on training sets (workflow detailed in Supplementary Fig. 1).\ Downstream analyses followed a modified version of the gnomAD quality-control framework and were\ primarily conducted with Hail. After WES and WGS data were merged in Hail, multiallelic variants and\ variants with genotype quality <20, read depth <10, allelic balance <0.2, or overlap with\ low-complexity regions were excluded.\
\\ At UCSC, V7 of the TSV.gz was obtained from the KOVA staff by email and converted to VCF. The file is not\ available for download from our site, but can be requested from the KOVA website.\ The makeDoc file for the varFreqs track documents how all source files were converted.\ Python scripts used for some of the tracks are available on GitHub.\
\ \\ Thanks to Insu Jang and the KOVA director for providing variant frequencies in TSV format.\
\ \\ Lee J, Lee J, Jeon S, Lee J, Jang I, Yang JO, Park S, Lee B, Choi J, Choi BO et al.\ \ A database of 5305 healthy Korean individuals reveals genetic and clinical implications for an East\ Asian population.\ Exp Mol Med. 2022 Nov;54(11):1862-1871.\ PMID: 36323850; PMC: PMC9628380\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_kova/kova.v7.vcf.gz\ dataVersion V7\ longLabel SNV Frequencies: KOVA Korea - 5305 samples, 1.9k WGS+3.4k WES\ parent varFreqs on\ priority 18\ shortLabel Korea KOVA 5.3k mixed\ tableBrowser off\ track kova\ type vcfTabix\ visibility hide\ LUAD LUAD bigLolly 12 + Lung adenocarcinoma 0 18 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/LUAD.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Lung adenocarcinoma\ parent gdcCancer off\ priority 18\ shortLabel LUAD\ track LUAD\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeRegDnaseUwNhlfPeak NHLF Pk narrowPeak NHLF lung fibroblast DNaseI Peaks from ENCODE 1 18 255 209 85 255 232 170 1 0 0 regulation 1 color 255,209,85\ longLabel NHLF lung fibroblast DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak on\ shortLabel NHLF Pk\ subGroups view=a_Peaks cellType=NHLF treatment=n_a tissue=lung cancer=normal\ track wgEncodeRegDnaseUwNhlfPeak\ wgEncodeRegDnaseUwNhlfWig NHLF Sg bigWig 0 11719.1 NHLF lung fibroblast DNaseI Signal from ENCODE 0 18 255 209 85 255 232 170 0 0 0 regulation 1 color 255,209,85\ longLabel NHLF lung fibroblast DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig on\ priority 1.156\ shortLabel NHLF Sg\ subGroups cellType=NHLF treatment=n_a tissue=lung cancer=normal\ table wgEncodeRegDnaseUwNhlfSignal\ track wgEncodeRegDnaseUwNhlfWig\ type bigWig 0 11719.1\ wgEncodeReg4DnasePancreas Pancreas bigWig Avg. DNase level of 12 pancreas experiments (tissues and primary cells only) 0 18 175 100 41 215 177 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPancreasDNase.bw\ color 175,100,41\ longLabel Avg. DNase level of 12 pancreas experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 18\ shortLabel Pancreas\ track wgEncodeReg4DnasePancreas\ type bigWig\ lincRNAsCTSkeletalMuscle SkeletalMuscle bed 5 + lincRNAs from skeletalmuscle 1 18 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from skeletalmuscle\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel SkeletalMuscle\ subGroups view=lincRNAsRefseqExp tissueType=skeletalmuscle\ track lincRNAsCTSkeletalMuscle\ wgEncodeReg4AtacAllSmallIntestine Small intestine (all biosamples) bigWig ATAC level of 1 small intestine experiment (all biosamples) 0 18 98 98 41 176 176 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/smallIntestineATAC.bw\ color 98,98,41\ longLabel ATAC level of 1 small intestine experiment (all biosamples)\ parent wgEncodeReg4Atac off\ priority 18\ shortLabel Small intestine (all biosamples)\ track wgEncodeReg4AtacAllSmallIntestine\ type bigWig\ Agilent_Human_Exon_V5_UTRs_Regions SureSel. V5+UTR T bigBed Agilent - SureSelect All Exon V5 + UTRs Target Regions 0 18 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S04380219_Regions.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V5 + UTRs Target Regions\ parent exomeProbesets off\ shortLabel SureSel. V5+UTR T\ track Agilent_Human_Exon_V5_UTRs_Regions\ type bigBed\ wgEncodeReg4MarkH3k4me3Testis Testis bigWig Avg. H3K4me3 level of 2 testis experiments (tissues and primary cells only) 0 18 139 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpTestisH3K4me3.bw\ color 139,140,140\ longLabel Avg. H3K4me3 level of 2 testis experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 18\ shortLabel Testis\ track wgEncodeReg4MarkH3k4me3Testis\ type bigWig\ chainPapAnu4 papAnu4 Chain chain papAnu4 Baboon (Apr. 2017 (Panu_3.0/papAnu4)) Chained Alignments 3 19 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Baboon (Apr. 2017 (Panu_3.0/papAnu4)) Chained Alignments\ otherDb papAnu4\ parent primateChainNetViewchain off\ shortLabel papAnu4 Chain\ subGroups view=chain species=s024 clade=c01\ track chainPapAnu4\ type chain papAnu4\ chainEnhLutNer1 Southern sea otter Chain chain enhLutNer1 Southern sea otter (Jun. 2019 (ASM641071v1/enhLutNer1)) Chained Alignments 3 19 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Southern sea otter (Jun. 2019 (ASM641071v1/enhLutNer1)) Chained Alignments\ otherDb enhLutNer1\ parent placentalChainNetViewchain off\ shortLabel Southern sea otter Chain\ subGroups view=chain species=s043a clade=c01\ track chainEnhLutNer1\ type chain enhLutNer1\ encTfChipPkENCFF167BKY A549 FOXA1 2 narrowPeak Transcription Factor ChIP-seq Peaks of FOXA1 in A549 from ENCODE 3 (ENCFF167BKY) 0 19 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of FOXA1 in A549 from ENCODE 3 (ENCFF167BKY)\ parent encTfChipPk off\ shortLabel A549 FOXA1 2\ subGroups cellType=A549 factor=FOXA1\ track encTfChipPkENCFF167BKY\ wgEncodeReg4MarkH3k27acAllAdrenalGland Adrenal gland (all biosamples) bigWig Avg. H3K27ac level of 8 adrenal gland experiments (all biosamples) 2 19 90 179 68 172 217 161 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adrenalGlandH3K27ac.bw\ color 90,179,68\ longLabel Avg. H3K27ac level of 8 adrenal gland experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 19\ shortLabel Adrenal gland (all biosamples)\ track wgEncodeReg4MarkH3k27acAllAdrenalGland\ type bigWig\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_ctss_fwd AorticSmsToFgf2_00hr45minBr2+ bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_forward 0 19 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep2%20%28LK11%29.CNhs13361.12743-135I7.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12743-135I7 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr45minBr2+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_tpm_fwd AorticSmsToFgf2_00hr45minBr2+ bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_forward 1 19 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep2%20%28LK11%29.CNhs13361.12743-135I7.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12743-135I7 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr45minBr2+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7\ urlLabel FANTOM5 Details:\ gtexCovBrainSpinalcordcervicalc-1 Brain Spinal cord cerv bigWig Brain Spinal cord cervical c-1 0 19 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-YFC4-0011-R9a-SM-4SOK4.Brain_Spinal_cord_cervical_c-1.RNAseq.bw\ color 238,238,0\ longLabel Brain Spinal cord cervical c-1\ parent gtexCov\ shortLabel Brain Spinal cord cerv\ track gtexCovBrainSpinalcordcervicalc-1\ cloneEndCH17 CH17 bed 12 CHORI BAC hydatidiform mole 0 19 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel CHORI BAC hydatidiform mole\ parent cloneEndSuper on\ priority 17\ shortLabel CH17\ subGroups source=chori\ track cloneEndCH17\ type bed 12\ visibility hide\ ENCFF226FAT_ENCFF582GHH_ENCFF441MGU_ENCFF897TLT ENCFF226FAT_ENCFF582GHH_ENCFF441MGU_ENCFF897TLT bigBed 9 + 5 Osteocyte, female embryo (5 days): (1) cCREs 4 19 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF226FAT_ENCFF582GHH_ENCFF441MGU_ENCFF897TLT.bb\ longLabel Osteocyte, female embryo (5 days): (1) cCREs\ mouseOver ID: ${name}\ The Mexico Biobank (MXB) project\ genotyped 6,011 individuals sampled across all 32 states of Mexico during the 2000 National\ Health Survey (ENSA 2000) conducted by the National Institute of Public Health (INSP).\ Genotyping used the Illumina Multi-Ethnic Global Array (MEGA, ~1.8M SNPs), which is\ optimized for admixed populations and enriched for ancestry-informative and medically relevant\ variants. Only autosomal, biallelic SNPs that passed quality control are included. Samples\ came from 898 recruitment sites, and indigenous language speakers were prioritized.\
\ \\ This track shows allele frequencies computed from the phased genotypes. The full\ phased genotype data with haplotype clustering display is available in the\ Mexico Biobank track under Phased Variants.\ Frequencies can also be plotted onto a map on the\ MexVar platform.\ The hg38 data was lifted from hg19 by UCSC (see below).\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ Allele frequencies by geographical state and ancestry are available via\ the MexVar platform.\ Raw genotype data are available under controlled access at the\ EGA (Study: EGAS00001005797; Dataset: EGAD00010002361). For the VCFs, email\ andres.moreno@cinvestav.mx to obtain the data.\
\ \\ Data processing included GenomeStudio → PLINK conversion, strand alignment, removal\ of duplicates, update of map positions using dbSNP Build 151 and low-quality\ variants/individuals, and relatedness filtering.\ At UCSC, the phased VCF was lifted from hg19 to hg38 with CrossMap, then allele counts\ (AC, AF, AN) were computed using bcftools fill-tags and genotypes were stripped to produce\ a sites-only frequency VCF.\
\ \\ The makeDoc file documents how the source files of the varFreqs track were converted.\ For some tracks, python scripts were needed and are also available from GitHub.\
\ \\ We thank the Center for Research and Advanced Studies (Cinvestav) of Mexico for\ generating and providing the frequency data, the National Institute of Medical\ Sciences and Nutrition (INCMNSZ) for DNA extraction, and the Ministry of Health\ together with the National Institute of Public Health (INSP) for the design and\ implementation of the National Health Survey 2000 (ENSA 2000). We also thank\ the ENSA-Genomics Consortium for their contributions to sample collection and\ data processing that made possible the construction of the MXB genomic resource.\
\ \\ Barberena-Jonas C, Medina-Muñoz SG, Cedillo-Castelán V, Sepúlveda-Morales T,\ Gonzaga-Jáuregui C, ENSA Genomics Consortium, García-García L, Ioannidis AG,\ Moreno-Estrada A.\ \ Clinical genetic variation across Hispanic populations in the Mexican Biobank.\ Nat Med. 2026 Jan 21;.\ DOI: 10.1038/s41591-025-04100-z; PMID: 41566040\
\ \\ Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Corés CD, Barberena-Jonas C, Medina-Muñoz SG,\ Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L et al.\ \ Mexican Biobank advances population and medical genomics of diverse ancestries.\ Nature. 2023 Oct;622(7984):775-783.\ PMID: 37821706; PMC: PMC10600006\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_mxb/mxb.freq.vcf.gz\ dataVersion Nov 2025 (hg38 lift)\ longLabel SNV Frequencies: Mexico Biobank - 6,011 individuals, genotyping array\ parent varFreqs on\ priority 19\ shortLabel Mexico Biobank 6k Array\ tableBrowser off\ track mxbFreq\ type vcfTabix\ visibility hide\ wgEncodeRegDnaseUwNhaPeak NH-A Pk narrowPeak NH-A astrocyte DNaseI Peaks from ENCODE 1 19 255 210 85 255 232 170 1 0 0 regulation 1 color 255,210,85\ longLabel NH-A astrocyte DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel NH-A Pk\ subGroups view=a_Peaks cellType=NH-A treatment=n_a tissue=brain cancer=normal\ track wgEncodeRegDnaseUwNhaPeak\ wgEncodeRegDnaseUwNhaWig NH-A Sg bigWig 0 9132.47 NH-A astrocyte DNaseI Signal from ENCODE 0 19 255 210 85 255 232 170 0 0 0 regulation 1 color 255,210,85\ longLabel NH-A astrocyte DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.15667\ shortLabel NH-A Sg\ subGroups cellType=NH-A treatment=n_a tissue=brain cancer=normal\ table wgEncodeRegDnaseUwNhaSignal\ track wgEncodeRegDnaseUwNhaWig\ type bigWig 0 9132.47\ wgEncodeReg4DnasePenis Penis bigWig Avg. DNase level of 2 penis experiments (tissues and primary cells only) 0 19 20 74 159 137 164 207 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPenisDNase.bw\ color 20,74,159\ longLabel Avg. DNase level of 2 penis experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 19\ shortLabel Penis\ track wgEncodeReg4DnasePenis\ type bigWig\ wgEncodeReg4AtacAllSpleen Spleen (all biosamples) bigWig ATAC level of 1 spleen experiment (all biosamples) 0 19 136 157 97 195 206 176 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/spleenATAC.bw\ color 136,157,97\ longLabel ATAC level of 1 spleen experiment (all biosamples)\ parent wgEncodeReg4Atac off\ priority 19\ shortLabel Spleen (all biosamples)\ track wgEncodeReg4AtacAllSpleen\ type bigWig\ Agilent_Human_Exon_V6_UTRs_Regions SureSel. V6 +UTR T bigBed Agilent - SureSelect All Exon V6 + UTR r2 Target Regions 0 19 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S07604624_Regions.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V6 + UTR r2 Target Regions\ parent exomeProbesets off\ shortLabel SureSel. V6 +UTR T\ track Agilent_Human_Exon_V6_UTRs_Regions\ type bigBed\ lincRNAsCTTestes Testes bed 5 + lincRNAs from testes 1 19 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from testes\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Testes\ subGroups view=lincRNAsRefseqExp tissueType=testes\ track lincRNAsCTTestes\ wgEncodeReg4MarkH3k4me3Uterus Uterus bigWig Avg. H3K4me3 level of 2 uterus experiments (tissues and primary cells only) 0 19 186 111 165 220 183 210 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpUterusH3K4me3.bw\ color 186,111,165\ longLabel Avg. H3K4me3 level of 2 uterus experiments (tissues and primary cells only)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 19\ shortLabel Uterus\ track wgEncodeReg4MarkH3k4me3Uterus\ type bigWig\ netPapAnu4 papAnu4 Net netAlign papAnu4 chainPapAnu4 Baboon (Apr. 2017 (Panu_3.0/papAnu4)) Alignment Net 1 20 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Baboon (Apr. 2017 (Panu_3.0/papAnu4)) Alignment Net\ otherDb papAnu4\ parent primateChainNetViewnet off\ shortLabel papAnu4 Net\ subGroups view=net species=s024 clade=c01\ track netPapAnu4\ type netAlign papAnu4 chainPapAnu4\ netEnhLutNer1 Southern sea otter Net netAlign enhLutNer1 chainEnhLutNer1 Southern sea otter (Jun. 2019 (ASM641071v1/enhLutNer1)) Alignment Net 1 20 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Southern sea otter (Jun. 2019 (ASM641071v1/enhLutNer1)) Alignment Net\ otherDb enhLutNer1\ parent placentalChainNetViewnet off\ shortLabel Southern sea otter Net\ subGroups view=net species=s043a clade=c01\ track netEnhLutNer1\ type netAlign enhLutNer1 chainEnhLutNer1\ encTfChipPkENCFF520GJC A549 GABPA narrowPeak Transcription Factor ChIP-seq Peaks of GABPA in A549 from ENCODE 3 (ENCFF520GJC) 0 20 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of GABPA in A549 from ENCODE 3 (ENCFF520GJC)\ parent encTfChipPk off\ shortLabel A549 GABPA\ subGroups cellType=A549 factor=GABPA\ track encTfChipPkENCFF520GJC\ wgEncodeReg4MarkH3k4me3AllAdipose Adipose (all biosamples) bigWig Avg. H3K4me3 level of 5 adipose experiments (all biosamples) 0 20 255 119 39 255 187 147 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adiposeH3K4me3.bw\ color 255,119,39\ longLabel Avg. H3K4me3 level of 5 adipose experiments (all biosamples)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 20\ shortLabel Adipose (all biosamples)\ track wgEncodeReg4MarkH3k4me3AllAdipose\ type bigWig\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_ctss_rev AorticSmsToFgf2_00hr45minBr2- bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_reverse 0 20 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep2%20%28LK11%29.CNhs13361.12743-135I7.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12743-135I7 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr45minBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_tpm_rev AorticSmsToFgf2_00hr45minBr2- bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_reverse 1 20 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep2%20%28LK11%29.CNhs13361.12743-135I7.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep2 (LK11)_CNhs13361_12743-135I7_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12743-135I7 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr45minBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep2LK11_CNhs13361_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12743-135I7\ urlLabel FANTOM5 Details:\ wgEncodeReg4MarkH3k27acAllBloodVessel Blood vessel (all biosamples) bigWig Avg. H3K27ac level of 12 blood vessel experiments (all biosamples) 2 20 255 37 41 255 146 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodVesselH3K27ac.bw\ color 255,37,41\ longLabel Avg. H3K27ac level of 12 blood vessel experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 20\ shortLabel Blood vessel (all biosamples)\ track wgEncodeReg4MarkH3k27acAllBloodVessel\ type bigWig\ gtexCovBrainSubstantianigra Brain Subst nigr bigWig Brain Substantia nigra 0 20 238 238 0 246 246 127 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-Z93S-0011-R2a-SM-4RGNG.Brain_Substantia_nigra.RNAseq.bw\ color 238,238,0\ longLabel Brain Substantia nigra\ parent gtexCov\ shortLabel Brain Subst nigr\ track gtexCovBrainSubstantianigra\ cloneEndCOR02 COR02 bed 12 NHGRI-CORIELLE CORIELL-02-F-39-40KB 0 20 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel NHGRI-CORIELLE CORIELL-02-F-39-40KB\ parent cloneEndSuper off\ priority 18\ shortLabel COR02\ subGroups source=corielle\ track cloneEndCOR02\ type bed 12\ visibility hide\ ENCFF146ZBO_ENCFF684UUJ_ENCFF900UMO_ENCFF327WOL ENCFF146ZBO_ENCFF684UUJ_ENCFF900UMO_ENCFF327WOL bigBed 9 + 5 NCI-H929: (1) cCREs 4 20 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF146ZBO_ENCFF684UUJ_ENCFF900UMO_ENCFF327WOL.bb\ longLabel NCI-H929: (1) cCREs\ mouseOver ID: ${name}\ The NCBI ALlele Frequency\ Aggregator (ALFA) pipeline computes allele frequencies from approved, unrestricted dbGaP studies\ and makes them publicly available through dbSNP. Its goal is to release frequency data from over\ one million dbGaP subjects to aid discoveries involving common and rare variants with biological\ or disease relevance. The R4 release aggregates allele frequencies from 408,709 subjects.\ After conversion to VCF and removal of zero-frequency entries, the UCSC track contains\ 163 million variants (146 million SNPs and 17 million indels), including hundreds of\ thousands of ClinVar variants.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is alfaVcf.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ We converted the NCBI track hub to VCF format; the data is freely available.\ Genotype and associated individual-level data are accessible through the dbGaP\ authorized access request system.\
\ \\ The ALFA pipeline processes genotype data from approved, unrestricted dbGaP studies, including\ chip array, exome, and genomic sequencing data. Selected study data undergoes quality assurance\ and transformation to standard VCF format. Variants are converted to SPDI notation and normalized\ using VOCA, then aggregated, remapped, and clustered to existing dbSNP rs identifiers or assigned\ new ones. Sample ancestries are validated using GRAF-pop and assigned to 12 major populations.\ QC exclusions include variants and subjects with call rate <95%, datasets failing Ancestry\ Informative Markers consistency checks, and array datasets with conflicting or flipped allele\ orientation.\
\\ The ALFA R4 bigBed files (904M variants) were converted to VCF using a custom script, retaining\ the 163M variants with non-zero allele frequency (146M SNPs, 17M indels).\ The makeDoc file documents how the source files of the varFreqs track were converted.\ For some tracks, python scripts were also needed; these are available on GitHub.\
\ \\ NCBI ALFA does not yet have a peer-reviewed primary publication. Cite the project as:\ Phan L, Jin Y, Zhang H, Qiang W, Shekhtman E, Shao D et al.\ \ ALFA: Allele Frequency Aggregator.\ National Center for Biotechnology Information, U.S. National Library of Medicine, 10 March 2020.\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/alfa/ALFA.vcf.gz\ dataVersion R4\ longLabel SNV Frequencies: NCBI ALFA (dbGaP data) - 408k mixed WGS/WES/array, 163M variants\ parent varFreqs on\ priority 20\ shortLabel NCBI ALFA 408k mixed\ track alfaVcf\ type vcfTabix\ url https://www.ncbi.nlm.nih.gov/snp/$$#frequency_tab\ urlLabel NCBI Variation Page\ visibility hide\ wgEncodeReg4MarkCtcfAllNerve Nerve (all biosamples) bigWig Avg. CTCF level of 4 nerve experiments (all biosamples) 0 20 160 156 0 207 205 127 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/nerveCTCF.bw\ color 160,156,0\ longLabel Avg. CTCF level of 4 nerve experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 20\ shortLabel Nerve (all biosamples)\ track wgEncodeReg4MarkCtcfAllNerve\ type bigWig\ wgEncodeReg4DnasePlacenta Placenta bigWig Avg. DNase level of 24 placenta experiments (tissues and primary cells only) 0 20 104 171 71 179 213 163 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpPlacentaDNase.bw\ color 104,171,71\ longLabel Avg. DNase level of 24 placenta experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 20\ shortLabel Placenta\ track wgEncodeReg4DnasePlacenta\ type bigWig\ wgEncodeReg4AtacAllStomach Stomach (all biosamples) bigWig Avg. ATAC level of 2 stomach experiments (all biosamples) 0 20 145 144 99 200 199 177 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/stomachATAC.bw\ color 145,144,99\ longLabel Avg. ATAC level of 2 stomach experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 20\ shortLabel Stomach (all biosamples)\ track wgEncodeReg4AtacAllStomach\ type bigWig\ Agilent_Human_Exon_V6_Covered SureSel. V6 P bigBed Agilent - SureSelect All Exon V6 r2 Covered by Probes 0 20 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S07604514_Covered.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V6 r2 Covered by Probes\ parent exomeProbesets off\ shortLabel SureSel. V6 P\ track Agilent_Human_Exon_V6_Covered\ type bigBed\ lincRNAsCTTestes_R Testes_R bed 5 + lincRNAs from testes_r 1 20 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from testes_r\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Testes_R\ subGroups view=lincRNAsRefseqExp tissueType=testes_r\ track lincRNAsCTTestes_R\ chainNeoSch1 Hawaiian monk seal Chain chain neoSch1 Hawaiian monk seal (Jun. 2017 (ASM220157v1/neoSch1)) Chained Alignments 3 21 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Hawaiian monk seal (Jun. 2017 (ASM220157v1/neoSch1)) Chained Alignments\ otherDb neoSch1\ parent placentalChainNetViewchain off\ shortLabel Hawaiian monk seal Chain\ subGroups view=chain species=s044 clade=c01\ track chainNeoSch1\ type chain neoSch1\ chainChlSab2 Green monkey Chain chain chlSab2 Green monkey (Mar. 2014 (Chlorocebus_sabeus 1.1/chlSab2)) Chained Alignments 3 21 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Green monkey (Mar. 2014 (Chlorocebus_sabeus 1.1/chlSab2)) Chained Alignments\ otherDb chlSab2\ parent primateChainNetViewchain off\ shortLabel Green monkey Chain\ subGroups view=chain species=s029 clade=c01\ track chainChlSab2\ type chain chlSab2\ encTfChipPkENCFF814DAF A549 HDAC2 narrowPeak Transcription Factor ChIP-seq Peaks of HDAC2 in A549 from ENCODE 3 (ENCFF814DAF) 0 21 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of HDAC2 in A549 from ENCODE 3 (ENCFF814DAF)\ parent encTfChipPk off\ shortLabel A549 HDAC2\ subGroups cellType=A549 factor=HDAC2\ track encTfChipPkENCFF814DAF\ wgEncodeReg4MarkH3k4me3AllAdrenalGland Adrenal gland (all biosamples) bigWig Avg. H3K4me3 level of 8 adrenal gland experiments (all biosamples) 0 21 90 179 68 172 217 161 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adrenalGlandH3K4me3.bw\ color 90,179,68\ longLabel Avg. H3K4me3 level of 8 adrenal gland experiments (all biosamples)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 21\ shortLabel Adrenal gland (all biosamples)\ track wgEncodeReg4MarkH3k4me3AllAdrenalGland\ type bigWig\ wgEncodeRegDnaseUwAg09319Peak AG09319 Pk narrowPeak AG09319 gingival fibroblast DNaseI Peaks from ENCODE 1 21 255 221 85 255 238 170 1 0 0 regulation 1 color 255,221,85\ longLabel AG09319 gingival fibroblast DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel AG09319 Pk\ subGroups view=a_Peaks cellType=AG09319 treatment=n_a tissue=periodontium cancer=normal\ track wgEncodeRegDnaseUwAg09319Peak\ wgEncodeRegDnaseUwAg09319Wig AG09319 Sg bigWig 0 28099 AG09319 gingival fibroblast DNaseI Signal from ENCODE 0 21 255 221 85 255 238 170 0 0 0 regulation 1 color 255,221,85\ longLabel AG09319 gingival fibroblast DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.17288\ shortLabel AG09319 Sg\ subGroups cellType=AG09319 treatment=n_a tissue=periodontium cancer=normal\ table wgEncodeRegDnaseUwAg09319Signal\ track wgEncodeRegDnaseUwAg09319Wig\ type bigWig 0 28099\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_ctss_fwd AorticSmsToFgf2_00hr45minBr3+ bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_forward 0 21 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep3%20%28LK12%29.CNhs13571.12841-137B6.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12841-137B6 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr45minBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_tpm_fwd AorticSmsToFgf2_00hr45minBr3+ bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_forward 1 21 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep3%20%28LK12%29.CNhs13571.12841-137B6.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12841-137B6 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr45minBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6\ urlLabel FANTOM5 Details:\ gtexCovBreastMammaryTissue Breast Mammary bigWig Breast Mammary Tissue 0 21 0 205 205 127 230 230 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-ZT9W-2026-SM-51MRA.Breast_Mammary_Tissue.RNAseq.bw\ color 0,205,205\ longLabel Breast Mammary Tissue\ parent gtexCov\ shortLabel Breast Mammary\ track gtexCovBreastMammaryTissue\ cloneEndCOR2A COR2A bed 12 NHGRI-CORIELLE CORIELL-02A-F-39-40KB 0 21 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel NHGRI-CORIELLE CORIELL-02A-F-39-40KB\ parent cloneEndSuper off\ priority 19\ shortLabel COR2A\ subGroups source=corielle\ track cloneEndCOR2A\ type bed 12\ visibility hide\ ENCFF280RMA_ENCFF651WOM_ENCFF262UEH_ENCFF850MLW ENCFF280RMA_ENCFF651WOM_ENCFF262UEH_ENCFF850MLW bigBed 9 + 5 SK-N-SH: (1) cCREs 4 21 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF280RMA_ENCFF651WOM_ENCFF262UEH_ENCFF850MLW.bb\ longLabel SK-N-SH: (1) cCREs\ mouseOver ID: ${name}\ The Genome of the Netherlands (GoNL) is a\ whole-genome sequencing project covering the Dutch population. The cohort was drawn from five\ Dutch biobanks and includes 250 parent-offspring families (231 trios and 19 quartets) from 11 of\ the 12 Dutch provinces. Samples were not selected by phenotype or disease status. This track\ shows allele counts and frequencies from the GRCh38 re-analysis of GoNL, restricted to the 498\ unrelated parents (250 fathers and 248 mothers; two mothers failed QC in the original release).\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API\ can be used; the track name is gonl.\ For bulk download, the VCF is available from\ our download\ server. The original file is also available from\ the GoNL download directory at MolGenis.\
\ \\ The track shown here uses the GRCh38 re-analysis (version 1.0). All samples were re-aligned from\ raw reads to a GRCh38 analysis set (GRCh38_no_alt_plus_hs38d1 with PhiX as decoy). The processing\ pipeline is documented in the\ README accompanying the data and differs from the original Nature Genetics\ pipeline (reference below). Per-library reads were trimmed with cutadapt 1.13, aligned with bwa mem 0.7.15, sorted\ with Picard SortSam 2.9.0, and base-quality-recalibrated with GATK BaseRecalibrator 3.7. Per-sample\ files were merged and deduplicated with sambamba 0.6.6, and variants were called per sample with\ GATK HaplotypeCaller 3.7. Per-family GVCFs were merged with GATK CombineGVCFs, and all families\ were jointly genotyped with GATK GenotypeGVCFs 3.7. The GRCh38 callset has not been filtered with\ VQSR and missing genotypes have not been imputed, so it is rougher than the original GRCh37\ release.\
\\ The file\ multisample.parents_only.info_only.vcf.gz was downloaded from\ https://download.molgeniscloud.org/downloads/gonl_public/variants/GoNL_GRCh38_1.0/.\ Of the 31,114,481 records in the source file, 30,904,161 were kept after dropping calls on the\ GRCh38 decoy contigs (chrUn_JTFH01* and similar) and the EBV contig, which are not part of the\ UCSC hg38 assembly. The original chromosome naming already uses the UCSC chr prefix, so\ no renaming was needed. The 2,629,361 multiallelic sites were then split with\ bcftools norm -m-any, with indels left-aligned against the hg38 reference, yielding\ 36,363,474 biallelic records (3,559,402 indels realigned). The maximum observed allele number\ (AN) is 996, which matches the 498 diploid parents in the cohort. Loading documentation is in the\ varFreqs makeDoc file; helper scripts for the broader varFreqs collection\ are in our GitHub scripts directory.\
\ \\ Data was generated by the Genome of the Netherlands Consortium and distributed via the\ MOLGENIS infrastructure at the\ University Medical Center Groningen. Thanks to the participants who donated samples and to the\ BBMRI-NL biobanks: LifeLines, Leiden Longevity Study, Netherlands Twin Registry, Rotterdam Study\ and Rucphen Study.\
\ \\ Genome of the Netherlands Consortium.\ \ Whole-genome sequence variation, population structure and demographic history of the Dutch\ population.\ Nat Genet. 2014 Aug;46(8):818-25.\ PMID: 24974849\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/gonl/gonl.vcf.gz\ dataVersion GRCh38 1.0\ longLabel SNV Frequencies: Genome of the Netherlands - 250 Dutch trios\ parent varFreqs on\ priority 21\ shortLabel Netherlands GoNL 498 WGS\ track gonl\ type vcfTabix\ visibility hide\ OV OV bigLolly 12 + Ovarian serous cystadenocarcinoma 0 21 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/OV.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Ovarian serous cystadenocarcinoma\ parent gdcCancer off\ priority 21\ shortLabel OV\ track OV\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4MarkCtcfAllOvary Ovary (all biosamples) bigWig Avg. CTCF level of 2 ovary experiments (all biosamples) 0 21 161 126 151 208 190 203 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/ovaryCTCF.bw\ color 161,126,151\ longLabel Avg. CTCF level of 2 ovary experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 21\ shortLabel Ovary (all biosamples)\ track wgEncodeReg4MarkCtcfAllOvary\ type bigWig\ wgEncodeReg4DnaseProstate Prostate bigWig Avg. DNase level of 2 prostate experiments (tissues and primary cells only) 0 21 140 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpProstateDNase.bw\ color 140,140,140\ longLabel Avg. DNase level of 2 prostate experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 21\ shortLabel Prostate\ track wgEncodeReg4DnaseProstate\ type bigWig\ Agilent_Human_Exon_V6_Regions SureSel. V6 T bigBed Agilent - SureSelect All Exon V6 r2 Target Regions 0 21 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S07604514_Regions.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V6 r2 Target Regions\ parent exomeProbesets off\ shortLabel SureSel. V6 T\ track Agilent_Human_Exon_V6_Regions\ type bigBed\ wgEncodeReg4AtacAllTestis Testis (all biosamples) bigWig ATAC level of 1 testis experiment (all biosamples) 0 21 139 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/testisATAC.bw\ color 139,140,140\ longLabel ATAC level of 1 testis experiment (all biosamples)\ parent wgEncodeReg4Atac off\ priority 21\ shortLabel Testis (all biosamples)\ track wgEncodeReg4AtacAllTestis\ type bigWig\ lincRNAsCTThyroid Thyroid bed 5 + lincRNAs from thyroid 1 21 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from thyroid\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel Thyroid\ subGroups view=lincRNAsRefseqExp tissueType=thyroid\ track lincRNAsCTThyroid\ netNeoSch1 Hawaiian monk seal Net netAlign neoSch1 chainNeoSch1 Hawaiian monk seal (Jun. 2017 (ASM220157v1/neoSch1)) Alignment Net 1 22 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Hawaiian monk seal (Jun. 2017 (ASM220157v1/neoSch1)) Alignment Net\ otherDb neoSch1\ parent placentalChainNetViewnet off\ shortLabel Hawaiian monk seal Net\ subGroups view=net species=s044 clade=c01\ track netNeoSch1\ type netAlign neoSch1 chainNeoSch1\ netChlSab2 Green monkey Net netAlign chlSab2 chainChlSab2 Green monkey (Mar. 2014 (Chlorocebus_sabeus 1.1/chlSab2)) Alignment Net 1 22 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Green monkey (Mar. 2014 (Chlorocebus_sabeus 1.1/chlSab2)) Alignment Net\ otherDb chlSab2\ parent primateChainNetViewnet off\ shortLabel Green monkey Net\ subGroups view=net species=s029 clade=c01\ track netChlSab2\ type netAlign chlSab2 chainChlSab2\ encTfChipPkENCFF127HJG A549 JUN narrowPeak Transcription Factor ChIP-seq Peaks of JUN in A549 from ENCODE 3 (ENCFF127HJG) 0 22 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of JUN in A549 from ENCODE 3 (ENCFF127HJG)\ parent encTfChipPk off\ shortLabel A549 JUN\ subGroups cellType=A549 factor=JUN\ track encTfChipPkENCFF127HJG\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_ctss_rev AorticSmsToFgf2_00hr45minBr3- bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_reverse 0 22 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep3%20%28LK12%29.CNhs13571.12841-137B6.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12841-137B6 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_00hr45minBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_tpm_rev AorticSmsToFgf2_00hr45minBr3- bigWig Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_reverse 1 22 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2000hr45min%2c%20biol_rep3%20%28LK12%29.CNhs13571.12841-137B6.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 00hr45min, biol_rep3 (LK12)_CNhs13571_12841-137B6_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12841-137B6 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_00hr45minBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF200hr45minBiolRep3LK12_CNhs13571_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12841-137B6\ urlLabel FANTOM5 Details:\ cloneEndbadEnds Bad end mappings bed 12 Clone end placements dropped at UCSC, map distance 3X median library size 0 22 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel Clone end placements dropped at UCSC, map distance 3X median library size\ parent cloneEndSuper off\ priority 24\ shortLabel Bad end mappings\ subGroups source=placements\ track cloneEndbadEnds\ type bed 12\ visibility hide\ wgEncodeReg4MarkH3k4me3AllBloodVessel Blood vessel (all biosamples) bigWig Avg. H3K4me3 level of 14 blood vessel experiments (all biosamples) 0 22 255 37 41 255 146 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodVesselH3K4me3.bw\ color 255,37,41\ longLabel Avg. H3K4me3 level of 14 blood vessel experiments (all biosamples)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 22\ shortLabel Blood vessel (all biosamples)\ track wgEncodeReg4MarkH3k4me3AllBloodVessel\ type bigWig\ gtexCovCellsEBV-transformedlymphocytes Cells EBV lymphoc bigWig Cells EBV-transformed lymphocytes 0 22 238 130 238 246 192 246 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-1122O-0003-SM-5Q5DL.Cells_EBV-transformed_lymphocytes.RNAseq.bw\ color 238,130,238\ longLabel Cells EBV-transformed lymphocytes\ parent gtexCov\ shortLabel Cells EBV lymphoc\ track gtexCovCellsEBV-transformedlymphocytes\ ENCFF286QGB_ENCFF835JIA_ENCFF618RAO_ENCFF700SCP ENCFF286QGB_ENCFF835JIA_ENCFF618RAO_ENCFF700SCP bigBed 9 + 5 Neural progenitor cell, female embryo (5 days): (1) cCREs 4 22 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF286QGB_ENCFF835JIA_ENCFF618RAO_ENCFF700SCP.bb\ longLabel Neural progenitor cell, female embryo (5 days): (1) cCREs\ mouseOver ID: ${name}\ NHLBI TOPMed (Trans-Omics for Precision\ Medicine) is a program launched by the U.S. National Heart, Lung, and Blood Institute that\ integrates whole-genome sequencing with molecular, clinical, and environmental data from large,\ well-phenotyped cohorts. Its goal is to uncover the biological mechanisms underlying heart, lung,\ blood, and sleep disorders to advance precision medicine and improve population health. Freeze 10\ contains 868,581,653 variants from 150,899 whole genomes.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ VCFs with summarized allele frequencies are available from\ the TOPMED BRAVO website. They require a\ login. The VCFs were downloaded from\ BRAVO.\
\ \\
TOPMed whole genome sequencing was performed at multiple NHLBI-funded sequencing centers\
using PCR-free library preparation with 150 bp paired-end reads on Illumina short-read\
platforms, targeting ≥30x mean coverage. Reads were aligned to the GRCh38 reference genome\
(hs38DH, including decoy sequences) using BWA-MEM, followed by duplicate marking with\
Picard MarkDuplicates and base quality score recalibration (BQSR) with GATK. Variant calling\
was performed using the TOPMed GotCloud pipeline (developed at the Center for Statistical\
Genetics, University of Michigan), comprising: (1) per-sample candidate variant detection with\
vt discover2 and normalization with vt normalize; (2) cross-sample variant site\
consolidation using cramore vcf-merge-candidate-variants; (3) joint genotyping across all\
samples; and (4) variant filtering using a Support Vector Machine (SVM) classifier\
(libsvm) trained on positive labels derived from HapMap 3.3 and 1000 Genomes Omni2.5\
array sites, and negative labels derived from Mendelian-inconsistent variants identified\
within the cohort's pedigree structure using vt milk-filter. Sample-level quality\
control included estimation of DNA contamination, genetic ancestry, and biological sex\
using cramore cram-verify-bam (verifyBamID2) and relative X/Y chromosomal depth. Full\
methods for TOPMed freeze 10 are available on the\
TOPMed WGS Methods page.\
\ Documentation on how all source files of the varFreqs track were converted is in the makeDoc file of the track.\ For some tracks, python scripts were necessary and are also available from GitHub.\
\ \\ Taliun D, Harris DN, Kessler MD, Carlson J, Szpiech ZA, Torres R, Taliun SAG, Corvelo A, Gogarten SM,\ Kang HM et al.\ \ Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program.\ Nature. 2021 Feb;590(7845):290-299.\ PMID: 33568819; PMC: PMC7875770\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_topmed/topmed10.vcf.gz\ dataVersion Freeze 10\ longLabel SNV Frequencies: NHLBI TOPMed - 151k WGS\ parent varFreqs on\ priority 22\ shortLabel NHLBI TOPMed 10 151k WGS\ tableBrowser off\ track topmed\ type vcfTabix\ visibility hide\ PAAD PAAD bigLolly 12 + Pancreatic adenocarcinoma 0 22 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/PAAD.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Pancreatic adenocarcinoma\ parent gdcCancer off\ priority 22\ shortLabel PAAD\ track PAAD\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4MarkCtcfAllParaythroidGland Parathyroid gland (all biosamples) bigWig Avg. CTCF level of 2 parathyroid gland experiments (all biosamples) 0 22 130 141 158 192 198 206 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/paraythroidGlandCTCF.bw\ color 130,141,158\ longLabel Avg. CTCF level of 2 parathyroid gland experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 22\ shortLabel Parathyroid gland (all biosamples)\ track wgEncodeReg4MarkCtcfAllParaythroidGland\ type bigWig\ wgEncodeReg4DnaseSkin Skin bigWig Avg. DNase level of 22 skin experiments (tissues and primary cells only) 0 22 127 133 209 191 194 232 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpSkinDNase.bw\ color 127,133,209\ longLabel Avg. DNase level of 22 skin experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 22\ shortLabel Skin\ track wgEncodeReg4DnaseSkin\ type bigWig\ smoothMuscMerged Smooth Muscle Merged bigWig Methylation Atlas: Smooth Muscle Merged Samples 2 22 205 92 92 230 173 173 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/smoothMuscMerged.bw\ color 205,92,92\ longLabel Methylation Atlas: Smooth Muscle Merged Samples\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 22\ shortLabel Smooth Muscle Merged\ subGroups cellType=Smooth-Musc dataType=Merged\ track smoothMuscMerged\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ Agilent_Human_Exon_V6_COSMIC_Covered SureSel. V6+COSMIC P bigBed Agilent - SureSelect All Exon V6 + COSMIC r2 Covered by Probes 0 22 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S07604715_Covered.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V6 + COSMIC r2 Covered by Probes\ parent exomeProbesets off\ shortLabel SureSel. V6+COSMIC P\ track Agilent_Human_Exon_V6_COSMIC_Covered\ type bigBed\ wgEncodeReg4AtacAllThyroid Thyroid (all biosamples) bigWig Avg. ATAC level of 3 thyroid experiments (all biosamples) 0 22 27 119 58 141 187 156 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/thyroidATAC.bw\ color 27,119,58\ longLabel Avg. ATAC level of 3 thyroid experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 22\ shortLabel Thyroid (all biosamples)\ track wgEncodeReg4AtacAllThyroid\ type bigWig\ lincRNAsCTWhiteBloodCell WhiteBloodCell bed 5 + lincRNAs from whitebloodcell 1 22 0 60 120 127 157 187 1 0 0 genes 1 longLabel lincRNAs from whitebloodcell\ origAssembly hg19\ parent lincRNAsAllCellType on\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel WhiteBloodCell\ subGroups view=lincRNAsRefseqExp tissueType=whitebloodcell\ track lincRNAsCTWhiteBloodCell\ chainSaiBol1 saiBol1 Chain chain saiBol1 Squirrel monkey (Oct. 2011 (Broad/saiBol1)) Chained Alignments 3 23 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Squirrel monkey (Oct. 2011 (Broad/saiBol1)) Chained Alignments\ otherDb saiBol1\ parent primateChainNetViewchain off\ shortLabel saiBol1 Chain\ subGroups view=chain species=s032 clade=c02\ track chainSaiBol1\ type chain saiBol1\ chainBosTau9 Cow Chain chain bosTau9 Cow (Apr. 2018 (ARS-UCD1.2/bosTau9)) Chained Alignments 3 23 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Cow (Apr. 2018 (ARS-UCD1.2/bosTau9)) Chained Alignments\ otherDb bosTau9\ parent placentalChainNetViewchain off\ shortLabel Cow Chain\ subGroups view=chain species=s051a clade=c02\ track chainBosTau9\ type chain bosTau9\ phastConsElements100way 100 Vert. El bed 5 . 100 vertebrates Conserved Elements 0 23 110 10 40 182 132 147 0 0 0 compGeno 1 color 110,10,40\ longLabel 100 vertebrates Conserved Elements\ noInherit on\ parent cons100wayViewelements off\ priority 23\ shortLabel 100 Vert. El\ subGroups view=elements\ track phastConsElements100way\ type bed 5 .\ phastConsElements30way 30-way El bed 5 . 30 mammals Conserved Elements (27 primates) 1 23 110 10 40 182 132 147 0 0 0 compGeno 1 color 110,10,40\ longLabel 30 mammals Conserved Elements (27 primates)\ noInherit on\ parent cons30wayViewelements on\ priority 23\ shortLabel 30-way El\ subGroups view=elements\ track phastConsElements30way\ type bed 5 .\ encTfChipPkENCFF587VEY A549 JUND narrowPeak Transcription Factor ChIP-seq Peaks of JUND in A549 from ENCODE 3 (ENCFF587VEY) 0 23 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of JUND in A549 from ENCODE 3 (ENCFF587VEY)\ parent encTfChipPk off\ shortLabel A549 JUND\ subGroups cellType=A549 factor=JUND\ track encTfChipPkENCFF587VEY\ aortaSmMusc41U Aorta - Smooth Muscle - Z0000041U bigWig Methylation Atlas: Aorta - Smooth Muscle - Z0000041U 2 23 205 92 92 230 173 173 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/aortaSmMusc41U.bw\ color 205,92,92\ longLabel Methylation Atlas: Aorta - Smooth Muscle - Z0000041U\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 23\ shortLabel Aorta - Smooth Muscle - Z0000041U\ subGroups cellType=Smooth-Musc dataType=Replicate\ track aortaSmMusc41U\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_ctss_fwd AorticSmsToFgf2_01hrBr1+ bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_forward 0 23 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep1%20%28LK13%29.CNhs12741.12646-134G9.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12646-134G9 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_01hrBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_tpm_fwd AorticSmsToFgf2_01hrBr1+ bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_forward 1 23 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep1%20%28LK13%29.CNhs12741.12646-134G9.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12646-134G9 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_01hrBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9\ urlLabel FANTOM5 Details:\ gtexCovCellsCulturedfibroblasts Cells fibrobl cult bigWig Cells Cultured fibroblasts 0 23 154 192 205 204 223 230 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-117XS-0008-SM-5Q5DQ.Cells_Cultured_fibroblasts.RNAseq.bw\ color 154,192,205\ longLabel Cells Cultured fibroblasts\ parent gtexCov\ shortLabel Cells fibrobl cult\ track gtexCovCellsCulturedfibroblasts\ cloneEndcoverageForward Coverage forward bigWig 0 5377 Clone end placements overlap coverage on the forward strand 2 23 0 0 0 127 127 127 0 0 0 map 0 alwaysZero on\ autoScale on\ longLabel Clone end placements overlap coverage on the forward strand\ maxHeightPixels 128:35:16\ parent cloneEndSuper off\ priority 25\ shortLabel Coverage forward\ subGroups source=placements\ track cloneEndcoverageForward\ type bigWig 0 5377\ visibility full\ windowingFunction mean\ ENCFF269VAY_ENCFF346LEZ_ENCFF118OBT_ENCFF536VOI ENCFF269VAY_ENCFF346LEZ_ENCFF118OBT_ENCFF536VOI bigBed 9 + 5 Glutamatergic neuron, male adult (53 years) male adult (53 years) nuclear fraction: (1) cCREs 4 23 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF269VAY_ENCFF346LEZ_ENCFF118OBT_ENCFF536VOI.bb\ longLabel Glutamatergic neuron, male adult (53 years) male adult (53 years) nuclear fraction: (1) cCREs\ mouseOver ID: ${name}\ Variant frequencies from 302 whole genomes at 30x coverage from the\ Saudi Genome Program. The genotyping data and imputations from 3,352\ individuals do not seem to be available publicly.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is saudi.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ The original data were downloaded from\ Figshare and converted to VCF.\
\ \\ Whole-genome sequencing of 302 Saudi Arabian individuals was performed on the Illumina HiSeq\ X Ten platform using TruSeq Nano DNA library preparation at 30x target coverage. Sequencing and\ initial bioinformatics processing were carried out by deCODE Genetics (Reykjavík, Iceland).\ Reads were aligned to the GRCh38 reference genome using BWA 0.7.10. Per-sample variants\ were called with GATK HaplotypeCaller, then jointly genotyped with CombineGVCFs and\ GenotypeGVCFs. Variant quality score recalibration (VQSR) was applied for both SNPs and indels.\ The final autosomal callset contains 25.5 million variants across the 302 individuals.\
\\ The variant data were downloaded from\ Figshare and converted to VCF format using a custom script.\ The makeDoc file documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were needed; these are also available from GitHub.\
\ \\ Malomane DK, Williams MP, Huber CD, Mangul S, Abedalthagafi M, Chiang CWK.\ \ Patterns of population structure and genetic variation within the Saudi Arabian population.\ bioRxiv. 2025 Jan 13;.\ PMID: 39868174; PMC: PMC11761371\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/saudi/saudi.vcf.gz\ dataVersion SHGP (figshare 51297884, 2025)\ longLabel SNV Frequencies: Saudi Genome Project - 302 WGS samples\ parent varFreqs on\ priority 23\ shortLabel Saudi Genome 302 WGS\ track saudi\ type vcfTabix\ visibility hide\ Agilent_Human_Exon_V6_COSMIC_Regions SureSel. V6+COSMIC T bigBed Agilent - SureSelect All Exon V6 + COSMIC r2 Target Regions 0 23 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S07604715_Regions.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V6 + COSMIC r2 Target Regions\ parent exomeProbesets off\ shortLabel SureSel. V6+COSMIC T\ track Agilent_Human_Exon_V6_COSMIC_Regions\ type bigBed\ wgEncodeReg4DnaseTestis Testis bigWig Avg. DNase level of 4 testis experiments (tissues and primary cells only) 0 23 139 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpTestisDNase.bw\ color 139,140,140\ longLabel Avg. DNase level of 4 testis experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 23\ shortLabel Testis\ track wgEncodeReg4DnaseTestis\ type bigWig\ wgEncodeReg4AtacAllUrinaryBladder Urinary bladder (all biosamples) bigWig ATAC level of 1 urinary bladder experiment (all biosamples) 0 23 194 33 39 224 144 147 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/urinaryBladderATAC.bw\ color 194,33,39\ longLabel ATAC level of 1 urinary bladder experiment (all biosamples)\ parent wgEncodeReg4Atac off\ priority 23\ shortLabel Urinary bladder (all biosamples)\ track wgEncodeReg4AtacAllUrinaryBladder\ type bigWig\ netSaiBol1 saiBol1 Net netAlign saiBol1 chainSaiBol1 Squirrel monkey (Oct. 2011 (Broad/saiBol1)) Alignment Net 1 24 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Squirrel monkey (Oct. 2011 (Broad/saiBol1)) Alignment Net\ otherDb saiBol1\ parent primateChainNetViewnet off\ shortLabel saiBol1 Net\ subGroups view=net species=s032 clade=c02\ track netSaiBol1\ type netAlign saiBol1 chainSaiBol1\ netBosTau9 Cow Net netAlign bosTau9 chainBosTau9 Cow (Apr. 2018 (ARS-UCD1.2/bosTau9)) Alignment Net 1 24 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Cow (Apr. 2018 (ARS-UCD1.2/bosTau9)) Alignment Net\ otherDb bosTau9\ parent placentalChainNetViewnet off\ shortLabel Cow Net\ subGroups view=net species=s051a clade=c02\ track netBosTau9\ type netAlign bosTau9 chainBosTau9\ encTfChipPkENCFF316CBQ A549 KDM1A narrowPeak Transcription Factor ChIP-seq Peaks of KDM1A in A549 from ENCODE 3 (ENCFF316CBQ) 0 24 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of KDM1A in A549 from ENCODE 3 (ENCFF316CBQ)\ parent encTfChipPk off\ shortLabel A549 KDM1A\ subGroups cellType=A549 factor=KDM1A\ track encTfChipPkENCFF316CBQ\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_ctss_rev AorticSmsToFgf2_01hrBr1- bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_reverse 0 24 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep1%20%28LK13%29.CNhs12741.12646-134G9.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12646-134G9 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_01hrBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_tpm_rev AorticSmsToFgf2_01hrBr1- bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_reverse 1 24 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep1%20%28LK13%29.CNhs12741.12646-134G9.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep1 (LK13)_CNhs12741_12646-134G9_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12646-134G9 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_01hrBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep1LK13_CNhs12741_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12646-134G9\ urlLabel FANTOM5 Details:\ bladderSmMusc41Z Bladder - Smooth Muscle - Z0000041Z bigWig Methylation Atlas: Bladder - Smooth Muscle - Z0000041Z 2 24 205 92 92 230 173 173 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/bladderSmMusc41Z.bw\ color 205,92,92\ longLabel Methylation Atlas: Bladder - Smooth Muscle - Z0000041Z\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 24\ shortLabel Bladder - Smooth Muscle - Z0000041Z\ subGroups cellType=Smooth-Musc dataType=Replicate\ track bladderSmMusc41Z\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ gtexCovCervixEctocervix Cervix Ectocerv bigWig Cervix Ectocervix 0 24 238 213 210 246 234 232 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-S341-1126-SM-4AD6T.Cervix_Ectocervix.RNAseq.bw\ color 238,213,210\ longLabel Cervix Ectocervix\ parent gtexCov\ shortLabel Cervix Ectocerv\ track gtexCovCervixEctocervix\ cloneEndcoverageReverse Coverage reverse bigWig 0 4112 Clone end placements overlap coverage on the reverse strand 2 24 0 0 0 127 127 127 0 0 0 map 0 alwaysZero on\ autoScale on\ longLabel Clone end placements overlap coverage on the reverse strand\ maxHeightPixels 128:35:16\ negateValues 1\ parent cloneEndSuper off\ priority 26\ shortLabel Coverage reverse\ subGroups source=placements\ track cloneEndcoverageReverse\ type bigWig 0 4112\ visibility full\ windowingFunction mean\ ENCFF386FNE_ENCFF768NPJ_ENCFF435NQW_ENCFF541XGP ENCFF386FNE_ENCFF768NPJ_ENCFF435NQW_ENCFF541XGP bigBed 9 + 5 Bipolar neuron (treated), male adult (53 years) treated with 0.5 μg/mL doxycycline hyclate for 4 days: (1) cCREs 4 24 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF386FNE_ENCFF768NPJ_ENCFF435NQW_ENCFF541XGP.bb\ longLabel Bipolar neuron (treated), male adult (53 years) treated with 0.5 μg/mL doxycycline hyclate for 4 days: (1) cCREs\ mouseOver ID: ${name}\ The SCHEMA (Schizophrenia Exome\ Meta-Analysis) consortium is an international collaboration that aggregated and harmonized\ whole-exome sequencing data to study the role of rare coding variants in schizophrenia.\ The dataset includes 24,248 cases and 97,322 controls from diverse global cohorts.\ SCHEMA identified genes with exome-wide significant rare variant burden in schizophrenia,\ which point to the biology of the disorder.\
\ \\ Since the data can be downloaded from the SCHEMA website, and does not seem to be under a license,\ we assume that we are allowed to redistribute it in VCF format.\ The data can be explored on our website interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is schema.\ For bulk download, the VCF file can be obtained from\ our download server.\
\\ Summary statistics and variant-level results are also available from the\ SCHEMA Browser.\
\ \\ The SCHEMA (Schizophrenia Exome Meta-Analysis) consortium aggregated whole-exome sequencing\ data from 24,248 schizophrenia cases and 97,322 controls (including non-psychiatric,\ non-neurological samples from the gnomAD consortium) across multiple international cohorts.\ Exome sequencing was performed using various capture platforms and Illumina sequencing\ instruments across cohorts sequenced over approximately a decade. Sequence data were\ uniformly reprocessed through the BWA-Picard-GATK best practices pipeline as part of the\ gnomAD v2 infrastructure, including alignment to GRCh37/hg19, duplicate marking, base\ quality score recalibration, and per-sample variant calling with GATK HaplotypeCaller,\ followed by joint genotyping across all samples. A novel exon-by-exon coverage estimation\ pipeline was developed to account for differences in capture technology across sequencing\ batches, and both site-level and genotype-level quality filters were applied. Protein-truncating\ variants (PTVs) were annotated using LOFTEE (Loss-Of-Function Transcript Effect Estimator),\ and missense variant deleteriousness was scored using MPC (Missense badness, PolyPhen-2,\ and Constraint). Gene-level association testing combined: (1) a case-control rare variant\ burden test aggregating ultra-rare PTVs (Class I: PTV and MPC > 3; Class II: missense\ MPC 2–3) across 18,321 protein-coding genes; and (2) de novo variant enrichment\ from 3,402 schizophrenia proband-parent trios assessed via a Poisson rate test against\ gnomAD-derived baseline mutation rates; with the two components combined using a weighted\ Z-score meta-analysis. This identified 10 genes at exome-wide significance (P < 2.14\ × 10-6) with odds ratios for PTVs ranging from 3 to 50, and 32 genes at\ FDR < 5%. Full data are available at\ schema.broadinstitute.org\ (Singh, Neale, Daly & the SCHEMA Consortium,\ Nature 2022).\
\\ We downloaded the TSV data from the SCHEMA website\ and converted it to VCF format using a custom Python script. The VCF was lifted to hg38 using our hg19ToHg38 chain\ file. \ We provide documentation that indicates how all source files of the varFreqs track were converted in the makeDoc file of the track.\ For some tracks, python scripts were necessary and are also available from GitHub.\
\ \\ Singh T, Poterba T, Curtis D, Akil H, Al Eissa M, Barchas JD, Bass N, Bigdeli TB, Breen G,\ Bromet EJ et al.\ \ Exome sequencing identifies rare coding variants in 10 genes which confer substantial risk for\ schizophrenia.\ Nature. 2022 Apr;604(7906):509-516.\ PMID: 35396579; PMC: PMC9392855\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/schema/SCHEMA_variant_results_withAF.vcf.gz\ dataVersion 2022\ longLabel SNV Frequencies: SCHEMA Schizophrenia Exome Meta-Analysis - WES 24k cases, 97k controls\ parent varFreqs on\ priority 24\ shortLabel SCHEMA 121k WES Sz\ track schema\ type vcfTabix\ url https://schema.broadinstitute.org/\ urlLabel SCHEMA Browser\ visibility hide\ wgEncodeReg4MarkCtcfAllSmallIntestine Small intestine (all biosamples) bigWig Avg. CTCF level of 4 small intestine experiments (all biosamples) 0 24 98 98 41 176 176 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/smallIntestineCTCF.bw\ color 98,98,41\ longLabel Avg. CTCF level of 4 small intestine experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 24\ shortLabel Small intestine (all biosamples)\ track wgEncodeReg4MarkCtcfAllSmallIntestine\ type bigWig\ Agilent_Human_Exon_V6_UTRs_Covered SureSel. V6+UTR P bigBed Agilent - SureSelect All Exon V6 + UTR r2 Covered by Probes 0 24 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S07604624_Covered.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V6 + UTR r2 Covered by Probes\ parent exomeProbesets off\ shortLabel SureSel. V6+UTR P\ track Agilent_Human_Exon_V6_UTRs_Covered\ type bigBed\ wgEncodeReg4DnaseUterus Uterus bigWig Avg. DNase level of 2 uterus experiments (tissues and primary cells only) 0 24 186 111 165 220 183 210 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/tpUterusDNase.bw\ color 186,111,165\ longLabel Avg. DNase level of 2 uterus experiments (tissues and primary cells only)\ parent wgEncodeReg4Dnase off\ priority 24\ shortLabel Uterus\ track wgEncodeReg4DnaseUterus\ type bigWig\ wgEncodeReg4AtacAllUterus Uterus (all biosamples) bigWig ATAC level of 1 uterus experiment (all biosamples) 0 24 186 111 165 220 183 210 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/uterusATAC.bw\ color 186,111,165\ longLabel ATAC level of 1 uterus experiment (all biosamples)\ parent wgEncodeReg4Atac off\ priority 24\ shortLabel Uterus (all biosamples)\ track wgEncodeReg4AtacAllUterus\ type bigWig\ chainOviAri4 Sheep Chain chain oviAri4 Sheep (Nov. 2015 (Oar_v4.0/oviAri4)) Chained Alignments 3 25 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Sheep (Nov. 2015 (Oar_v4.0/oviAri4)) Chained Alignments\ otherDb oviAri4\ parent placentalChainNetViewchain off\ shortLabel Sheep Chain\ subGroups view=chain species=s065b clade=c02\ track chainOviAri4\ type chain oviAri4\ chainCalJac4 Marmoset Chain chain calJac4 Marmoset (May 2020 (Callithrix_jacchus_cj1700_1.1/calJac4)) Chained Alignments 3 25 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Marmoset (May 2020 (Callithrix_jacchus_cj1700_1.1/calJac4)) Chained Alignments\ otherDb calJac4\ parent primateChainNetViewchain off\ shortLabel Marmoset Chain\ subGroups view=chain species=s034a clade=c02\ track chainCalJac4\ type chain calJac4\ encTfChipPkENCFF149INM A549 KDM5A narrowPeak Transcription Factor ChIP-seq Peaks of KDM5A in A549 from ENCODE 3 (ENCFF149INM) 0 25 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of KDM5A in A549 from ENCODE 3 (ENCFF149INM)\ parent encTfChipPk off\ shortLabel A549 KDM5A\ subGroups cellType=A549 factor=KDM5A\ track encTfChipPkENCFF149INM\ wgEncodeReg4DnaseAllAdipose Adipose (all biosamples) bigWig Avg. DNase level of 3 adipose experiments (all biosamples) 0 25 255 119 39 255 187 147 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adiposeDNase.bw\ color 255,119,39\ longLabel Avg. DNase level of 3 adipose experiments (all biosamples)\ parent wgEncodeReg4Dnase off\ priority 25\ shortLabel Adipose (all biosamples)\ track wgEncodeReg4DnaseAllAdipose\ type bigWig\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_ctss_fwd AorticSmsToFgf2_01hrBr3+ bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_forward 0 25 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep3%20%28LK15%29.CNhs13683.12842-137B7.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12842-137B7 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_01hrBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_tpm_fwd AorticSmsToFgf2_01hrBr3+ bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_forward 1 25 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep3%20%28LK15%29.CNhs13683.12842-137B7.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12842-137B7 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_01hrBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7\ urlLabel FANTOM5 Details:\ wgEncodeReg4AtacAllBlood Blood (all biosamples) bigWig Avg. ATAC level of 209 blood experiments (all biosamples) 0 25 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodATAC.bw\ color 254,75,173\ longLabel Avg. ATAC level of 209 blood experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 25\ shortLabel Blood (all biosamples)\ track wgEncodeReg4AtacAllBlood\ type bigWig\ gtexCovCervixEndocervix Cervix Endocerv bigWig Cervix Endocervix 0 25 238 213 210 246 234 232 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-ZPIC-1326-SM-DO91Y.Cervix_Endocervix.RNAseq.bw\ color 238,213,210\ longLabel Cervix Endocervix\ parent gtexCov\ shortLabel Cervix Endocerv\ track gtexCovCervixEndocervix\ coronArtSmMusc420 Coronary Artery - Smooth Muscle - Z00000420 bigWig Methylation Atlas: Coronary Artery - Smooth Muscle - Z00000420 2 25 205 92 92 230 173 173 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/coronArtSmMusc420.bw\ color 205,92,92\ longLabel Methylation Atlas: Coronary Artery - Smooth Muscle - Z00000420\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 25\ shortLabel Coronary Artery - Smooth Muscle - Z00000420\ subGroups cellType=Smooth-Musc dataType=Replicate\ track coronArtSmMusc420\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ ENCFF926MIK_ENCFF153BJG_ENCFF751GCN_ENCFF569HGW ENCFF926MIK_ENCFF153BJG_ENCFF751GCN_ENCFF569HGW bigBed 9 + 5 Astrocyte, male adult (53 years): (1) cCREs 4 25 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF926MIK_ENCFF153BJG_ENCFF751GCN_ENCFF569HGW.bb\ longLabel Astrocyte, male adult (53 years): (1) cCREs\ mouseOver ID: ${name}\ The Simons Foundation Autism Research\ Initiative (SFARI) recruited a large cohort of families with autistic children who provided\ DNA samples and phenotypes. 54,558 families, parents and their children were sequenced, a total\ of 142,357 individuals with whole-exome (WES) and 12,519 with whole-genome sequencing (WGS).\ The data contains 32,559 trios and 8,895 quads (one sibling without autism), and 824 twins.\
\ \\ The same frequencies shown here are also available publicly on the\ SFARI Genome Browser.\ See (SPARK et al, Neuron 2018) for details.\
\ \\ In addition to the overall allele count (AC), allele number (AN), and allele\ frequency (AF), each variant record carries counts split by autism status\ (the asd column of the SPARK individual registration file):\
\\ A small minority of samples have a blank asd value and so contribute\ only to the overall AC/AN/AF, not to either group total.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ Allele frequencies can also be displayed on the\ SFARI Genome Browser.\ Full CRAMs and VCFs with genotypes are available from\ SFARI Base.\ They require a data access request, which is usually reviewed quickly. More information is\ available in the\ SPARK Welcome Packet.\
\ \The genome browser track project was approved by the Simons Foundation under request\ number 14584.1. WES and WGS data were downloaded from\ SFARI Base.\ pVCFs were downloaded, anonymized with a script using bcftools and its "fill-tags" plugin and\ normalized. There was no minimum allele frequency cutoff.\ The ASD-status sample-group file derived from the SPARK individuals_registration\ TSV was passed to fill-tags via its -S option, which adds the per-group\ AC_AUT/AN_AUT/AF_AUT and AC_NON_AUT/AN_NON_AUT/AF_NON_AUT\ tags alongside the overall AC/AN/AF.
\ \The methods are documented as follows by SFARI:
\\ The makeDoc file documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were necessary and are also available from GitHub.\
\ \\ SPARK Consortium. Electronic address: pfeliciano@simonsfoundation.org, SPARK Consortium.\ \ SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research.\ Neuron. 2018 Feb 7;97(3):488-493.\ PMID: 29420931; PMC: PMC7444276\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_sfari/wgs_12519_genome.deepvariant.norm.vcf.gz\ dataVersion iWGS v1.1\ html sfariSparkExomes\ longLabel SNV Frequencies: SFARI SPARK - 12,519 WGS\ parent varFreqs on\ priority 25\ shortLabel SFARI SPARK 12k WGS\ tableBrowser off\ track sfariSparkWgs\ type vcfTabix\ visibility hide\ wgEncodeReg4MarkCtcfAllSpinalCord Spinal cord (all biosamples) bigWig CTCF level of 1 spinal cord experiment (all biosamples) 0 25 130 141 158 192 198 206 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/spinalCordCTCF.bw\ color 130,141,158\ longLabel CTCF level of 1 spinal cord experiment (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 25\ shortLabel Spinal cord (all biosamples)\ track wgEncodeReg4MarkCtcfAllSpinalCord\ type bigWig\ Agilent_Human_Exon_V7_Covered SureSel. V7 P bigBed Agilent - SureSelect All Exon V7 Covered by Probes 0 25 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S31285117_Covered.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V7 Covered by Probes\ parent exomeProbesets on\ shortLabel SureSel. V7 P\ track Agilent_Human_Exon_V7_Covered\ type bigBed\ netOviAri4 Sheep Net netAlign oviAri4 chainOviAri4 Sheep (Nov. 2015 (Oar_v4.0/oviAri4)) Alignment Net 1 26 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Sheep (Nov. 2015 (Oar_v4.0/oviAri4)) Alignment Net\ otherDb oviAri4\ parent placentalChainNetViewnet off\ shortLabel Sheep Net\ subGroups view=net species=s065b clade=c02\ track netOviAri4\ type netAlign oviAri4 chainOviAri4\ netCalJac4 Marmoset Net netAlign calJac4 chainCalJac4 Marmoset (May 2020 (Callithrix_jacchus_cj1700_1.1/calJac4)) Alignment Net 1 26 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Marmoset (May 2020 (Callithrix_jacchus_cj1700_1.1/calJac4)) Alignment Net\ otherDb calJac4\ parent primateChainNetViewnet on\ shortLabel Marmoset Net\ subGroups view=net species=s034a clade=c02\ track netCalJac4\ type netAlign calJac4 chainCalJac4\ encTfChipPkENCFF813WJW A549 MAFK narrowPeak Transcription Factor ChIP-seq Peaks of MAFK in A549 from ENCODE 3 (ENCFF813WJW) 0 26 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of MAFK in A549 from ENCODE 3 (ENCFF813WJW)\ parent encTfChipPk off\ shortLabel A549 MAFK\ subGroups cellType=A549 factor=MAFK\ track encTfChipPkENCFF813WJW\ wgEncodeReg4DnaseAllAdrenalGland Adrenal gland (all biosamples) bigWig Avg. DNase level of 14 adrenal gland experiments (all biosamples) 0 26 90 179 68 172 217 161 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/adrenalGlandDNase.bw\ color 90,179,68\ longLabel Avg. DNase level of 14 adrenal gland experiments (all biosamples)\ parent wgEncodeReg4Dnase off\ priority 26\ shortLabel Adrenal gland (all biosamples)\ track wgEncodeReg4DnaseAllAdrenalGland\ type bigWig\ wgEncodeRegDnaseUwAoafPeak AoAF Pk narrowPeak AoAF aorta fibroblast DNaseI Peaks from ENCODE 1 26 255 236 85 255 245 170 1 0 0 regulation 1 color 255,236,85\ longLabel AoAF aorta fibroblast DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel AoAF Pk\ subGroups view=a_Peaks cellType=AoAF treatment=n_a tissue=blood_vessel cancer=normal\ track wgEncodeRegDnaseUwAoafPeak\ wgEncodeRegDnaseUwAoafWig AoAF Sg bigWig 0 10369.5 AoAF aorta fibroblast DNaseI Signal from ENCODE 0 26 255 236 85 255 245 170 0 0 0 regulation 1 color 255,236,85\ longLabel AoAF aorta fibroblast DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.19489\ shortLabel AoAF Sg\ subGroups cellType=AoAF treatment=n_a tissue=blood_vessel cancer=normal\ table wgEncodeRegDnaseUwAoafSignal\ track wgEncodeRegDnaseUwAoafWig\ type bigWig 0 10369.5\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_ctss_rev AorticSmsToFgf2_01hrBr3- bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_reverse 0 26 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep3%20%28LK15%29.CNhs13683.12842-137B7.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12842-137B7 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_01hrBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_tpm_rev AorticSmsToFgf2_01hrBr3- bigWig Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_reverse 1 26 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2001hr%2c%20biol_rep3%20%28LK15%29.CNhs13683.12842-137B7.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 01hr, biol_rep3 (LK15)_CNhs13683_12842-137B7_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12842-137B7 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_01hrBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF201hrBiolRep3LK15_CNhs13683_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12842-137B7\ urlLabel FANTOM5 Details:\ wgEncodeReg4AtacAllBrain Brain (all biosamples) bigWig Avg. ATAC level of 14 brain experiments (all biosamples) 0 26 155 155 18 205 205 136 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/brainATAC.bw\ color 155,155,18\ longLabel Avg. ATAC level of 14 brain experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 26\ shortLabel Brain (all biosamples)\ track wgEncodeReg4AtacAllBrain\ type bigWig\ gtexCovColonSigmoid Colon Sigmoid bigWig Colon Sigmoid 0 26 205 183 158 230 219 206 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-1KXAM-1926-SM-D3LAG.Colon_Sigmoid.RNAseq.bw\ color 205,183,158\ longLabel Colon Sigmoid\ parent gtexCov\ shortLabel Colon Sigmoid\ track gtexCovColonSigmoid\ ENCFF963PFR_ENCFF577BWJ_ENCFF643ZMC_ENCFF714NPP ENCFF963PFR_ENCFF577BWJ_ENCFF643ZMC_ENCFF714NPP bigBed 9 + 5 Astrocyte: (1) cCREs 4 26 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF963PFR_ENCFF577BWJ_ENCFF643ZMC_ENCFF714NPP.bb\ longLabel Astrocyte: (1) cCREs\ mouseOver ID: ${name}\ The Simons Foundation Autism Research\ Initiative (SFARI) recruited a large cohort of families with autistic children who provided\ DNA samples and phenotypes. 54,558 families, parents and their children were sequenced, a total\ of 142,357 individuals with whole-exome (WES) and 12,519 with whole-genome sequencing (WGS).\ The data contains 32,559 trios and 8,895 quads (one sibling without autism), and 824 twins.\
\ \\ The same frequencies shown here are also available publicly on the\ SFARI Genome Browser.\ See (SPARK et al, Neuron 2018) for details.\
\ \\ In addition to the overall allele count (AC), allele number (AN), and allele\ frequency (AF), each variant record carries counts split by autism status\ (the asd column of the SPARK individual registration file):\
\\ A small minority of samples have a blank asd value and so contribute\ only to the overall AC/AN/AF, not to either group total.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ Allele frequencies can also be displayed on the\ SFARI Genome Browser.\ Full CRAMs and VCFs with genotypes are available from\ SFARI Base.\ They require a data access request, which is usually reviewed quickly. More information is\ available in the\ SPARK Welcome Packet.\
\ \The genome browser track project was approved by the Simons Foundation under request\ number 14584.1. WES and WGS data were downloaded from\ SFARI Base.\ pVCFs were downloaded, anonymized with a script using bcftools and its "fill-tags" plugin and\ normalized. There was no minimum allele frequency cutoff.\ The ASD-status sample-group file derived from the SPARK individuals_registration\ TSV was passed to fill-tags via its -S option, which adds the per-group\ AC_AUT/AN_AUT/AF_AUT and AC_NON_AUT/AN_NON_AUT/AF_NON_AUT\ tags alongside the overall AC/AN/AF.
\ \The methods are documented as follows by SFARI:
\\ The makeDoc file documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were necessary and are also available from GitHub.\
\ \\ SPARK Consortium. Electronic address: pfeliciano@simonsfoundation.org, SPARK Consortium.\ \ SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research.\ Neuron. 2018 Feb 7;97(3):488-493.\ PMID: 29420931; PMC: PMC7444276\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_sfari/SPARK.iWES_v3.2024_08.deepvariant.norm.vcf.gz\ dataVersion iWES v3 2024_08\ longLabel SNV Frequencies: SFARI SPARK - 140k WES\ parent varFreqs on\ priority 26\ shortLabel SFARI SPARK 140k WES\ tableBrowser off\ track sfariSparkExomes\ type vcfTabix\ visibility hide\ wgEncodeReg4MarkH3k27acAllSmallIntestine Small intestine (all biosamples) bigWig Avg. H3K27ac level of 10 small intestine experiments (all biosamples) 2 26 98 98 41 176 176 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/smallIntestineH3K27ac.bw\ color 98,98,41\ longLabel Avg. H3K27ac level of 10 small intestine experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 26\ shortLabel Small intestine (all biosamples)\ track wgEncodeReg4MarkH3k27acAllSmallIntestine\ type bigWig\ wgEncodeReg4MarkCtcfAllSpleen Spleen (all biosamples) bigWig Avg. CTCF level of 9 spleen experiments (all biosamples) 0 26 136 157 97 195 206 176 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/spleenCTCF.bw\ color 136,157,97\ longLabel Avg. CTCF level of 9 spleen experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 26\ shortLabel Spleen (all biosamples)\ track wgEncodeReg4MarkCtcfAllSpleen\ type bigWig\ Agilent_Human_Exon_V7_Regions SureSel. V7 T bigBed Agilent - SureSelect All Exon V7 Target Regions 0 26 255 36 36 255 145 145 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/S31285117_Regions.bb\ color 255,36,36\ longLabel Agilent - SureSelect All Exon V7 Target Regions\ parent exomeProbesets on\ shortLabel SureSel. V7 T\ track Agilent_Human_Exon_V7_Regions\ type bigBed\ chainSusScr11 Pig Chain chain susScr11 Pig (Feb. 2017 (Sscrofa11.1/susScr11)) Chained Alignments 3 27 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Pig (Feb. 2017 (Sscrofa11.1/susScr11)) Chained Alignments\ otherDb susScr11\ parent placentalChainNetViewchain off\ shortLabel Pig Chain\ subGroups view=chain species=s069 clade=c02\ track chainSusScr11\ type chain susScr11\ chainTarSyr2 Tarsier Chain chain tarSyr2 Tarsier (Sep. 2013 (Tarsius_syrichta-2.0.1/tarSyr2)) Chained Alignments 3 27 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Tarsier (Sep. 2013 (Tarsius_syrichta-2.0.1/tarSyr2)) Chained Alignments\ otherDb tarSyr2\ parent primateChainNetViewchain off\ shortLabel Tarsier Chain\ subGroups view=chain species=s037 clade=c02\ track chainTarSyr2\ type chain tarSyr2\ encTfChipPkENCFF542GMN A549 MYC narrowPeak Transcription Factor ChIP-seq Peaks of MYC in A549 from ENCODE 3 (ENCFF542GMN) 0 27 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of MYC in A549 from ENCODE 3 (ENCFF542GMN)\ parent encTfChipPk off\ shortLabel A549 MYC\ subGroups cellType=A549 factor=MYC\ track encTfChipPkENCFF542GMN\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_ctss_fwd AorticSmsToFgf2_02hrBr1+ bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_forward 0 27 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep1%20%28LK16%29.CNhs13344.12647-134H1.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12647-134H1 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_02hrBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_tpm_fwd AorticSmsToFgf2_02hrBr1+ bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_forward 1 27 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep1%20%28LK16%29.CNhs13344.12647-134H1.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12647-134H1 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_02hrBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1\ urlLabel FANTOM5 Details:\ wgEncodeReg4DnaseAllBloodVessel Blood vessel (all biosamples) bigWig Avg. DNase level of 25 blood vessel experiments (all biosamples) 0 27 255 37 41 255 146 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodVesselDNase.bw\ color 255,37,41\ longLabel Avg. DNase level of 25 blood vessel experiments (all biosamples)\ parent wgEncodeReg4Dnase off\ priority 27\ shortLabel Blood vessel (all biosamples)\ track wgEncodeReg4DnaseAllBloodVessel\ type bigWig\ wgEncodeReg4AtacAllBreast Breast (all biosamples) bigWig Avg. ATAC level of 4 breast experiments (all biosamples) 0 27 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/breastATAC.bw\ color 65,171,173\ longLabel Avg. ATAC level of 4 breast experiments (all biosamples)\ parent wgEncodeReg4Atac off\ priority 27\ shortLabel Breast (all biosamples)\ track wgEncodeReg4AtacAllBreast\ type bigWig\ gtexCovColonTransverse Colon Transverse bigWig Colon Transverse 0 27 238 197 145 246 226 200 0 0 0 expression 0 bigDataUrl /gbdb/hg38/gtex/cov/GTEX-1IDJC-1326-SM-CL53H.Colon_Transverse.RNAseq.bw\ color 238,197,145\ longLabel Colon Transverse\ parent gtexCov\ shortLabel Colon Transverse\ track gtexCovColonTransverse\ ENCFF796XMI_ENCFF679AWS_ENCFF703DMY_ENCFF782LSR ENCFF796XMI_ENCFF679AWS_ENCFF703DMY_ENCFF782LSR bigBed 9 + 5 Middle frontal area 46, female adult (90 or above years): (1) cCREs 4 27 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF796XMI_ENCFF679AWS_ENCFF703DMY_ENCFF782LSR.bb\ longLabel Middle frontal area 46, female adult (90 or above years): (1) cCREs\ mouseOver ID: ${name}\ The Simons Genome Diversity Project (SGDP), funded by the Simons Foundation,\ sequenced high-coverage genomes from 300 individuals (279 in this track) representing 142 diverse\ and often indigenous populations worldwide. Its goal was to capture the full range of human\ genetic diversity to better understand population history, migration, and adaptation. The\ sampling was designed to cover as much anthropological, linguistic and cultural diversity\ as possible, so it includes many deeply divergent human populations that are not well\ represented in other datasets.\
\ \\ This track shows allele frequencies only. The full phased genotype data with haplotype\ clustering display is available in the\ SGDP track under Phased Variants.\ Not all SGDP data is public, so this track contains only 279 genomes.\ The hg38 data was lifted from hg19.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For programmatic access, our REST API can be used; the\ track name is sgdpFreq.\ For bulk download, the VCF file can be obtained from\ our download server.\
\ \The original source VCFs are available from\ https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/.\
\ \\ High-coverage whole-genome sequencing of 300 individuals (279 publicly available) from 142\ diverse populations was performed on Illumina instruments using PCR-free library preparation at\ an average depth of 43x. Reads were aligned to the hs37d5 reference (GRCh37 with decoy\ sequences) using BWA-MEM 0.7.12. SNP genotyping was performed using GATK\ HaplotypeCaller with joint genotyping across all samples. (The Mallick 2016 release also\ includes an independent indel callset generated with FermiKit; indels are not carried in\ this track.)\
\\ The per-sample VCFs were merged with bcftools and lifted to hg38 with CrossMap. At UCSC,\ genotypes were stripped to produce a sites-only frequency VCF that keeps the AC, AF, and AN\ INFO fields. The deployed file contains 44,756,737 SNV records (601,775 of which represent\ multiallelic sites split into separate biallelic records). Indels from the source callset\ are not included.\ The conversion steps for all source files are documented in the makeDoc file of the track.\ Python scripts are also available from GitHub.\
\ \\ This project was funded by the Simons Foundation. Thanks to David Reich and Swapan\ Mallick for help with importing the data.\
\ \\ Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S,\ Tandon A et al.\ \ The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.\ Nature. 2016 Oct 13;538(7624):201-206.\ PMID: 27654912; PMC: PMC5161557\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/sgdpFreq/sgdp.freq.vcf.gz\ dataVersion 2016-12-07 (hg38 lift)\ longLabel SNV Frequencies: Simons Genome Diversity Project - 279 WGS, 142 populations\ parent varFreqs on\ priority 27\ shortLabel SGDP 279 WGS\ track sgdpFreq\ type vcfTabix\ visibility hide\ SKCM SKCM bigLolly 12 + Skin Cutaneous Melanoma 0 27 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/SKCM.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Skin Cutaneous Melanoma\ parent gdcCancer off\ priority 27\ shortLabel SKCM\ track SKCM\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4MarkH3k27acAllSpinalCord Spinal cord (all biosamples) bigWig H3K27ac level of 1 spinal cord experiment (all biosamples) 2 27 130 141 158 192 198 206 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/spinalCordH3K27ac.bw\ color 130,141,158\ longLabel H3K27ac level of 1 spinal cord experiment (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 27\ shortLabel Spinal cord (all biosamples)\ track wgEncodeReg4MarkH3k27acAllSpinalCord\ type bigWig\ wgEncodeReg4MarkCtcfAllStomach Stomach (all biosamples) bigWig Avg. CTCF level of 4 stomach experiments (all biosamples) 0 27 145 144 99 200 199 177 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/stomachCTCF.bw\ color 145,144,99\ longLabel Avg. CTCF level of 4 stomach experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 27\ shortLabel Stomach (all biosamples)\ track wgEncodeReg4MarkCtcfAllStomach\ type bigWig\ Twist_Comp_Exome_Target Twist Compr. T bigBed Twist - Comprehensive Exome Panel Target Regions 1 27 254 97 0 254 176 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/Twist_ComprehensiveExome_targets_hg38.bb\ color 254,97,0\ longLabel Twist - Comprehensive Exome Panel Target Regions\ parent exomeProbesets on\ shortLabel Twist Compr. T\ track Twist_Comp_Exome_Target\ type bigBed\ visibility dense\ cloneEndWI2 WI2 bed 12 WIBR-2 Fosmid library 0 27 0 0 0 127 127 127 0 0 0 map 1 colorByStrand 0,0,128 0,128,0\ longLabel WIBR-2 Fosmid library\ parent cloneEndSuper off\ priority 22\ shortLabel WI2\ subGroups source=wibr\ track cloneEndWI2\ type bed 12\ visibility hide\ netSusScr11 Pig Net netAlign susScr11 chainSusScr11 Pig (Feb. 2017 (Sscrofa11.1/susScr11)) Alignment Net 1 28 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Pig (Feb. 2017 (Sscrofa11.1/susScr11)) Alignment Net\ otherDb susScr11\ parent placentalChainNetViewnet on\ shortLabel Pig Net\ subGroups view=net species=s069 clade=c02\ track netSusScr11\ type netAlign susScr11 chainSusScr11\ netTarSyr2 Tarsier Net netAlign tarSyr2 chainTarSyr2 Tarsier (Sep. 2013 (Tarsius_syrichta-2.0.1/tarSyr2)) Alignment Net 1 28 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Tarsier (Sep. 2013 (Tarsius_syrichta-2.0.1/tarSyr2)) Alignment Net\ otherDb tarSyr2\ parent primateChainNetViewnet off\ shortLabel Tarsier Net\ subGroups view=net species=s037 clade=c02\ track netTarSyr2\ type netAlign tarSyr2 chainTarSyr2\ encTfChipPkENCFF418TUX A549 NFE2L2 narrowPeak Transcription Factor ChIP-seq Peaks of NFE2L2 in A549 from ENCODE 3 (ENCFF418TUX) 0 28 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of NFE2L2 in A549 from ENCODE 3 (ENCFF418TUX)\ parent encTfChipPk off\ shortLabel A549 NFE2L2\ subGroups cellType=A549 factor=NFE2L2\ track encTfChipPkENCFF418TUX\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_ctss_rev AorticSmsToFgf2_02hrBr1- bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_reverse 0 28 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep1%20%28LK16%29.CNhs13344.12647-134H1.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12647-134H1 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_02hrBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_tpm_rev AorticSmsToFgf2_02hrBr1- bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_reverse 1 28 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep1%20%28LK16%29.CNhs13344.12647-134H1.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep1 (LK16)_CNhs13344_12647-134H1_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12647-134H1 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_02hrBr1-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep1LK16_CNhs13344_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12647-134H1\ urlLabel FANTOM5 Details:\ ENCFF799QGM_ENCFF436OWL_ENCFF484YUA_ENCFF685MPU ENCFF799QGM_ENCFF436OWL_ENCFF484YUA_ENCFF685MPU bigBed 9 + 5 Middle frontal area 46 (mild cognitive impairment), female adult (90 or above years) with mild cognitive impairment: (1) cCREs 4 28 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF799QGM_ENCFF436OWL_ENCFF484YUA_ENCFF685MPU.bb\ longLabel Middle frontal area 46 (mild cognitive impairment), female adult (90 or above years) with mild cognitive impairment: (1) cCREs\ mouseOver ID: ${name}\ The National Precision Medicine (NPM) program\ in Singapore sequenced 9,770 whole genomes, mostly of Chinese, Indian and Malay ancestry.\ A minimum allele count cutoff of >5 was applied. CNV data is also available.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ VCF download can be requested on the Chorus Browser website, which requires an\ account and data access request.\
\ \\ Whole Genome Sequencing (WGS) data processing followed GATK4 best practices. The GATK4 germline\ variant analysis workflow written in WDL was adapted to Nextflow and deployed at the National\ Supercomputing Centre, Singapore (NSCC). WGS reads were aligned against GRCh38 with the BWA-MEM\ algorithm and used as input to GATK HaplotypeCaller to produce single sample gVCFs. The gVCF files\ were joint-called then loaded in Hail. Low-quality WGS libraries and low-quality variants were\ removed. QC-ed variants were functionally annotated with Ensembl Variant Effect Predictor (VEP)\ (version 95). For variants that affect protein-coding regions, the annotations also include\ information on potential changes to the cognate protein's 3D structure and drug binding ability.\
\\ Our data access request was approved by the NPM data access committee. It can be contacted at contact_npco@a-star.edu.sg.\ We downloaded the data from the NPM Chorus browser download section.\ The makeDoc file of the track documents how all source files of the varFreqs track were converted.\ For some tracks, python scripts were necessary and are also available from GitHub.\
\ \\ Thanks to the NPM Data Access Committee and Eleanor for granting our data request.\ By browsing the data, you agree to use the data only for academic, non-commercial\ research to improve human health (biology/disease). We request all data users\ agree to protect the confidentiality of the data subjects in any research papers or publications\ that they may prepare, by taking all reasonable care to limit the possibility\ of identification. In particular, the data users shall not use, or attempt\ to use, the data to deliberately compromise or otherwise infringe the\ confidentiality of information on data subjects and their right to privacy.\ If you use any of the data obtained from the CHORUS variant browser, we request\ that you cite the NPM flagship paper (Wong et al, 2023). All data users of the\ data must take note that the data provider and relevant SG10K_Health cohort\ owners bear no responsibility for the further analysis or interpretation of the data.\
\ \\ Wong E, Bertin N, Hebrard M, Tirado-Magallanes R, Bellis C, Lim WK, Chua CY, Tong PML, Chua R, Mak K\ et al.\ \ The Singapore National Precision Medicine Strategy.\ Nat Genet. 2023 Feb;55(2):178-186.\ PMID: 36658435\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_npm/SG10K_Health_r5.3.2.sites.vcf.bgz\ dataVersion r5.3.2\ longLabel SNV Frequencies: NPM Singapore - 9,770 WGS samples\ parent varFreqs on\ priority 28\ shortLabel Singapore NPM 9.7k WGS\ tableBrowser off\ track npm\ type vcfTabix\ visibility hide\ wgEncodeReg4MarkH3k4me3AllSmallIntestine Small intestine (all biosamples) bigWig Avg. H3K4me3 level of 11 small intestine experiments (all biosamples) 0 28 98 98 41 176 176 148 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/smallIntestineH3K4me3.bw\ color 98,98,41\ longLabel Avg. H3K4me3 level of 11 small intestine experiments (all biosamples)\ parent wgEncodeReg4MarkH3k4me3 off\ priority 28\ shortLabel Small intestine (all biosamples)\ track wgEncodeReg4MarkH3k4me3AllSmallIntestine\ type bigWig\ wgEncodeReg4MarkH3k27acAllSpleen Spleen (all biosamples) bigWig Avg. H3K27ac level of 11 spleen experiments (all biosamples) 2 28 136 157 97 195 206 176 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/spleenH3K27ac.bw\ color 136,157,97\ longLabel Avg. H3K27ac level of 11 spleen experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 28\ shortLabel Spleen (all biosamples)\ track wgEncodeReg4MarkH3k27acAllSpleen\ type bigWig\ STAD STAD bigLolly 12 + Stomach adenocarcinoma 0 28 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/STAD.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Stomach adenocarcinoma\ parent gdcCancer off\ priority 28\ shortLabel STAD\ track STAD\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4MarkCtcfAllTestis Testis (all biosamples) bigWig Avg. CTCF level of 2 testis experiments (all biosamples) 0 28 139 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/testisCTCF.bw\ color 139,140,140\ longLabel Avg. CTCF level of 2 testis experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 28\ shortLabel Testis (all biosamples)\ track wgEncodeReg4MarkCtcfAllTestis\ type bigWig\ Twist_Exome_Target Twist Core T bigBed Twist - Bioscience - Core Exome Panel Target Regions 1 28 254 97 0 254 176 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/Twist_Exome_Target_hg38.bb\ color 254,97,0\ longLabel Twist - Bioscience - Core Exome Panel Target Regions\ parent exomeProbesets off\ shortLabel Twist Core T\ track Twist_Exome_Target\ type bigBed\ visibility dense\ chainManPen1 Chinese pangolin Chain chain manPen1 Chinese pangolin (Aug 2014 (M_pentadactyla-1.1.1/manPen1)) Chained Alignments 3 29 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Chinese pangolin (Aug 2014 (M_pentadactyla-1.1.1/manPen1)) Chained Alignments\ otherDb manPen1\ parent placentalChainNetViewchain off\ shortLabel Chinese pangolin Chain\ subGroups view=chain species=s092 clade=c04\ track chainManPen1\ type chain manPen1\ chainMicMur2 Mouse lemur Chain chain micMur2 Mouse lemur (May 2015 (Mouse lemur/micMur2)) Chained Alignments 3 29 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Mouse lemur (May 2015 (Mouse lemur/micMur2)) Chained Alignments\ otherDb micMur2\ parent primateChainNetViewchain off\ shortLabel Mouse lemur Chain\ subGroups view=chain species=s043 clade=c03\ track chainMicMur2\ type chain micMur2\ encTfChipPkENCFF714KXI A549 NR3C1 1 narrowPeak Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF714KXI) 0 29 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF714KXI)\ parent encTfChipPk off\ shortLabel A549 NR3C1 1\ subGroups cellType=A549 factor=NR3C1\ track encTfChipPkENCFF714KXI\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_ctss_fwd AorticSmsToFgf2_02hrBr2+ bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_forward 0 29 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep2%20%28LK17%29.CNhs13363.12745-135I9.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12745-135I9 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_02hrBr2+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_tpm_fwd AorticSmsToFgf2_02hrBr2+ bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_forward 1 29 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep2%20%28LK17%29.CNhs13363.12745-135I9.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12745-135I9 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_02hrBr2+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9\ urlLabel FANTOM5 Details:\ ENCFF013AMD_ENCFF563YFA_ENCFF336MIJ_ENCFF302UYV ENCFF013AMD_ENCFF563YFA_ENCFF336MIJ_ENCFF302UYV bigBed 9 + 5 Middle frontal area 46 (Alzheimers disease), female adult (88 years) with Alzheimers disease: (1) cCREs 4 29 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF013AMD_ENCFF563YFA_ENCFF336MIJ_ENCFF302UYV.bb\ longLabel Middle frontal area 46 (Alzheimers disease), female adult (88 years) with Alzheimers disease: (1) cCREs\ mouseOver ID: ${name}\ This track shows small-variant (single-nucleotide variant and short-indel)\ allele frequencies from 101 samples released as part of the\ GWAS\ SVatalog tool (Chirmade et al. 2026). The same 101-sample cohort\ underlies the structural-variant sibling track\ SVatalog 101 SVs in the Long-read\ SV collection; this track provides the companion small-variant allele\ frequencies that SVatalog uses to compute linkage disequilibrium between\ SNPs and SVs.\
\\ The callset contains about 8.8 million sites across the autosomes\ and chromosome X. Each site reports the alternate allele frequency in the\ 101 samples, the gnomAD v3.1 non-Finnish European allele frequency (when\ annotated in the source release), and a dbSNP rsID when one was available.\
\ \\ The track uses the standard VCF display. Variants appear as colored marks\ along the genome; clicking an item opens the detail page with per-site\ INFO fields: AF, AC, AN, the gnomAD v3.1 NFE allele frequency\ (GNOMAD_NFE_AF) and the dbSNP rsID (RSID).\
\\ Note on AC/AN: the source allele-frequency release only ships AF. For this\ track we synthesize AC and AN by assuming the full 2x101 = 202-allele\ denominator (AN=202, AC=round(AF x 202)), so the values are approximate\ at sites where some samples had missing genotypes.\
\ \\ Small variants were called from 10X Genomics linked-read (paired-end\ short-read) whole-genome sequencing of the 101 SVatalog samples with\ GATK\ HaplotypeCaller v4.0.0.0 using default parameters. Calls were phased\ across the cohort with\ SHAPEIT\ v4.2.0, and per-site alternate allele frequencies were computed on\ the resulting joint callset. Structural variants, released as a separate\ lrSv subtrack, were called from long-read data and merged with these\ SNPs for the LD analyses reported by GWAS SVatalog.\
\\ For display here, the per-chromosome allele-frequency text files\ (chr{1..22,X}_allele_freq.txt) were converted to a single\ sites-only VCF with approximate AC/AN fields and bgzipped / tabix\ indexed. The step-by-step build commands are recorded in the UCSC\ makeDoc\ \ doc/hg38/varFreqs.txt; the converter script lives in\ \ makeDb/scripts/varFreqs.\
\ \\ The VCF file for this track is available from\ our\ download server as svatalog.vcf.gz (with .tbi index).\ Regions can be extracted with tabix:\ tabix http://hgdownload.soe.ucsc.edu/gbdb/hg38/varFreqs/svatalog/svatalog.vcf.gz chr21:1-100000000.\
\\ The original per-chromosome allele-frequency tables and the accompanying\ LD statistics used by the SVatalog tool are available from the\ companion Zenodo deposit:\ zenodo.org/records/13367574.\ The SVatalog web tool itself is at\ svatalog.research.sickkids.ca.\
\ \\ Thanks to Chirmade, Strug and colleagues at The Hospital for Sick\ Children and the University of Toronto for releasing this annotated\ SNP frequency callset alongside the GWAS SVatalog tool.\
\ \\ Chirmade S, Wang Z, Mastromatteo S, Sanders E, Thiruvahindrapuram B, Nalpathamkalam T, Pellecchia G,\ Lin F, Keenan K, Patel RV et al.\ \ GWAS SVatalog: a visualization tool to aid fine-mapping of GWAS loci with structural variations.\ Heredity (Edinb). 2026 Mar;135(3):199-210.\ PMID: 41203876; PMC: PMC13031531\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/svatalog/svatalog.vcf.gz\ dataVersion Chirmade 2025 release\ longLabel SNV Frequencies: GWAS SVatalog - 101 samples, 10X Genomics linked-read SNPs\ parent varFreqs on\ priority 29\ shortLabel SVatalog 101 WGS\ track svatalogSnv\ type vcfTabix\ visibility hide\ TGCT TGCT bigLolly 12 + Testicular Germ Cell Tumors 0 29 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/TGCT.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Testicular Germ Cell Tumors\ parent gdcCancer off\ priority 29\ shortLabel TGCT\ track TGCT\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4MarkCtcfAllThyroid Thyroid (all biosamples) bigWig Avg. CTCF level of 4 thyroid experiments (all biosamples) 0 29 27 119 58 141 187 156 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/thyroidCTCF.bw\ color 27,119,58\ longLabel Avg. CTCF level of 4 thyroid experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 29\ shortLabel Thyroid (all biosamples)\ track wgEncodeReg4MarkCtcfAllThyroid\ type bigWig\ Twist_Exome_Target2 Twist Exome 2.0 bigBed Twist - Exome 2.0 Panel Target Regions 1 29 254 97 0 254 176 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/TwistExome21.bb\ color 254,97,0\ longLabel Twist - Exome 2.0 Panel Target Regions\ parent exomeProbesets on\ shortLabel Twist Exome 2.0\ track Twist_Exome_Target2\ type bigBed\ visibility dense\ netManPen1 Chinese pangolin Net netAlign manPen1 chainManPen1 Chinese pangolin (Aug 2014 (M_pentadactyla-1.1.1/manPen1)) Alignment Net 1 30 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Chinese pangolin (Aug 2014 (M_pentadactyla-1.1.1/manPen1)) Alignment Net\ otherDb manPen1\ parent placentalChainNetViewnet off\ shortLabel Chinese pangolin Net\ subGroups view=net species=s092 clade=c04\ track netManPen1\ type netAlign manPen1 chainManPen1\ netMicMur2 Mouse lemur Net netAlign micMur2 chainMicMur2 Mouse lemur (May 2015 (Mouse lemur/micMur2)) Alignment Net 1 30 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Mouse lemur (May 2015 (Mouse lemur/micMur2)) Alignment Net\ otherDb micMur2\ parent primateChainNetViewnet off\ shortLabel Mouse lemur Net\ subGroups view=net species=s043 clade=c03\ track netMicMur2\ type netAlign micMur2 chainMicMur2\ encTfChipPkENCFF514IGC A549 NR3C1 2 narrowPeak Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF514IGC) 0 30 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF514IGC)\ parent encTfChipPk off\ shortLabel A549 NR3C1 2\ subGroups cellType=A549 factor=NR3C1\ track encTfChipPkENCFF514IGC\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_ctss_rev AorticSmsToFgf2_02hrBr2- bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_reverse 0 30 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep2%20%28LK17%29.CNhs13363.12745-135I9.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12745-135I9 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_02hrBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_tpm_rev AorticSmsToFgf2_02hrBr2- bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_reverse 1 30 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep2%20%28LK17%29.CNhs13363.12745-135I9.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep2 (LK17)_CNhs13363_12745-135I9_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12745-135I9 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_02hrBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep2LK17_CNhs13363_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12745-135I9\ urlLabel FANTOM5 Details:\ ENCFF987RXP_ENCFF419XND_ENCFF224LYA_ENCFF081IRZ ENCFF987RXP_ENCFF419XND_ENCFF224LYA_ENCFF081IRZ bigBed 9 + 5 Middle frontal area 46 (cognitive impairment), female adult (81 years) with Cognitive impairment: (1) cCREs 4 30 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF987RXP_ENCFF419XND_ENCFF224LYA_ENCFF081IRZ.bb\ longLabel Middle frontal area 46 (cognitive impairment), female adult (81 years) with Cognitive impairment: (1) cCREs\ mouseOver ID: ${name}\ SweGen provides\ whole-genome sequencing variant frequencies for 1,000 Swedish individuals.\ The 1,000 individuals represent a cross-section of the Swedish population and no disease\ information was used for the selection. The frequency data may therefore include genetic variants\ that are associated with, or causative of, disease. SweGen also provides SV calls, TEs, MELT\ results for TEs, HLAs and a FASTA file with new sequence not in hg38. There is\ also a version for the T2T CHM13 assembly. The full dataset can be browsed at\ the\ SweGen Browser.\
\\ The mobile element insertions called by MELT on the same 1,000 SweGen\ samples are loaded as a separate track,\ SweGen 1000 MEIs, in the\ Mobile Element Insertions collection.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ VCF files can be requested at\ SweGen via a form. The request\ needs manual approval, which is usually quick. If there is no reply, email SweGen directly.\
\ \\ Fragment size 350bp on a Covaris E220. Paired-end sequencing with 150bp read length was performed\ on Illumina HiSeq X (HiSeq Control Software 3.3.39/RTA 2.7.1) with v2.5 sequencing chemistry.\ Raw whole-genome reads were aligned to the GRCh37 reference using BWA-MEM v0.7.12, then sorted and\ indexed with samtools v0.1.19 and assessed with qualimap v2.2.20; per-sample alignments from\ multiple lanes and flow cells were merged using Picard MergeSamFiles v1.120. Processing followed\ GATK best practices with GATK v3.3, including indel realignment (RealignerTargetCreator,\ IndelRealigner), duplicate marking (Picard MarkDuplicates v1.120), and base quality score\ recalibration (BaseRecalibrator), producing one finalized BAM per sample. Per-sample gVCFs were\ generated with GATK HaplotypeCaller v3.3 using reference files from the GATK v2.8 resource bundle,\ with all steps coordinated via Piper v1.4.0. Joint genotyping of 1,000 samples was performed by\ merging gVCFs in five batches of 200 using GATK CombineGVCFs, followed by cohort genotyping with\ GATK GenotypeGVCFs and variant quality score recalibration for SNVs and indels using\ VariantRecalibrator and ApplyRecalibration.\
\\ At UCSC, the hg38 VCF was downloaded from\ SweFreq and loaded as-is.\ The file that we use is swegen_frequencies_fixploidy_GRCh38_20190204.vcf.gz.\ The conversion steps for all source files of the varFreqs track are documented in the track's makeDoc file.\ For some tracks, python scripts were needed; these are also available from GitHub.\
\ \\ The SweGen allele frequency data was generated by Science for Life Laboratory. \ Any redistributed data derived from the SweGen data set must follow the SweGen terms and conditions.\ The data may not be used to attempt to identify any individual in this or other studies.\ Thanks to the SweGen patients and SciLifeLab for making the data available.\
\ \\ Ameur A, Dahlberg J, Olason P, Vezzi F, Karlsson R, Martin M, Viklund J, Kähäri AK,\ Lundin P, Che H et al.\ \ SweGen: a whole-genome data resource of genetic variability in a cross-section of the Swedish\ population.\ Eur J Hum Genet. 2017 Nov;25(11):1253-1260.\ PMID: 28832569; PMC: PMC5765326\
\ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_swefreq/swegen_frequencies_fixploidy_GRCh38_20190204.vcf.gz\ dataVersion 20251201\ longLabel SNV Frequencies: Sweden SweGen - 1k WGS\ parent varFreqs on\ priority 30\ shortLabel Sweden SweGen 1k WGS\ tableBrowser off\ track swefreq\ type vcfTabix\ visibility hide\ wgEncodeReg4MarkH3k27acAllTestis Testis (all biosamples) bigWig Avg. H3K27ac level of 2 testis experiments (all biosamples) 2 30 139 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/testisH3K27ac.bw\ color 139,140,140\ longLabel Avg. H3K27ac level of 2 testis experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 30\ shortLabel Testis (all biosamples)\ track wgEncodeReg4MarkH3k27acAllTestis\ type bigWig\ THCA THCA bigLolly 12 + Thyroid carcinoma 0 30 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/THCA.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Thyroid carcinoma\ parent gdcCancer off\ priority 30\ shortLabel THCA\ track THCA\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ Twist_Exome_RefSeq_Targets Twist RefSeq T bigBed Twist - RefSeq Exome Panel Target Regions 1 30 254 97 0 254 176 127 0 0 0 map 1 bigDataUrl /gbdb/hg38/exomeProbesets/Twist_Exome_RefSeq_targets_hg38.bb\ color 254,97,0\ longLabel Twist - RefSeq Exome Panel Target Regions\ parent exomeProbesets off\ shortLabel Twist RefSeq T\ track Twist_Exome_RefSeq_Targets\ type bigBed\ visibility dense\ wgEncodeReg4MarkCtcfAllVagina Vagina (all biosamples) bigWig Avg. CTCF level of 2 vagina experiments (all biosamples) 0 30 255 101 174 255 178 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/vaginaCTCF.bw\ color 255,101,174\ longLabel Avg. CTCF level of 2 vagina experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 30\ shortLabel Vagina (all biosamples)\ track wgEncodeReg4MarkCtcfAllVagina\ type bigWig\ chainEquCab3 Horse Chain chain equCab3 Horse (Jan. 2018 (EquCab3.0/equCab3)) Chained Alignments 3 31 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Horse (Jan. 2018 (EquCab3.0/equCab3)) Chained Alignments\ otherDb equCab3\ parent placentalChainNetViewchain off\ shortLabel Horse Chain\ subGroups view=chain species=s096a clade=c05\ track chainEquCab3\ type chain equCab3\ chainOtoGar3 Bushbaby Chain chain otoGar3 Bushbaby (Mar. 2011 (Broad/otoGar3)) Chained Alignments 3 31 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Bushbaby (Mar. 2011 (Broad/otoGar3)) Chained Alignments\ otherDb otoGar3\ parent primateChainNetViewchain off\ shortLabel Bushbaby Chain\ subGroups view=chain species=s045 clade=c03\ track chainOtoGar3\ type chain otoGar3\ encTfChipPkENCFF963CGV A549 NR3C1 3 narrowPeak Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF963CGV) 0 31 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF963CGV)\ parent encTfChipPk off\ shortLabel A549 NR3C1 3\ subGroups cellType=A549 factor=NR3C1\ track encTfChipPkENCFF963CGV\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_ctss_fwd AorticSmsToFgf2_02hrBr3+ bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_forward 0 31 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep3%20%28LK18%29.CNhs13572.12843-137B8.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12843-137B8 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_02hrBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_tpm_fwd AorticSmsToFgf2_02hrBr3+ bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_forward 1 31 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep3%20%28LK18%29.CNhs13572.12843-137B8.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12843-137B8 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_02hrBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8\ urlLabel FANTOM5 Details:\ wgEncodeReg4MarkCtcfAllBlood Blood (all biosamples) bigWig Avg. CTCF level of 25 blood experiments (all biosamples) 0 31 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodCTCF.bw\ color 254,75,173\ longLabel Avg. CTCF level of 25 blood experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 31\ shortLabel Blood (all biosamples)\ track wgEncodeReg4MarkCtcfAllBlood\ type bigWig\ ENCFF521HEY_ENCFF862YHY_ENCFF194KAZ_ENCFF263VJQ ENCFF521HEY_ENCFF862YHY_ENCFF194KAZ_ENCFF263VJQ bigBed 9 + 5 Middle frontal area 46 (Alzheimers disease), female adult (85 years) with Alzheimers disease: (1) cCREs 4 31 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF521HEY_ENCFF862YHY_ENCFF194KAZ_ENCFF263VJQ.bb\ longLabel Middle frontal area 46 (Alzheimers disease), female adult (85 years) with Alzheimers disease: (1) cCREs\ mouseOver ID: ${name}\ This track shows allele frequencies for 672,843 variants from the\ Taiwan\ Precision Medicine Initiative (TPMI), a large cohort of people of\ Han Chinese ancestry recruited in Taiwan. The frequencies come from the\ publicly released annotation of the Axiom TPM1 SNP array, the\ population-optimized chip that TPMI used to genotype 165,596 of its\ participants. Variants are positioned on hg38 (GRCh38). About 80% of\ the sites are biallelic SNVs; the remainder are short insertions or\ deletions and a small number of multi-nucleotide variants.\
\ \\ TPMI is one of the largest non-European cohorts in genetic research,\ with 565,390 enrolled participants as of the v37 data freeze. Han\ Chinese people are nearly 20% of the world's population but are\ under-represented in genetic studies. A cohort of this size is useful\ for population-specific allele frequency reference, GWAS replication,\ and clinical variant interpretation in East Asian populations.\
\ \\ The track uses the standard UCSC VCF display. Hovering a variant shows\ the cohort allele frequency (AF), the derived allele count\ (AC), the assumed total allele number (AN), the TPMI\ NGS concordance score from the chip annotation, and the Affymetrix\ probe set ID.\
\ \\ TPMI participants were recruited from 16 partner medical centres (33\ affiliated hospitals) across Taiwan, who together serve about 40% of the\ Taiwanese population. Each participant donated a blood sample and\ consented to access of their electronic medical records. Genomic DNA\ was extracted with the QIAsymphony DSP DNA Mini Kit and genotyped on\ two custom Axiom arrays (TPMv1 and TPMv2; Thermo Fisher Scientific)\ designed to optimally tag Han Chinese variation. Genotype calling was\ done with Applied Biosystems Array Power Tools using the Best Practices\ Workflow at the National Center for Genome Medicine, Academia Sinica.\ After QC, the TPMv1 array had been used on 165,596 participants and\ TPMv2 on 321,360 (486,956 with both genotype and EMR). The cohort has\ broad coverage of Han Chinese subgroups as well as Indigenous Taiwanese\ populations. See the TPMI Nature paper (in References) for sample\ recruitment, calling, imputation and quality control details.\
\\ The source data for this track is the Axiom TPM1 chip annotation file\ TPM1_Array_Annotation.csv distributed by Thermo Fisher\ Scientific (create date 2022-06-01), which embeds the TPMI cohort allele\ frequency in a column named Allele Frequency alongside the\ probe-design metadata. The chip annotation declares hg38 coordinates,\ so no liftover was needed. We converted the CSV to VCF with the script\ tpmiToVcf.py:\ rows on alt or random contigs were dropped, rows flagged as TPMI\ blacklist or with no reported allele frequency were dropped, and indels\ encoded with - for the empty allele were rewritten in\ VCF-compatible form by prepending an anchor base read from the hg38\ reference with twoBitToFa. The resulting VCF was sorted and\ indexed with bcftools sort and tabix. The full\ recipe is in the\ makeDoc\ file.\
\\ The source publishes only allele frequencies, not allele counts. To\ make the track usable in count-based aggregate views, we derived\ AC = round(AF * AN) with AN = 100,000. This AN value\ was chosen because every reported AF in the file is an exact integer\ multiple of 1/100,000, so the source data was rounded to that\ precision. The TPMv1 chip was used on 165,596 participants (~330,000\ chromosomes for autosomes), so the true AN may be roughly three times\ larger; the AC values published here are therefore proportional to the\ true counts but not equal to them. The assumption is documented in the\ VCF header.\
\ \\ Of 752,921 rows in the source CSV, 672,843 were emitted to the VCF.\ The skipped rows are: 80,034 rows with no reported allele frequency\ (the chip carries probe annotations for some sites that the TPMI cohort\ did not type or quality-filter, including the entire chrY content of\ the chip); 36 rows on alt or random contigs; 8 rows with no defined\ reference allele in the source. About 61,000 rows are also flagged as\ TPMI blacklist; none of those have a published allele frequency, so\ they are filtered out by the no-AF rule.\
\\ The TPM2 chip annotation (~755,000 SNPs) is not represented in this\ track because its public annotation does not embed a TPMI cohort allele\ frequency column. It only carries the 1000 Genomes / HapMap CEU/CHB/JPT/YRI\ frequencies that ship with all Affymetrix Axiom chips, which are already\ available through dbSNP. About 234,255 SNPs are shared between TPM1 and\ TPM2, so the TPM1-only track still covers most of the cohort-typed\ content.\
\\ The TPMI authors note that allele frequencies on the TPMv1 chip are\ reliable for variants with MAF above about 0.1%; rarer sites are\ reported but should be interpreted cautiously because SNP arrays have\ higher genotyping error at low MAF.\
\ \\ Due to license restrictions, the data for this track cannot be downloaded from the UCSC\ Genome Browser. The Table Browser, Data Integrator, and download server are not available\ for this track.\
\\ The original Axiom TPM1 chip annotation CSV is distributed by Thermo Fisher Scientific;\ search their support site for "Axiom TPM1 Annotation" to download the matching version\ (we used the 2022-06-01 release).\
\ \\ Thanks to the TPMI participants and to the Academia Sinica and Thermo\ Fisher Scientific teams that designed and curated the Axiom TPMv1 SNP\ array and published the chip annotation file.\
\ \\ Yang HC, Kwok PY, Li LH, Liu YM, Jong YJ, Lee KY, Wang DW, Tsai MF, Yang JH, Chen CH et al.\ \ The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies.\ Nature. 2025 Dec;648(8092):117-127.\ PMID: 41092961; PMC: PMC12675286\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/_tpmi/tpmi.vcf.gz\ dataVersion Axiom TPM1 2022-06\ longLabel SNV Frequencies: Taiwan Precision Medicine Initiative - Axiom TPM1 chip, Han Chinese\ parent varFreqs on\ priority 31\ shortLabel Taiwan TPMI Axiom array\ tableBrowser off\ track tpmi\ type vcfTabix\ visibility hide\ THYM THYM bigLolly 12 + Thymoma 0 31 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/THYM.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Thymoma\ parent gdcCancer off\ priority 31\ shortLabel THYM\ track THYM\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ wgEncodeReg4MarkH3k27acAllThymus Thymus (all biosamples) bigWig Avg. H3K27ac level of 2 thymus experiments (all biosamples) 2 31 142 124 195 198 189 225 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/thymusH3K27ac.bw\ color 142,124,195\ longLabel Avg. H3K27ac level of 2 thymus experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 31\ shortLabel Thymus (all biosamples)\ track wgEncodeReg4MarkH3k27acAllThymus\ type bigWig\ netEquCab3 Horse Net netAlign equCab3 chainEquCab3 Horse (Jan. 2018 (EquCab3.0/equCab3)) Alignment Net 1 32 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Horse (Jan. 2018 (EquCab3.0/equCab3)) Alignment Net\ otherDb equCab3\ parent placentalChainNetViewnet off\ shortLabel Horse Net\ subGroups view=net species=s096a clade=c05\ track netEquCab3\ type netAlign equCab3 chainEquCab3\ netOtoGar3 Bushbaby Net netAlign otoGar3 chainOtoGar3 Bushbaby (Mar. 2011 (Broad/otoGar3)) Alignment Net 1 32 0 0 0 255 255 0 0 0 0 compGeno 0 longLabel Bushbaby (Mar. 2011 (Broad/otoGar3)) Alignment Net\ otherDb otoGar3\ parent primateChainNetViewnet on\ shortLabel Bushbaby Net\ subGroups view=net species=s045 clade=c03\ track netOtoGar3\ type netAlign otoGar3 chainOtoGar3\ encTfChipPkENCFF114SRD A549 NR3C1 4 narrowPeak Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF114SRD) 0 32 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF114SRD)\ parent encTfChipPk off\ shortLabel A549 NR3C1 4\ subGroups cellType=A549 factor=NR3C1\ track encTfChipPkENCFF114SRD\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_ctss_rev AorticSmsToFgf2_02hrBr3- bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_reverse 0 32 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep3%20%28LK18%29.CNhs13572.12843-137B8.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12843-137B8 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_02hrBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_tpm_rev AorticSmsToFgf2_02hrBr3- bigWig Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_reverse 1 32 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2002hr%2c%20biol_rep3%20%28LK18%29.CNhs13572.12843-137B8.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to FGF2, 02hr, biol_rep3 (LK18)_CNhs13572_12843-137B8_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12843-137B8 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_02hrBr3-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=reverse\ track AorticSmoothMuscleCellResponseToFGF202hrBiolRep3LK18_CNhs13572_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12843-137B8\ urlLabel FANTOM5 Details:\ wgEncodeReg4MarkCtcfAllBrain Brain (all biosamples) bigWig Avg. CTCF level of 69 brain experiments (all biosamples) 0 32 155 155 18 205 205 136 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/brainCTCF.bw\ color 155,155,18\ longLabel Avg. CTCF level of 69 brain experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 32\ shortLabel Brain (all biosamples)\ track wgEncodeReg4MarkCtcfAllBrain\ type bigWig\ ENCFF541ZVM_ENCFF889QTE_ENCFF480FCW_ENCFF796CNP ENCFF541ZVM_ENCFF889QTE_ENCFF480FCW_ENCFF796CNP bigBed 9 + 5 Middle frontal area 46 (Alzheimers disease), female adult (81 years) with Alzheimers disease: (1) cCREs 4 32 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF541ZVM_ENCFF889QTE_ENCFF480FCW_ENCFF796CNP.bb\ longLabel Middle frontal area 46 (Alzheimers disease), female adult (81 years) with Alzheimers disease: (1) cCREs\ mouseOver ID: ${name}\ This track shows allele frequencies and imputation quality scores for\ 13,743,085 variants observed in 361,194 UK Biobank participants of white\ British ancestry. The\ UK Biobank\ is a prospective study of around 500,000 adults aged 40-69 at recruitment\ in the UK, with linked genotype, imaging and health-record data. The\ allele counts shown here are taken from the Neale Lab's open release of\ imputed-v3 GWAS results, which the Lab made freely available as a\ companion to their large phenotype-wide GWAS of UK Biobank (Round 2 of\ the\ Neale Lab\ UK Biobank GWAS).\
\ \\ The Neale Lab pipeline restricts to white British ancestry to limit\ population-stratification confounding in the GWAS. As a consequence the\ frequencies in this track are not representative of the multi-ancestry UK\ Biobank cohort. They describe a single population subset. The\ gnomAD HGDP+1kG,\ ToMMo Japan,\ AllOfUs and other tracks in this\ collection provide complementary frequencies from other populations.\
\ \\ The track uses the standard UCSC VCF display. Hover over a variant to see\ the allele frequency, imputation INFO score, HWE p-value, hom-ref / het /\ hom-alt sample counts and the most-severe VEP consequence reported by\ the Neale Lab.\
\ \\ UK Biobank participants were genotyped on the UK Biobank Axiom and UK\ BiLEVE Axiom arrays. The Wellcome Trust Centre for Human Genetics imputed\ the array data against a combined reference panel of the Haplotype\ Reference Consortium, UK10K and 1000 Genomes Phase 3. This produced\ approximately 90 million imputed SNPs. The Neale Lab Round 2 (imputed-v3)\ analysis started from the 487,409 individuals with phased and imputed\ genotype data, filtered to 361,194 unrelated samples of white British\ ancestry, and retained variants with imputation INFO score above 0.8,\ minor allele frequency above 0.001 (or 1e-6 for coding variants) and\ HWE p-value above 1e-10. The final set has 13.7 million SNPs and short\ indels on chromosomes 1-22 and X. Variant consequences are from Ensembl VEP. See\ the Neale Lab\ data\ processing blog post and the\ UK_Biobank_GWAS\ GitHub repository for the full pipeline.\
\ \\ The variant manifest\ variants.tsv.bgz was downloaded from the Neale Lab\ UK Biobank\ GWAS results page. The Neale Lab release uses GRCh37 coordinates and\ provides chromosome, position, reference and alternate alleles, dbSNP\ rsID, VEP consequence, imputation INFO score, allele count and\ frequency, Hardy-Weinberg p-value and per-genotype sample counts. We\ converted the TSV to a sites-only VCF using a custom Python script and\ lifted the coordinates to GRCh38 with CrossMap and the UCSC\ hg19ToHg38.over.chain. 39,659 rows with allele count zero (variants\ present only in the imputation panel) were dropped, 6,889 failed\ liftOver and 1,834 mapped to alt/random/fix contigs, leaving 13,743,085\ variants in the final file. AN was set to twice the\ n_called field, per the Neale Lab convention.\ The full pipeline is documented in the\ makeDoc\ file of the track, and the conversion script is available from\ our\ GitHub repository.\
\ \\ The variant frequencies can be explored with the\ Table Browser or the\ Data Integrator, and exported to\ spreadsheet or tab-separated tables. From scripts, data can be accessed\ via our REST API\ with track=ukbb.\
\\ The VCF file is also available from\ our\ download server as ukbb.vcf.gz. Individual regions can be\ extracted with tabix, for example\ tabix http://hgdownload.soe.ucsc.edu/gbdb/hg38/varFreqs/ukbb/ukbb.vcf.gz chr21:1-100000000.\ The original Neale Lab manifest variants.tsv.bgz is linked from\ the\ Neale Lab UK\ Biobank GWAS results page and is distributed under UK Biobank's\ data-access conditions.\
\ \\ Thanks to the UK Biobank participants and to Benjamin Neale, Liam\ Abbott, Raymond Walters, Duncan Palmer and the rest of the Neale Lab for\ making the Round 2 imputed-v3 GWAS results, including the variant\ manifest used here, publicly available.\
\ \\ Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, Motyer A, Vukcevic D, Delaneau O,\ O'Connell J et al.\ \ The UK Biobank resource with deep phenotyping and genomic data.\ Nature. 2018 Oct;562(7726):203-209.\ PMID: 30305743; PMC: PMC6786975\
\ \ varRep 1 bigDataUrl /gbdb/hg38/varFreqs/ukbb/ukbb.vcf.gz\ dataVersion Neale Lab R2 08-2018\ longLabel SNV Frequencies: UK Biobank Genotypes - 361k White British, Neale Lab Round 2 imputed\ parent varFreqs on\ priority 32\ shortLabel UK Biobank 361k imputed\ track ukbb\ type vcfTabix\ visibility hide\ chainDasNov3 Armadillo Chain chain dasNov3 Armadillo (Dec. 2011 (Baylor/dasNov3)) Chained Alignments 3 33 0 0 0 255 255 0 1 0 0 compGeno 1 longLabel Armadillo (Dec. 2011 (Baylor/dasNov3)) Chained Alignments\ otherDb dasNov3\ parent placentalChainNetViewchain off\ shortLabel Armadillo Chain\ subGroups view=chain species=s100 clade=c06\ track chainDasNov3\ type chain dasNov3\ encTfChipPkENCFF463DJO A549 NR3C1 5 narrowPeak Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF463DJO) 0 33 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of NR3C1 in A549 from ENCODE 3 (ENCFF463DJO)\ parent encTfChipPk off\ shortLabel A549 NR3C1 5\ subGroups cellType=A549 factor=NR3C1\ track encTfChipPkENCFF463DJO\ AorticSmoothMuscleCellResponseToFGF203hrBiolRep1LK19_CNhs13345_ctss_fwd AorticSmsToFgf2_03hrBr1+ bigWig Aortic smooth muscle cell response to FGF2, 03hr, biol_rep1 (LK19)_CNhs13345_12648-134H2_forward 0 33 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12648-134H2 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2003hr%2c%20biol_rep1%20%28LK19%29.CNhs13345.12648-134H2.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 03hr, biol_rep1 (LK19)_CNhs13345_12648-134H2_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12648-134H2 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_03hrBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF203hrBiolRep1LK19_CNhs13345_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12648-134H2\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF203hrBiolRep1LK19_CNhs13345_tpm_fwd AorticSmsToFgf2_03hrBr1+ bigWig Aortic smooth muscle cell response to FGF2, 03hr, biol_rep1 (LK19)_CNhs13345_12648-134H2_forward 1 33 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12648-134H2 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2003hr%2c%20biol_rep1%20%28LK19%29.CNhs13345.12648-134H2.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 03hr, biol_rep1 (LK19)_CNhs13345_12648-134H2_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12648-134H2 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_03hrBr1+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF203hrBiolRep1LK19_CNhs13345_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12648-134H2\ urlLabel FANTOM5 Details:\ wgEncodeReg4AtacAllBoneMarrow Bone marrow (all biosamples) bigWig ATAC level of 1 bone marrow experiment (all biosamples) 0 33 184 120 120 219 187 187 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/boneMarrowATAC.bw\ color 184,120,120\ longLabel ATAC level of 1 bone marrow experiment (all biosamples)\ parent wgEncodeReg4Atac off\ priority 33\ shortLabel Bone marrow (all biosamples)\ track wgEncodeReg4AtacAllBoneMarrow\ type bigWig\ wgEncodeRegDnaseUwBonemarrowmscPeak bonemarrow_MSC Pk narrowPeak bone_marrow_MSC bone marrow fibroblastoid DNaseI Peaks from ENCODE 1 33 228 255 85 241 255 170 1 0 0 regulation 1 color 228,255,85\ longLabel bone_marrow_MSC bone marrow fibroblastoid DNaseI Peaks from ENCODE\ parent wgEncodeRegDnasePeak off\ shortLabel bonemarrow_MSC Pk\ subGroups view=a_Peaks cellType=bone_marrow_MSC treatment=n_a tissue=bone_marrow cancer=normal\ track wgEncodeRegDnaseUwBonemarrowmscPeak\ wgEncodeRegDnaseUwBonemarrowmscWig bonemarrow_MSC Sg bigWig 0 3047.47 bone_marrow_MSC bone marrow fibroblastoid DNaseI Signal from ENCODE 0 33 228 255 85 241 255 170 0 0 0 regulation 1 color 228,255,85\ longLabel bone_marrow_MSC bone marrow fibroblastoid DNaseI Signal from ENCODE\ parent wgEncodeRegDnaseWig off\ priority 1.26303\ shortLabel bonemarrow_MSC Sg\ subGroups cellType=bone_marrow_MSC treatment=n_a tissue=bone_marrow cancer=normal\ table wgEncodeRegDnaseUwBonemarrowmscSignal\ track wgEncodeRegDnaseUwBonemarrowmscWig\ type bigWig 0 3047.47\ wgEncodeReg4MarkCtcfAllBreast Breast (all biosamples) bigWig Avg. CTCF level of 9 breast experiments (all biosamples) 0 33 65 171 173 160 213 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/breastCTCF.bw\ color 65,171,173\ longLabel Avg. CTCF level of 9 breast experiments (all biosamples)\ parent wgEncodeReg4MarkCtcf off\ priority 33\ shortLabel Breast (all biosamples)\ track wgEncodeReg4MarkCtcfAllBreast\ type bigWig\ ENCFF753DPM_ENCFF353SJI_ENCFF649LLS_ENCFF554FTX ENCFF753DPM_ENCFF353SJI_ENCFF649LLS_ENCFF554FTX bigBed 9 + 5 Middle frontal area 46, female adult (90 or above years): (1) cCREs 4 33 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF753DPM_ENCFF353SJI_ENCFF649LLS_ENCFF554FTX.bb\ longLabel Middle frontal area 46, female adult (90 or above years): (1) cCREs\ mouseOver ID: ${name}\ The GENCODE Genes track (version 49, Sept 2025) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 49 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 49 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 49 corresponds to Ensembl 115.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V49 (Ensembl 115)\ maxTransEnabled on\ priority 34.156\ shortLabel All GENCODE V49\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV49\ type genePred\ visibility pack\ wgEncodeGencodeVersion 49\ wgEncodeGencodeV49ViewGenes Genes genePred All GENCODE annotations from V49 (Ensembl 115) 3 34.156 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV49 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV49\ longLabel All GENCODE annotations from V49 (Ensembl 115)\ parent wgEncodeGencodeV49\ shortLabel Genes\ track wgEncodeGencodeV49ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV49ViewPolya PolyA genePred All GENCODE annotations from V49 (Ensembl 115) 0 34.156 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V49 (Ensembl 115)\ parent wgEncodeGencodeV49\ shortLabel PolyA\ track wgEncodeGencodeV49ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV48 All GENCODE V48 genePred All GENCODE annotations from V48 (Ensembl 114) 0 34.157 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 48, May 2025) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 48 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 48 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 48 corresponds to Ensembl 114.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V48 (Ensembl 114)\ maxTransEnabled on\ priority 34.157\ shortLabel All GENCODE V48\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV48\ type genePred\ visibility hide\ wgEncodeGencodeVersion 48\ wgEncodeGencodeV48ViewGenes Genes genePred All GENCODE annotations from V48 (Ensembl 114) 3 34.157 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV48 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV48\ longLabel All GENCODE annotations from V48 (Ensembl 114)\ parent wgEncodeGencodeV48\ shortLabel Genes\ track wgEncodeGencodeV48ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV48ViewPolya PolyA genePred All GENCODE annotations from V48 (Ensembl 114) 0 34.157 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V48 (Ensembl 114)\ parent wgEncodeGencodeV48\ shortLabel PolyA\ track wgEncodeGencodeV48ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV47 All GENCODE V47 genePred All GENCODE annotations from V47 (Ensembl 113) 0 34.158 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 47, Oct 2024) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 47 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 47 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 47 corresponds to Ensembl 113.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V47 (Ensembl 113)\ maxTransEnabled on\ priority 34.158\ shortLabel All GENCODE V47\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV47\ type genePred\ visibility hide\ wgEncodeGencodeVersion 47\ wgEncodeGencodeV47ViewGenes Genes genePred All GENCODE annotations from V47 (Ensembl 113) 3 34.158 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,annotation_in_progress,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,annotation_in_progress,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV47 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV47\ longLabel All GENCODE annotations from V47 (Ensembl 113)\ parent wgEncodeGencodeV47\ shortLabel Genes\ track wgEncodeGencodeV47ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV47ViewPolya PolyA genePred All GENCODE annotations from V47 (Ensembl 113) 0 34.158 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V47 (Ensembl 113)\ parent wgEncodeGencodeV47\ shortLabel PolyA\ track wgEncodeGencodeV47ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV46 All GENCODE V46 genePred All GENCODE annotations from V46 (Ensembl 112) 0 34.159 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 46, May 2024) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 46 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 46 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 46 corresponds to Ensembl 112.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V46 (Ensembl 112)\ maxTransEnabled on\ priority 34.159\ shortLabel All GENCODE V46\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV46\ type genePred\ visibility hide\ wgEncodeGencodeVersion 46\ wgEncodeGencodeV46ViewGenes Genes genePred All GENCODE annotations from V46 (Ensembl 112) 3 34.159 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,annotation_in_progress,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,annotation_in_progress,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,confirm_experimentally,dotter_confirmed,downstream_ATG,Ensembl_canonical,EnsEMBL_merge_exception,exp_conf,fragmented_locus,fragmented_mixed_strand_locus,GENCODE_Primary,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,polymorphic_pseudogene_no_stop,precursor_RNA,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,Selenoprotein,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV46 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV46\ longLabel All GENCODE annotations from V46 (Ensembl 112)\ parent wgEncodeGencodeV46\ shortLabel Genes\ track wgEncodeGencodeV46ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV46ViewPolya PolyA genePred All GENCODE annotations from V46 (Ensembl 112) 0 34.159 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V46 (Ensembl 112)\ parent wgEncodeGencodeV46\ shortLabel PolyA\ track wgEncodeGencodeV46ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV45 All GENCODE V45 genePred All GENCODE annotations from V45 (Ensembl 111) 0 34.16 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 45, Jan 2024) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 45 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 45 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 45 corresponds to Ensembl 111.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V45 (Ensembl 111)\ maxTransEnabled on\ priority 34.160\ shortLabel All GENCODE V45\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV45\ type genePred\ visibility hide\ wgEncodeGencodeVersion 45\ wgEncodeGencodeV45ViewGenes Genes genePred All GENCODE annotations from V45 (Ensembl 111) 3 34.16 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV45 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV45\ longLabel All GENCODE annotations from V45 (Ensembl 111)\ parent wgEncodeGencodeV45\ shortLabel Genes\ track wgEncodeGencodeV45ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV45ViewPolya PolyA genePred All GENCODE annotations from V45 (Ensembl 111) 0 34.16 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V45 (Ensembl 111)\ parent wgEncodeGencodeV45\ shortLabel PolyA\ track wgEncodeGencodeV45ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV44 All GENCODE V44 genePred All GENCODE annotations from V44 (Ensembl 110) 0 34.161 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 44, July 2023) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 44 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 44 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 44 corresponds to Ensembl 110.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V44 (Ensembl 110)\ maxTransEnabled on\ priority 34.161\ shortLabel All GENCODE V44\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV44\ type genePred\ visibility hide\ wgEncodeGencodeVersion 44\ wgEncodeGencodeV44ViewGenes Genes genePred All GENCODE annotations from V44 (Ensembl 110) 3 34.161 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV44 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV44\ longLabel All GENCODE annotations from V44 (Ensembl 110)\ parent wgEncodeGencodeV44\ shortLabel Genes\ track wgEncodeGencodeV44ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV44ViewPolya PolyA genePred All GENCODE annotations from V44 (Ensembl 110) 0 34.161 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V44 (Ensembl 110)\ parent wgEncodeGencodeV44\ shortLabel PolyA\ track wgEncodeGencodeV44ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV43 All GENCODE V43 genePred All GENCODE annotations from V43 (Ensembl 109) 0 34.162 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 43, Feb 2023) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 43 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 43 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 43 corresponds to Ensembl 109.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V43 (Ensembl 109)\ maxTransEnabled on\ priority 34.162\ shortLabel All GENCODE V43\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV43\ type genePred\ visibility hide\ wgEncodeGencodeVersion 43\ wgEncodeGencodeV43ViewGenes Genes genePred All GENCODE annotations from V43 (Ensembl 109) 3 34.162 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,PAR,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,PAR,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV43 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV43\ longLabel All GENCODE annotations from V43 (Ensembl 109)\ parent wgEncodeGencodeV43\ shortLabel Genes\ track wgEncodeGencodeV43ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV43ViewPolya PolyA genePred All GENCODE annotations from V43 (Ensembl 109) 0 34.162 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V43 (Ensembl 109)\ parent wgEncodeGencodeV43\ shortLabel PolyA\ track wgEncodeGencodeV43ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV42 All GENCODE V42 genePred All GENCODE annotations from V42 (Ensembl 108) 0 34.163 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 42, Oct 2022) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 42 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 42 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 42 corresponds to Ensembl 108.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V42 (Ensembl 108)\ maxTransEnabled on\ priority 34.163\ shortLabel All GENCODE V42\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes bPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV42\ type genePred\ visibility hide\ wgEncodeGencodeVersion 42\ wgEncodeGencodeV42ViewGenes Genes genePred All GENCODE annotations from V42 (Ensembl 108) 3 34.163 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,PAR,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_CDS_not_defined,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,overlaps_pseudogene,PAR,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV42 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV42\ longLabel All GENCODE annotations from V42 (Ensembl 108)\ parent wgEncodeGencodeV42\ shortLabel Genes\ track wgEncodeGencodeV42ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV42ViewPolya PolyA genePred All GENCODE annotations from V42 (Ensembl 108) 0 34.163 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V42 (Ensembl 108)\ parent wgEncodeGencodeV42\ shortLabel PolyA\ track wgEncodeGencodeV42ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV41View2Way 2-Way genePred All GENCODE annotations from V41 (Ensembl 107) 0 34.164 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V41 (Ensembl 107)\ parent wgEncodeGencodeV41\ shortLabel 2-Way\ track wgEncodeGencodeV41View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV41 All GENCODE V41 genePred All GENCODE annotations from V41 (Ensembl 107) 0 34.164 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 41, July 2022) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 41 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 41 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 41 corresponds to Ensembl 107.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V41 (Ensembl 107)\ priority 34.164\ shortLabel All GENCODE V41\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper pack\ track wgEncodeGencodeV41\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV41\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV41\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV41\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV41\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV41\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV41\ wgEncodeGencodePdb wgEncodeGencodePdbV41\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV41\ wgEncodeGencodePubMed wgEncodeGencodePubMedV41\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV41\ wgEncodeGencodeTag wgEncodeGencodeTagV41\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV41\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV41\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV41\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV41\ wgEncodeGencodeVersion 41\ wgEncodeGencodeV41ViewGenes Genes genePred All GENCODE annotations from V41 (Ensembl 107) 3 34.164 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=artifact,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,processed_pseudogene,processed_transcript,protein_coding,protein_coding_LoF,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,artifactual_duplication,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_gene,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV41 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV41\ longLabel All GENCODE annotations from V41 (Ensembl 107)\ parent wgEncodeGencodeV41\ shortLabel Genes\ track wgEncodeGencodeV41ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV41ViewPolya PolyA genePred All GENCODE annotations from V41 (Ensembl 107) 0 34.164 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V41 (Ensembl 107)\ parent wgEncodeGencodeV41\ shortLabel PolyA\ track wgEncodeGencodeV41ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV40View2Way 2-Way genePred All GENCODE annotations from V40 (Ensembl 106) 0 34.165 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V40 (Ensembl 106)\ parent wgEncodeGencodeV40\ shortLabel 2-Way\ track wgEncodeGencodeV40View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV40 All GENCODE V40 genePred All GENCODE annotations from V40 (Ensembl 106) 0 34.165 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 40, Feb 2022) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 40 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 40 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 40 corresponds to Ensembl 106.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V40 (Ensembl 106)\ priority 34.165\ shortLabel All GENCODE V40\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV40\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV40\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV40\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV40\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV40\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV40\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV40\ wgEncodeGencodePdb wgEncodeGencodePdbV40\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV40\ wgEncodeGencodePubMed wgEncodeGencodePubMedV40\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV40\ wgEncodeGencodeTag wgEncodeGencodeTagV40\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV40\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV40\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV40\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV40\ wgEncodeGencodeVersion 40\ wgEncodeGencodeV40ViewGenes Genes genePred All GENCODE annotations from V40 (Ensembl 106) 3 34.165 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV40 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV40\ longLabel All GENCODE annotations from V40 (Ensembl 106)\ parent wgEncodeGencodeV40\ shortLabel Genes\ track wgEncodeGencodeV40ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV40ViewPolya PolyA genePred All GENCODE annotations from V40 (Ensembl 106) 0 34.165 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V40 (Ensembl 106)\ parent wgEncodeGencodeV40\ shortLabel PolyA\ track wgEncodeGencodeV40ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV39View2Way 2-Way genePred All GENCODE annotations from V39 (Ensembl 105) 0 34.166 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V39 (Ensembl 105)\ parent wgEncodeGencodeV39\ shortLabel 2-Way\ track wgEncodeGencodeV39View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV39 All GENCODE V39 genePred All GENCODE annotations from V39 (Ensembl 105) 0 34.166 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 39, Oct 2021) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 39 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 39 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 39 corresponds to Ensembl 105.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V39 (Ensembl 105)\ priority 34.166\ shortLabel All GENCODE V39\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV39\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV39\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV39\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV39\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV39\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV39\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV39\ wgEncodeGencodePdb wgEncodeGencodePdbV39\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV39\ wgEncodeGencodePubMed wgEncodeGencodePubMedV39\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV39\ wgEncodeGencodeTag wgEncodeGencodeTagV39\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV39\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV39\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV39\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV39\ wgEncodeGencodeVersion 39\ wgEncodeGencodeV39ViewGenes Genes genePred All GENCODE annotations from V39 (Ensembl 105) 3 34.166 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV39 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV39\ longLabel All GENCODE annotations from V39 (Ensembl 105)\ parent wgEncodeGencodeV39\ shortLabel Genes\ track wgEncodeGencodeV39ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV39ViewPolya PolyA genePred All GENCODE annotations from V39 (Ensembl 105) 0 34.166 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V39 (Ensembl 105)\ parent wgEncodeGencodeV39\ shortLabel PolyA\ track wgEncodeGencodeV39ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV38View2Way 2-Way genePred All GENCODE annotations from V38 (Ensembl 104) 0 34.167 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V38 (Ensembl 104)\ parent wgEncodeGencodeV38\ shortLabel 2-Way\ track wgEncodeGencodeV38View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV38 All GENCODE V38 genePred All GENCODE annotations from V38 (Ensembl 104) 0 34.167 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 38, May 2021) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 38 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 38 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 38 corresponds to Ensembl 104.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V38 (Ensembl 104)\ priority 34.167\ shortLabel All GENCODE V38\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV38\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV38\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV38\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV38\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV38\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV38\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV38\ wgEncodeGencodePdb wgEncodeGencodePdbV38\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV38\ wgEncodeGencodePubMed wgEncodeGencodePubMedV38\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV38\ wgEncodeGencodeTag wgEncodeGencodeTagV38\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV38\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV38\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV38\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV38\ wgEncodeGencodeVersion 38\ wgEncodeGencodeV38ViewGenes Genes genePred All GENCODE annotations from V38 (Ensembl 104) 3 34.167 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,Ensembl_canonical,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV38 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV38\ longLabel All GENCODE annotations from V38 (Ensembl 104)\ parent wgEncodeGencodeV38\ shortLabel Genes\ track wgEncodeGencodeV38ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV38ViewPolya PolyA genePred All GENCODE annotations from V38 (Ensembl 104) 0 34.167 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V38 (Ensembl 104)\ parent wgEncodeGencodeV38\ shortLabel PolyA\ track wgEncodeGencodeV38ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV37View2Way 2-Way genePred All GENCODE annotations from V37 (Ensembl 103) 0 34.168 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V37 (Ensembl 103)\ parent wgEncodeGencodeV37\ shortLabel 2-Way\ track wgEncodeGencodeV37View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV37 All GENCODE V37 genePred All GENCODE annotations from V37 (Ensembl 103) 0 34.168 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 37, Feb 2021) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 37 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 37 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 37 corresponds to Ensembl 103.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V37 (Ensembl 103)\ priority 34.168\ shortLabel All GENCODE V37\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV37\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV37\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV37\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV37\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV37\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV37\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV37\ wgEncodeGencodePdb wgEncodeGencodePdbV37\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV37\ wgEncodeGencodePubMed wgEncodeGencodePubMedV37\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV37\ wgEncodeGencodeTag wgEncodeGencodeTagV37\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV37\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV37\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV37\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV37\ wgEncodeGencodeVersion 37\ wgEncodeGencodeV37ViewGenes Genes genePred All GENCODE annotations from V37 (Ensembl 103) 3 34.168 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Plus_Clinical,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV37 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV37\ longLabel All GENCODE annotations from V37 (Ensembl 103)\ parent wgEncodeGencodeV37\ shortLabel Genes\ track wgEncodeGencodeV37ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV37ViewPolya PolyA genePred All GENCODE annotations from V37 (Ensembl 103) 0 34.168 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V37 (Ensembl 103)\ parent wgEncodeGencodeV37\ shortLabel PolyA\ track wgEncodeGencodeV37ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV36View2Way 2-Way genePred All GENCODE annotations from V36 (Ensembl 102) 0 34.169 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V36 (Ensembl 102)\ parent wgEncodeGencodeV36\ shortLabel 2-Way\ track wgEncodeGencodeV36View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV36 All GENCODE V36 genePred All GENCODE annotations from V36 (Ensembl 102) 0 34.169 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 36, Nov 2020) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 36 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 36 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 36 corresponds to Ensembl 102.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V36 (Ensembl 102)\ priority 34.169\ shortLabel All GENCODE V36\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV36\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV36\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV36\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV36\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV36\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV36\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV36\ wgEncodeGencodePdb wgEncodeGencodePdbV36\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV36\ wgEncodeGencodePubMed wgEncodeGencodePubMedV36\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV36\ wgEncodeGencodeTag wgEncodeGencodeTagV36\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV36\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV36\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV36\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV36\ wgEncodeGencodeVersion 36\ wgEncodeGencodeV36ViewGenes Genes genePred All GENCODE annotations from V36 (Ensembl 102) 3 34.169 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV36 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV36\ longLabel All GENCODE annotations from V36 (Ensembl 102)\ parent wgEncodeGencodeV36\ shortLabel Genes\ track wgEncodeGencodeV36ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV36ViewPolya PolyA genePred All GENCODE annotations from V36 (Ensembl 102) 0 34.169 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V36 (Ensembl 102)\ parent wgEncodeGencodeV36\ shortLabel PolyA\ track wgEncodeGencodeV36ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV35View2Way 2-Way genePred All GENCODE annotations from V35 (Ensembl 101) 0 34.17 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V35 (Ensembl 101)\ parent wgEncodeGencodeV35\ shortLabel 2-Way\ track wgEncodeGencodeV35View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV35 All GENCODE V35 genePred All GENCODE annotations from V35 (Ensembl 101) 0 34.17 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 35, Aug 2020) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 35 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 35 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 35 corresponds to Ensembl 101.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V35 (Ensembl 101)\ priority 34.170\ shortLabel All GENCODE V35\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV35\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV35\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV35\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV35\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV35\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV35\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV35\ wgEncodeGencodePdb wgEncodeGencodePdbV35\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV35\ wgEncodeGencodePubMed wgEncodeGencodePubMedV35\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV35\ wgEncodeGencodeTag wgEncodeGencodeTagV35\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV35\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV35\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV35\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV35\ wgEncodeGencodeVersion 35\ wgEncodeGencodeV35ViewGenes Genes genePred All GENCODE annotations from V35 (Ensembl 101) 3 34.17 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vault_RNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV35 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV35\ longLabel All GENCODE annotations from V35 (Ensembl 101)\ parent wgEncodeGencodeV35\ shortLabel Genes\ track wgEncodeGencodeV35ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV35ViewPolya PolyA genePred All GENCODE annotations from V35 (Ensembl 101) 0 34.17 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V35 (Ensembl 101)\ parent wgEncodeGencodeV35\ shortLabel PolyA\ track wgEncodeGencodeV35ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV34View2Way 2-Way genePred All GENCODE annotations from V34 (Ensembl 100) 0 34.171 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V34 (Ensembl 100)\ parent wgEncodeGencodeV34\ shortLabel 2-Way\ track wgEncodeGencodeV34View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV34 All GENCODE V34 genePred All GENCODE annotations from V34 (Ensembl 100) 0 34.171 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 34, April 2020) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 34 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 34 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 34 corresponds to Ensembl 100.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V34 (Ensembl 100)\ priority 34.171\ shortLabel All GENCODE V34\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV34\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV34\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV34\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV34\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV34\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV34\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV34\ wgEncodeGencodePdb wgEncodeGencodePdbV34\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV34\ wgEncodeGencodePubMed wgEncodeGencodePubMedV34\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV34\ wgEncodeGencodeTag wgEncodeGencodeTagV34\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV34\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV34\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV34\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV34\ wgEncodeGencodeVersion 34\ wgEncodeGencodeV34ViewGenes Genes genePred All GENCODE annotations from V34 (Ensembl 100) 3 34.171 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV34 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV34\ longLabel All GENCODE annotations from V34 (Ensembl 100)\ parent wgEncodeGencodeV34\ shortLabel Genes\ track wgEncodeGencodeV34ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV34ViewPolya PolyA genePred All GENCODE annotations from V34 (Ensembl 100) 0 34.171 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V34 (Ensembl 100)\ parent wgEncodeGencodeV34\ shortLabel PolyA\ track wgEncodeGencodeV34ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV33View2Way 2-Way genePred All GENCODE annotations from V33 (Ensembl 99) 0 34.172 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V33 (Ensembl 99)\ parent wgEncodeGencodeV33\ shortLabel 2-Way\ track wgEncodeGencodeV33View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV33 All GENCODE V33 genePred All GENCODE annotations from V33 (Ensembl 99) 0 34.172 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 33, Jan 2020) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 33 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 33 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 33 corresponds to Ensembl 99.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V33 (Ensembl 99)\ priority 34.172\ shortLabel All GENCODE V33\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV33\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV33\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV33\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV33\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV33\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV33\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV33\ wgEncodeGencodePdb wgEncodeGencodePdbV33\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV33\ wgEncodeGencodePubMed wgEncodeGencodePubMedV33\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV33\ wgEncodeGencodeTag wgEncodeGencodeTagV33\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV33\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV33\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV33\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV33\ wgEncodeGencodeVersion 33\ wgEncodeGencodeV33ViewGenes Genes genePred All GENCODE annotations from V33 (Ensembl 99) 3 34.172 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV33 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV33\ longLabel All GENCODE annotations from V33 (Ensembl 99)\ parent wgEncodeGencodeV33\ shortLabel Genes\ track wgEncodeGencodeV33ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV33ViewPolya PolyA genePred All GENCODE annotations from V33 (Ensembl 99) 0 34.172 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V33 (Ensembl 99)\ parent wgEncodeGencodeV33\ shortLabel PolyA\ track wgEncodeGencodeV33ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV32View2Way 2-Way genePred All GENCODE annotations from V32 (Ensembl 98) 0 34.173 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V32 (Ensembl 98)\ parent wgEncodeGencodeV32\ shortLabel 2-Way\ track wgEncodeGencodeV32View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV32 All GENCODE V32 genePred All GENCODE annotations from V32 (Ensembl 98) 0 34.173 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 32, Sept 2019) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 32 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\GENCODE GFF3 and GTF files are available from the\ GENCODE release 32 site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 32 corresponds to Ensembl 98.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V32 (Ensembl 98)\ priority 34.173\ shortLabel All GENCODE V32\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV32\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV32\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV32\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV32\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV32\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV32\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV32\ wgEncodeGencodePdb wgEncodeGencodePdbV32\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV32\ wgEncodeGencodePubMed wgEncodeGencodePubMedV32\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV32\ wgEncodeGencodeTag wgEncodeGencodeTagV32\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV32\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV32\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV32\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV32\ wgEncodeGencodeVersion 32\ wgEncodeGencodeV32ViewGenes Genes genePred All GENCODE annotations from V32 (Ensembl 98) 3 34.173 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,stop_codon_readthrough,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV32 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV32\ longLabel All GENCODE annotations from V32 (Ensembl 98)\ parent wgEncodeGencodeV32\ shortLabel Genes\ track wgEncodeGencodeV32ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV32ViewPolya PolyA genePred All GENCODE annotations from V32 (Ensembl 98) 0 34.173 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V32 (Ensembl 98)\ parent wgEncodeGencodeV32\ shortLabel PolyA\ track wgEncodeGencodeV32ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV31View2Way 2-Way genePred All GENCODE annotations from V31 (Ensembl 97) 0 34.174 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V31 (Ensembl 97)\ parent wgEncodeGencodeV31\ shortLabel 2-Way\ track wgEncodeGencodeV31View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV31 All GENCODE V31 genePred All GENCODE annotations from V31 (Ensembl 97) 0 34.174 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 31, June 2019) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 31 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ GENCODE GFF3 and GTF files are available from the\ GENCODE release 31\ site.
\ \\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 31 corresponds to Ensembl 97.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V31 (Ensembl 97)\ priority 34.174\ shortLabel All GENCODE V31\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV31\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV31\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV31\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV31\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV31\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV31\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV31\ wgEncodeGencodePdb wgEncodeGencodePdbV31\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV31\ wgEncodeGencodePubMed wgEncodeGencodePubMedV31\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV31\ wgEncodeGencodeTag wgEncodeGencodeTagV31\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV31\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV31\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV31\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV31\ wgEncodeGencodeVersion 31\ wgEncodeGencodeV31ViewGenes Genes genePred All GENCODE annotations from V31 (Ensembl 97) 3 34.174 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,TAGENE,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV31 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV31\ longLabel All GENCODE annotations from V31 (Ensembl 97)\ parent wgEncodeGencodeV31\ shortLabel Genes\ track wgEncodeGencodeV31ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV31ViewPolya PolyA genePred All GENCODE annotations from V31 (Ensembl 97) 0 34.174 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V31 (Ensembl 97)\ parent wgEncodeGencodeV31\ shortLabel PolyA\ track wgEncodeGencodeV31ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV30View2Way 2-Way genePred All GENCODE annotations from V30 (Ensembl 96) 0 34.175 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V30 (Ensembl 96)\ parent wgEncodeGencodeV30\ shortLabel 2-Way\ track wgEncodeGencodeV30View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV30 All GENCODE V30 genePred All GENCODE annotations from V30 (Ensembl 96) 0 34.175 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 30, Apr 2019) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 30 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 30 corresponds to Ensembl 96.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V30 (Ensembl 96)\ priority 34.175\ shortLabel All GENCODE V30\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV30\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV30\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV30\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV30\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV30\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV30\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV30\ wgEncodeGencodePdb wgEncodeGencodePdbV30\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV30\ wgEncodeGencodePubMed wgEncodeGencodePubMedV30\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV30\ wgEncodeGencodeTag wgEncodeGencodeTagV30\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV30\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV30\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV30\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV30\ wgEncodeGencodeVersion 30\ wgEncodeGencodeV30ViewGenes Genes genePred All GENCODE annotations from V30 (Ensembl 96) 3 34.175 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,MANE_Select,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV30 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV30\ longLabel All GENCODE annotations from V30 (Ensembl 96)\ parent wgEncodeGencodeV30\ shortLabel Genes\ track wgEncodeGencodeV30ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV30ViewPolya PolyA genePred All GENCODE annotations from V30 (Ensembl 96) 0 34.175 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V30 (Ensembl 96)\ parent wgEncodeGencodeV30\ shortLabel PolyA\ track wgEncodeGencodeV30ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV29View2Way 2-Way genePred All GENCODE annotations from V29 (Ensembl 94) 0 34.176 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V29 (Ensembl 94)\ parent wgEncodeGencodeV29\ shortLabel 2-Way\ track wgEncodeGencodeV29View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV29 All GENCODE V29 genePred All GENCODE annotations from V29 (Ensembl 94) 0 34.176 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 29, Oct 2018) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 29 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 29 corresponds to Ensembl 94.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V29 (Ensembl 94)\ priority 34.176\ shortLabel All GENCODE V29\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV29\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV29\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV29\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV29\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV29\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV29\ wgEncodeGencodeHgnc wgEncodeGencodeHgncV29\ wgEncodeGencodePdb wgEncodeGencodePdbV29\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV29\ wgEncodeGencodePubMed wgEncodeGencodePubMedV29\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV29\ wgEncodeGencodeTag wgEncodeGencodeTagV29\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV29\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV29\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV29\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV29\ wgEncodeGencodeVersion 29\ wgEncodeGencodeV29ViewGenes Genes genePred All GENCODE annotations from V29 (Ensembl 94) 3 34.176 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,orphan,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,rRNA_pseudogene,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,orphan,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV29 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV29\ longLabel All GENCODE annotations from V29 (Ensembl 94)\ parent wgEncodeGencodeV29\ shortLabel Genes\ track wgEncodeGencodeV29ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV29ViewPolya PolyA genePred All GENCODE annotations from V29 (Ensembl 94) 0 34.176 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V29 (Ensembl 94)\ parent wgEncodeGencodeV29\ shortLabel PolyA\ track wgEncodeGencodeV29ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV28View2Way 2-Way genePred All GENCODE annotations from V28 (Ensembl 92) 0 34.177 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V28 (Ensembl 92)\ parent wgEncodeGencodeV28\ shortLabel 2-Way\ track wgEncodeGencodeV28View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV28 All GENCODE V28 genePred All GENCODE annotations from V28 (Ensembl 92) 0 34.177 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 28, Apr 2018) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 28 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 28 corresponds to Ensembl 92.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V28 (Ensembl 92)\ priority 34.177\ shortLabel All GENCODE V28\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV28\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV28\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV28\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV28\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV28\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV28\ wgEncodeGencodePdb wgEncodeGencodePdbV28\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV28\ wgEncodeGencodePubMed wgEncodeGencodePubMedV28\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV28\ wgEncodeGencodeTag wgEncodeGencodeTagV28\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV28\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV28\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV28\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV28\ wgEncodeGencodeVersion 28\ wgEncodeGencodeV28ViewGenes Genes genePred All GENCODE annotations from V28 (Ensembl 92) 3 34.177 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,orphan,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CAGE_supported_TSS,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,orphan,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV28 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV28\ longLabel All GENCODE annotations from V28 (Ensembl 92)\ parent wgEncodeGencodeV28\ shortLabel Genes\ track wgEncodeGencodeV28ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV28ViewPolya PolyA genePred All GENCODE annotations from V28 (Ensembl 92) 0 34.177 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V28 (Ensembl 92)\ parent wgEncodeGencodeV28\ shortLabel PolyA\ track wgEncodeGencodeV28ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV27View2Way 2-Way genePred All GENCODE annotations from V27 (Ensembl 90) 0 34.178 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V27 (Ensembl 90)\ parent wgEncodeGencodeV27\ shortLabel 2-Way\ track wgEncodeGencodeV27View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV27 All GENCODE V27 genePred All GENCODE annotations from V27 (Ensembl 90) 0 34.178 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 27, Aug 2017) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 27 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 27 corresponds to Ensembl 90.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V27 (Ensembl 90)\ priority 34.178\ shortLabel All GENCODE V27\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV27\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV27\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV27\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV27\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV27\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV27\ wgEncodeGencodePdb wgEncodeGencodePdbV27\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV27\ wgEncodeGencodePubMed wgEncodeGencodePubMedV27\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV27\ wgEncodeGencodeTag wgEncodeGencodeTagV27\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV27\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV27\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV27\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV27\ wgEncodeGencodeVersion 27\ wgEncodeGencodeV27ViewGenes Genes genePred All GENCODE annotations from V27 (Ensembl 90) 3 34.178 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense_RNA,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,orphan,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense_RNA,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,fragmented_locus,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,ncRNA_host,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,orphan,overlapping_locus,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,reference_genome_error,retained_intron_CDS,retained_intron_final,retained_intron_first,retrogene,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,semi_processed,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV27 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV27\ longLabel All GENCODE annotations from V27 (Ensembl 90)\ parent wgEncodeGencodeV27\ shortLabel Genes\ track wgEncodeGencodeV27ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV27ViewPolya PolyA genePred All GENCODE annotations from V27 (Ensembl 90) 0 34.178 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V27 (Ensembl 90)\ parent wgEncodeGencodeV27\ shortLabel PolyA\ track wgEncodeGencodeV27ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV26View2Way 2-Way genePred All GENCODE annotations from V26 (Ensembl 88) 0 34.179 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V26 (Ensembl 88)\ parent wgEncodeGencodeV26\ shortLabel 2-Way\ track wgEncodeGencodeV26View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV26 All GENCODE V26 genePred All GENCODE annotations from V26 (Ensembl 88) 0 34.179 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 26, March 2017) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The 26 annotation was carried out on genome assembly GRCh38 (hg38).\
\\ The Ensembl human and mouse data sets are the same gene annotations as GENCODE for the\ corresponding release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 26 corresponds to Ensembl 88.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE annotations from V26 (Ensembl 88)\ priority 34.179\ shortLabel All GENCODE V26\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV26\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV26\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV26\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV26\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV26\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV26\ wgEncodeGencodePdb wgEncodeGencodePdbV26\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV26\ wgEncodeGencodePubMed wgEncodeGencodePubMedV26\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV26\ wgEncodeGencodeTag wgEncodeGencodeTagV26\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV26\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV26\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV26\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV26\ wgEncodeGencodeVersion 26\ wgEncodeGencodeV26ViewGenes Genes genePred All GENCODE annotations from V26 (Ensembl 88) 3 34.179 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=3_nested_supported_extension,3_standard_supported_extension,454_RNA_Seq_supported,5_nested_supported_extension,5_standard_supported_extension,alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,nested_454_RNA_Seq_supported,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,pseudo_consens,readthrough_transcript,retained_intron_CDS,retained_intron_final,retained_intron_first,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_D_pseudogene,IG_J_gene,IG_LV_gene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,pseudo_consens,readthrough_transcript,retained_intron_CDS,retained_intron_final,retained_intron_first,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV26 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV26\ longLabel All GENCODE annotations from V26 (Ensembl 88)\ parent wgEncodeGencodeV26\ shortLabel Genes\ track wgEncodeGencodeV26ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV26ViewPolya PolyA genePred All GENCODE annotations from V26 (Ensembl 88) 0 34.179 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE annotations from V26 (Ensembl 88)\ parent wgEncodeGencodeV26\ shortLabel PolyA\ track wgEncodeGencodeV26ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV25View2Way 2-Way genePred All GENCODE transcripts including comprehensive set V25 0 34.18 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V25\ parent wgEncodeGencodeV25\ shortLabel 2-Way\ track wgEncodeGencodeV25View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV25 All GENCODE V25 genePred All GENCODE transcripts including comprehensive set V25 0 34.18 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 25, July 2016) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The annotation was carried out on genome assembly GRCh38 (hg38).\
\\ As of GENCODE Version 11, Ensembl and GENCODE have converged. The gene\ annotations in the GENCODE comprehensive set are the same as the corresponding\ Ensembl release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 25 corresponds to Ensembl 85.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE transcripts including comprehensive set V25\ priority 34.180\ shortLabel All GENCODE V25\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV25\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV25\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV25\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV25\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV25\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV25\ wgEncodeGencodePdb wgEncodeGencodePdbV25\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV25\ wgEncodeGencodePubMed wgEncodeGencodePubMedV25\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV25\ wgEncodeGencodeTag wgEncodeGencodeTagV25\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV25\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV25\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV25\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV25\ wgEncodeGencodeVersion 25\ wgEncodeGencodeV25ViewGenes Genes genePred All GENCODE transcripts including comprehensive set V25 3 34.18 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,pseudo_consens,readthrough_transcript,retained_intron_CDS,retained_intron_final,retained_intron_first,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncRNA,antisense,bidirectional_promoter_lncRNA,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,scRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,bicistronic,CCDS,cds_end_NF,cds_start_NF,dotter_confirmed,downstream_ATG,exp_conf,inferred_exon_combination,inferred_transcript_model,low_sequence_quality,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,non_submitted_evidence,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,pseudo_consens,readthrough_transcript,retained_intron_CDS,retained_intron_final,retained_intron_first,RNA_Seq_supported_only,RNA_Seq_supported_partial,RP_supported_TIS,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV25 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV25\ longLabel All GENCODE transcripts including comprehensive set V25\ parent wgEncodeGencodeV25\ shortLabel Genes\ track wgEncodeGencodeV25ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV25ViewPolya PolyA genePred All GENCODE transcripts including comprehensive set V25 0 34.18 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V25\ parent wgEncodeGencodeV25\ shortLabel PolyA\ track wgEncodeGencodeV25ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV24View2Way 2-Way genePred All GENCODE transcripts including comprehensive set V24 0 34.181 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V24\ parent wgEncodeGencodeV24\ shortLabel 2-Way\ track wgEncodeGencodeV24View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV24 All GENCODE V24 genePred All GENCODE transcripts including comprehensive set V24 0 34.181 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 24, December 2015) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The annotation was carried out on genome assembly GRCh38 (hg38).\
\\ As of GENCODE Version 11, Ensembl and GENCODE have converged. The gene\ annotations in the GENCODE comprehensive set are the same as the corresponding\ Ensembl release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 24 corresponds to Ensembl 84.
\ \See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE transcripts including comprehensive set V24\ priority 34.181\ shortLabel All GENCODE V24\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV24\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV24\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV24\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV24\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV24\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV24\ wgEncodeGencodePdb wgEncodeGencodePdbV24\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV24\ wgEncodeGencodePubMed wgEncodeGencodePubMedV24\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV24\ wgEncodeGencodeTag wgEncodeGencodeTagV24\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV24\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV24\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV24\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV24\ wgEncodeGencodeVersion 24\ wgEncodeGencodeV24ViewGenes Genes genePred All GENCODE transcripts including comprehensive set V24 3 34.181 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,CCDS,cds_end_NF,cds_start_NF,downstream_ATG,exp_conf,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,CCDS,cds_end_NF,cds_start_NF,downstream_ATG,exp_conf,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV24 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV24\ longLabel All GENCODE transcripts including comprehensive set V24\ parent wgEncodeGencodeV24\ shortLabel Genes\ track wgEncodeGencodeV24ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV24ViewPolya PolyA genePred All GENCODE transcripts including comprehensive set V24 0 34.181 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V24\ parent wgEncodeGencodeV24\ shortLabel PolyA\ track wgEncodeGencodeV24ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV23View2Way 2-Way genePred All GENCODE transcripts including comprehensive set V23 0 34.182 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V23\ parent wgEncodeGencodeV23\ shortLabel 2-Way\ track wgEncodeGencodeV23View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV23 All GENCODE V23 genePred All GENCODE transcripts including comprehensive set V23 0 34.182 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 23, March 2015) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The annotation was carried out on genome assembly GRCh38 (hg38).\
\\ As of GENCODE Version 11, Ensembl and GENCODE have converged. The gene\ annotations in the GENCODE comprehensive set are the same as the corresponding\ Ensembl release.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 23 corresponds to Ensembl 81 and 82.
\ \See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE transcripts including comprehensive set V23\ priority 34.182\ shortLabel All GENCODE V23\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV23\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV23\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV23\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV23\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV23\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV23\ wgEncodeGencodePdb wgEncodeGencodePdbV23\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV23\ wgEncodeGencodePubMed wgEncodeGencodePubMedV23\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV23\ wgEncodeGencodeTag wgEncodeGencodeTagV23\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV23\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV23\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV23\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV23\ wgEncodeGencodeVersion 23\ wgEncodeGencodeV23ViewGenes Genes genePred All GENCODE transcripts including comprehensive set V23 3 34.182 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,CCDS,cds_end_NF,cds_start_NF,downstream_ATG,exp_conf,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,CCDS,cds_end_NF,cds_start_NF,downstream_ATG,exp_conf,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV23 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV23\ longLabel All GENCODE transcripts including comprehensive set V23\ parent wgEncodeGencodeV23\ shortLabel Genes\ track wgEncodeGencodeV23ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV23ViewPolya PolyA genePred All GENCODE transcripts including comprehensive set V23 0 34.182 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V23\ parent wgEncodeGencodeV23\ shortLabel PolyA\ track wgEncodeGencodeV23ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV22View2Way 2-Way genePred All GENCODE transcripts including comprehensive set V22 0 34.183 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V22\ parent wgEncodeGencodeV22\ shortLabel 2-Way\ track wgEncodeGencodeV22View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV22 All GENCODE V22 genePred All GENCODE transcripts including comprehensive set V22 0 34.183 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 22, March 2015) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The annotation was carried out on genome assembly GRCh38 (hg38).\
\\ As of GENCODE Version 11, Ensembl and GENCODE have converged. The gene\ annotations in the GENCODE comprehensive set are the same as the corresponding\ Ensembl release.\
\\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 22 corresponds to Ensembl 79.
\See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel All GENCODE transcripts including comprehensive set V22\ priority 34.183\ shortLabel All GENCODE V22\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV22\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV22\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV22\ wgEncodeGencodeEntrezGene wgEncodeGencodeEntrezGeneV22\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV22\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV22\ wgEncodeGencodePdb wgEncodeGencodePdbV22\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV22\ wgEncodeGencodePubMed wgEncodeGencodePubMedV22\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV22\ wgEncodeGencodeTag wgEncodeGencodeTagV22\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV22\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV22\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV22\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV22\ wgEncodeGencodeVersion 22\ wgEncodeGencodeV22ViewGenes Genes genePred All GENCODE transcripts including comprehensive set V22 3 34.183 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,CCDS,cds_end_NF,cds_start_NF,downstream_ATG,exp_conf,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,macro_lncRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_coding,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,ribozyme,rRNA,scaRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,sRNA,TEC,transcribed_processed_pseudogene,transcribed_unitary_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,translated_unprocessed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene,vaultRNA tag:Tag=alternative_3_UTR,alternative_5_UTR,appris_alternative_1,appris_alternative_2,appris_principal_1,appris_principal_2,appris_principal_3,appris_principal_4,appris_principal_5,basic,CCDS,cds_end_NF,cds_start_NF,downstream_ATG,exp_conf,mRNA_end_NF,mRNA_start_NF,NAGNAG_splice_site,NMD_exception,NMD_likely_if_extended,non_ATG_start,non_canonical_conserved,non_canonical_genome_sequence_error,non_canonical_other,non_canonical_polymorphism,non_canonical_TEC,non_canonical_U12,not_best_in_genome_evidence,not_organism_supported,overlapping_uORF,PAR,pseudo_consens,readthrough_transcript,seleno,sequence_error,upstream_ATG,upstream_uORF supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV22 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV22\ longLabel All GENCODE transcripts including comprehensive set V22\ parent wgEncodeGencodeV22\ shortLabel Genes\ track wgEncodeGencodeV22ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV22ViewPolya PolyA genePred All GENCODE transcripts including comprehensive set V22 0 34.183 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel All GENCODE transcripts including comprehensive set V22\ parent wgEncodeGencodeV22\ shortLabel PolyA\ track wgEncodeGencodeV22ViewPolya\ type genePred\ view cPolya\ visibility hide\ wgEncodeGencodeV20View2Way 2-Way genePred Gene Annotations from GENCODE Version 20 (Ensembl 76) 0 34.185 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel Gene Annotations from GENCODE Version 20 (Ensembl 76)\ parent wgEncodeGencodeV20\ shortLabel 2-Way\ track wgEncodeGencodeV20View2Way\ type genePred\ view b2-way\ visibility hide\ wgEncodeGencodeV20 GENCODE V20 (Ensembl 76) genePred Gene Annotations from GENCODE Version 20 (Ensembl 76) 0 34.185 0 0 0 127 127 127 0 0 0\ The GENCODE Genes track (version 20, August 2014) shows high-quality manual\ annotations merged with evidence-based automated annotations across the entire\ human genome generated by the\ GENCODE project.\ The GENCODE gene set presents a full merge\ between HAVANA manual annotation process and Ensembl automatic annotation pipeline.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations.\ The annotation was carried out on genome assembly GRCh38 (hg38).\
\\ As of GENCODE Version 11, Ensembl and GENCODE have converged. The gene\ annotations in the GENCODE comprehensive set are the same as the corresponding\ Ensembl release. UCSC will continue to provide a separate Ensembl track on\ Human in the same format as the Ensembl tracks on other organisms.\
\ \\ This track is a multi-view composite track that contains differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ To show only selected subtracks, uncheck the boxes next to the tracks that\ you wish to hide.
\ Views available on this track are:\\ Maximum number of transcripts to display\ is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks.\ Starting with the GENCODE human V42 and mouse VM31 releases, \ transcripts are assigned rank within the gene. The ranks may be used to filter the number of transcripts\ displayed in a principled manner. Transcript ranking is not available in the lift37 releases.\ See Methods for details of rank assignment.\
\ \Filtering is available for the items in the GENCODE Basic, Comprehensive and Pseudogene tracks\ using the following criteria:
\Coloring for the gene annotations is based on the annotation type:
\\ The GENCODE project aims to annotate all evidence-based gene features on the \ human and mouse reference sequence with high accuracy by integrating \ computational approaches (including comparative methods), manual\ annotation and targeted experimental verification. This goal includes identifying \ all protein-coding loci with associated alternative variants, non-coding\ loci which have transcript evidence, and pseudogenes. \ For a detailed description of the methods and references used, see\ Harrow et al. (2006).\
\ \\ GENCODE Basic Set selection:\ The GENCODE Basic Set is intended to provide a simplified subset of\ the GENCODE transcript annotations that will be useful to the majority of\ users. The goal was to have a high-quality basic set that also covered all loci. \ Selection of GENCODE annotations for inclusion in the basic set\ was determined independently for the coding and non-coding transcripts at each\ gene locus.\
\\ Non-coding transcript categorization: \ Non-coding transcripts are categorized using\ their biotype\ and the following criteria:\
\Transcript ranking:\ Within each gene, transcripts have been ranked according to the \ following criteria. The ranking approach is preliminary and will\ change is future releases.\
\ \\ Transcription Support Level (TSL):\ It is important that users understand how to assess transcript annotations\ that they see in GENCODE. While some transcript models have a high level of\ support through the full length of their exon structure, there are also\ transcripts that are poorly supported and that should be considered\ speculative. The Transcription Support Level (TSL) is a method to highlight the\ well-supported and poorly-supported transcript models for users. The method\ relies on the primary data that can support full-length transcript\ structure: mRNA and EST alignments supplied by UCSC and Ensembl.
\ \The mRNA and EST alignments are compared to the GENCODE transcripts and the\ transcripts are scored according to how well the alignment matches over its\ full length. \ The GENCODE TSL provides a consistent method of evaluating the\ level of support that a GENCODE transcript annotation is\ actually expressed in mouse. Mouse transcript sequences from the \ International Nucleotide\ Sequence Database Collaboration (GenBank, ENA, and DDBJ) are used as\ the evidence for this analysis.\ \ Exonerate RNA alignments from Ensembl,\ BLAT RNA and EST alignments from the UCSC Genome Browser Database are used in\ the analysis. Erroneous transcripts and libraries identified in lists\ maintained by the Ensembl, UCSC, HAVANA and RefSeq groups are flagged as\ suspect. GENCODE annotations for protein-coding and non-protein-coding\ transcripts are compared with the evidence alignments.
\ \Annotations in the MHC region and other immunological genes are not\ evaluated, as automatic alignments tend to be very problematic. \ Methods for evaluating single-exon genes are still being developed and \ they are not included\ in the current analysis. Multi-exon GENCODE annotations are evaluated using\ the criteria that all introns are supported by an evidence alignment and the\ evidence alignment does not indicate that there are unannotated exons. Small\ insertions and deletions in evidence alignments are assumed to be due to\ polymorphisms and not considered as differing from the annotations. All\ intron boundaries must match exactly. The transcript start and end locations\ are allowed to differ.
\ \The following categories are assigned to each of the evaluated annotations:
\ \APPRIS\ is a system to annotate alternatively spliced transcripts based on a range of computational\ methods. It provides value to the annotations of the human, mouse, zebrafish, rat, and pig genomes.\ APPRIS has selected a single CDS variant for each gene as the 'PRINCIPAL' isoform. Principal\ isoforms are tagged with the numbers 1 to 5, with 1 being the most reliable.
\\ Selected transcript models are verified experimentally by RT-PCR amplification followed by sequencing.\ Those experiments can be found at GEO:
\See Harrow et al. (2006) for information on verification\ techniques.\
\ \\ GENCODE version 20 corresponds to Ensembl 76 and Vega 56.
\ \See also: The GENCODE Project\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 1 allButtonPair on\ compositeTrack on\ configurable off\ dragAndDrop subTracks\ fileSortOrder labVersion=Contents dccAccession=UCSC_Accession\ group genes\ longLabel Gene Annotations from GENCODE Version 20 (Ensembl 76)\ priority 34.185\ shortLabel GENCODE V20 (Ensembl 76)\ sortOrder name=+ view=+\ subGroup1 view View aGenes=Genes b2-way=2-way cPolya=PolyA\ subGroup2 name Name Basic=Basic Comprehensive=Comprehensive Pseudogenes=Pseudogenes yTwo-way=2-way_Pseudogenes zPolyA=PolyA\ superTrack wgEncodeGencodeSuper hide\ track wgEncodeGencodeV20\ type genePred\ visibility hide\ wgEncodeGencodeAnnotationRemark wgEncodeGencodeAnnotationRemarkV20\ wgEncodeGencodeAttrs wgEncodeGencodeAttrsV20\ wgEncodeGencodeExonSupport wgEncodeGencodeExonSupportV20\ wgEncodeGencodeGeneSource wgEncodeGencodeGeneSourceV20\ wgEncodeGencodePdb wgEncodeGencodePdbV20\ wgEncodeGencodePolyAFeature wgEncodeGencodePolyAFeatureV20\ wgEncodeGencodePubMed wgEncodeGencodePubMedV20\ wgEncodeGencodeRefSeq wgEncodeGencodeRefSeqV20\ wgEncodeGencodeTag wgEncodeGencodeTagV20\ wgEncodeGencodeTranscriptSource wgEncodeGencodeTranscriptSourceV20\ wgEncodeGencodeTranscriptSupport wgEncodeGencodeTranscriptSupportV20\ wgEncodeGencodeTranscriptionSupportLevel wgEncodeGencodeTranscriptionSupportLevelV20\ wgEncodeGencodeUniProt wgEncodeGencodeUniProtV20\ wgEncodeGencodeVersion 20\ wgEncodeGencodeV20ViewGenes Genes genePred Gene Annotations from GENCODE Version 20 (Ensembl 76) 3 34.185 0 0 0 127 127 127 0 0 0 genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ cdsDrawDefault genomic\\ codons\ configurable on\ filterBy attrs.transcriptClass:Transcript_Class=coding,nonCoding,pseudo,problem transcriptMethod:Transcript_Annotation_Method=manual,automatic,manual_only,automatic_only attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,rRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,transcribed_processed_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA\ gClass_coding 12,12,120\ gClass_nonCoding 0,153,0\ gClass_problem 254,0,0\ gClass_pseudo 255,51,255\ geneClasses coding nonCoding pseudo problem\ highlightBy supportLevel:Support_Level=tsl1,tsl2,tsl3,tsl4,tsl5,tslNA attrs.transcriptType:Transcript_Biotype=3prime_overlapping_ncrna,antisense,IG_C_gene,IG_C_pseudogene,IG_D_gene,IG_J_gene,IG_J_pseudogene,IG_V_gene,IG_V_pseudogene,lincRNA,miRNA,misc_RNA,Mt_rRNA,Mt_tRNA,nonsense_mediated_decay,non_stop_decay,polymorphic_pseudogene,processed_pseudogene,processed_transcript,protein_coding,pseudogene,retained_intron,rRNA,sense_intronic,sense_overlapping,snoRNA,snRNA,transcribed_processed_pseudogene,transcribed_unprocessed_pseudogene,translated_processed_pseudogene,TR_C_gene,TR_D_gene,TR_J_gene,TR_J_pseudogene,TR_V_gene,TR_V_pseudogene,unitary_pseudogene,unprocessed_pseudogene\ highlightColor 255,255,0\ idXref wgEncodeGencodeAttrsV20 transcriptId geneId\ itemClassClassColumn transcriptClass\ itemClassNameColumn transcriptId\ itemClassTbl wgEncodeGencodeAttrsV20\ longLabel Gene Annotations from GENCODE Version 20 (Ensembl 76)\ parent wgEncodeGencodeV20\ shortLabel Genes\ track wgEncodeGencodeV20ViewGenes\ type genePred\ view aGenes\ visibility pack\ wgEncodeGencodeV20ViewPolya PolyA genePred Gene Annotations from GENCODE Version 20 (Ensembl 76) 0 34.185 0 0 0 127 127 127 0 0 0 genes 1 configurable off\ longLabel Gene Annotations from GENCODE Version 20 (Ensembl 76)\ parent wgEncodeGencodeV20\ shortLabel PolyA\ track wgEncodeGencodeV20ViewPolya\ type genePred\ view cPolya\ visibility hide\ encTfChipPkENCFF915LKZ A549 POLR2A 1 narrowPeak Transcription Factor ChIP-seq Peaks of POLR2A in A549 from ENCODE 3 (ENCFF915LKZ) 0 35 254 93 85 254 174 170 0 0 0 regulation 1 color 254,93,85\ longLabel Transcription Factor ChIP-seq Peaks of POLR2A in A549 from ENCODE 3 (ENCFF915LKZ)\ parent encTfChipPk off\ shortLabel A549 POLR2A 1\ subGroups cellType=A549 factor=POLR2A\ track encTfChipPkENCFF915LKZ\ AorticSmoothMuscleCellResponseToFGF203hrBiolRep2LK20_CNhs13364_ctss_fwd AorticSmsToFgf2_03hrBr2+ bigWig Aortic smooth muscle cell response to FGF2, 03hr, biol_rep2 (LK20)_CNhs13364_12746-136A1_forward 0 35 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12746-136A1 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2003hr%2c%20biol_rep2%20%28LK20%29.CNhs13364.12746-136A1.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 03hr, biol_rep2 (LK20)_CNhs13364_12746-136A1_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12746-136A1 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToFgf2_03hrBr2+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF203hrBiolRep2LK20_CNhs13364_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12746-136A1\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToFGF203hrBiolRep2LK20_CNhs13364_tpm_fwd AorticSmsToFgf2_03hrBr2+ bigWig Aortic smooth muscle cell response to FGF2, 03hr, biol_rep2 (LK20)_CNhs13364_12746-136A1_forward 1 35 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12746-136A1 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20FGF2%2c%2003hr%2c%20biol_rep2%20%28LK20%29.CNhs13364.12746-136A1.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to FGF2, 03hr, biol_rep2 (LK20)_CNhs13364_12746-136A1_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12746-136A1 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToFgf2_03hrBr2+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_FGF2 strand=forward\ track AorticSmoothMuscleCellResponseToFGF203hrBiolRep2LK20_CNhs13364_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12746-136A1\ urlLabel FANTOM5 Details:\ wgEncodeReg4MarkH3k27acAllBlood Blood (all biosamples) bigWig Avg. H3K27ac level of 152 blood experiments (all biosamples) 2 35 254 75 173 254 165 214 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/bloodH3K27ac.bw\ color 254,75,173\ longLabel Avg. H3K27ac level of 152 blood experiments (all biosamples)\ parent wgEncodeReg4MarkH3k27ac off\ priority 35\ shortLabel Blood (all biosamples)\ track wgEncodeReg4MarkH3k27acAllBlood\ type bigWig\ ENCFF278VYR_ENCFF971OSG_ENCFF238JTO_ENCFF264VOP ENCFF278VYR_ENCFF971OSG_ENCFF238JTO_ENCFF264VOP bigBed 9 + 5 Middle frontal area 46 (mild cognitive impairment), female adult (90 or above years) with mild cognitive impairment: (1) cCREs 4 35 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF278VYR_ENCFF971OSG_ENCFF238JTO_ENCFF264VOP.bb\ longLabel Middle frontal area 46 (mild cognitive impairment), female adult (90 or above years) with mild cognitive impairment: (1) cCREs\ mouseOver ID: ${name}\ This track shows alignments between human expressed sequence tags \ (ESTs) in GenBank and the genome. ESTs are single-read sequences, \ typically about 500 bases in length, that usually represent fragments of \ transcribed genes.
\\ NOTE: As of April, 2007, we no longer include GenBank sequences \ that contain the following URL as part of the record:\
\ http://fulllength.invitrogen.com\\ Some of these entries are the result of alignment to pseudogenes,\ followed by "correction" of the EST to match the genomic sequence. \ It is therefore not the sequence of the actual EST and makes it appear that \ the EST is transcribed. Invitrogen no longer sells the clones.\ \ \
\ This track follows the display conventions for \ PSL alignment tracks. In dense display mode, the items that\ are more darkly shaded indicate matches of better quality.
\\ The strand information (+/-) indicates the\ direction of the match between the EST and the matching\ genomic sequence. It bears no relationship to the direction\ of transcription of the RNA with which it might be associated.
\\ The description page for this track has a filter that can be used to change \ the display mode, alter the color, and include/exclude a subset of items \ within the track. This may be helpful when many items are shown in the track \ display, especially when only some are relevant to the current task.
\\ To use the filter:\
\ This track may also be configured to display base labeling, a feature that\ allows the user to display all bases in the aligning sequence or only those \ that differ from the genomic sequence. For more information about this option,\ click \ here.\ Several types of alignment gap may also be colored; \ for more information, click \ here.\
\ \\ To make an EST, RNA is isolated from cells and reverse\ transcribed into cDNA. Typically, the cDNA is cloned\ into a plasmid vector and a read is taken from the 5'\ and/or 3' primer. For most — but not all — ESTs, the\ reverse transcription is primed by an oligo-dT, which\ hybridizes with the poly-A tail of mature mRNA. The\ reverse transcriptase may or may not make it to the 5'\ end of the mRNA, which may or may not be degraded.
\\ In general, the 3' ESTs mark the end of transcription\ reasonably well, but the 5' ESTs may end at any point\ within the transcript. Some of the newer cap-selected\ libraries cover transcription start reasonably well. Before the \ cap-selection techniques\ emerged, some projects used random rather than poly-A\ priming in an attempt to retrieve sequence distant from the\ 3' end. These projects were successful at this, but as\ a side effect also deposited sequences from unprocessed\ mRNA and perhaps even genomic sequences into the EST databases.\ Even outside of the random-primed projects, there is a\ degree of non-mRNA contamination. Because of this, a\ single unspliced EST should be viewed with considerable\ skepticism.
\\ To generate this track, human ESTs from GenBank were aligned \ against the genome using blat. Note that the maximum intron length\ allowed by blat is 750,000 bases, which may eliminate some ESTs with very \ long introns that might otherwise align. When a single \ EST aligned in multiple places, the alignment having the \ highest base identity was identified. Only alignments having\ a base identity level within 0.5% of the best and at least 96% base identity \ with the genomic sequence were kept.
\ \\ This track was produced at UCSC from EST sequence data\ submitted to the international public sequence databases by \ scientists worldwide.
\ \\ Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Wheeler DL.\ GenBank: update. Nucleic Acids Res.\ 2004 Jan 1;32(Database issue):D23-6.
\\ Kent WJ.\ BLAT - The BLAST-Like Alignment Tool.\ Genome Res. 2002 Apr;12(4):656-64.
\ rna 1 baseColorUseSequence genbank\ group rna\ indelDoubleInsert on\ indelQueryInsert on\ intronGap 30\ longLabel Human ESTs Including Unspliced\ maxItems 300\ shortLabel Human ESTs\ spectrum on\ table all_est\ track est\ type psl est\ visibility hide\ mrna Human mRNAs psl . Human mRNAs from GenBank 0 100 0 0 0 127 127 127 1 0 0\ The mRNA track shows alignments between human mRNAs\ in \ GenBank and the genome.
\ \\ This track follows the display conventions for\ \ PSL alignment tracks. In dense display mode, the items that\ are more darkly shaded indicate matches of better quality.\
\ \\ The description page for this track has a filter that can be used to change\ the display mode, alter the color, and include/exclude a subset of items\ within the track. This may be helpful when many items are shown in the track\ display, especially when only some are relevant to the current task.\
\ \\ To use the filter:\
\ This track may also be configured to display codon coloring, a feature that\ allows the user to quickly compare mRNAs against the genomic sequence. For more\ information about this option, go to the\ \ Codon and Base Coloring for Alignment Tracks page.\ Several types of alignment gap may also be colored;\ for more information, go to the\ \ Alignment Insertion/Deletion Display Options page.\
\ \\ GenBank human mRNAs were aligned against the genome using the\ blat program. When a single mRNA aligned in multiple places,\ the alignment having the highest base identity was found.\ Only alignments having a base identity level within 0.5% of\ the best and at least 96% base identity with the genomic sequence were kept.\
\ \\ The mRNA track was produced at UCSC from mRNA sequence data\ submitted to the international public sequence databases by\ scientists worldwide.\
\ \\ Benson DA, Cavanaugh M, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW.\ \ GenBank.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D36-42.\ PMID: 23193287; PMC: PMC3531190\
\ \\ Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Wheeler DL.\ GenBank: update.\ Nucleic Acids Res. 2004 Jan 1;32(Database issue):D23-6.\ PMID: 14681350; PMC: PMC308779\
\ \\ Kent WJ.\ BLAT - the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ rna 1 baseColorDefault diffCodons\ baseColorUseCds genbank\ baseColorUseSequence genbank\ group rna\ indelDoubleInsert on\ indelPolyA on\ indelQueryInsert on\ longLabel Human mRNAs from GenBank\ shortLabel Human mRNAs\ showDiffBasesAllScales .\ spectrum on\ table all_mrna\ track mrna\ type psl .\ visibility hide\ tgpArchive 1000 Genomes 1000 Genomes Phase 3 0 100 0 0 0 127 127 127 0 0 0\ This supertrack is a collection of tracks from the\ 1000 Genomes Project showing\ paired-end accessible regions and integrated variant calls. More information about display\ conventions, methods, credits, and references can be found on each subtrack's description page.\
\\ For more details, see:
\\ Thanks to the International Genome Sample Resource (IGSR) for making these variant calls\ freely available.
\ varRep 0 cartVersion 2\ group varRep\ html ../tgpArchive\ longLabel 1000 Genomes Phase 3\ shortLabel 1000 Genomes\ superTrack on\ track tgpArchive\ visibility hide\ tgpTrios 1000 Genomes Trios vcfPhasedTrio Thousand Genomes Project Family VCF Trios 3 100 0 0 0 127 127 127 0 0 23 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX,\ This track shows approximately 4.5 million single nucleotide variants (SNVs) and\ 0.6 million short insertions/deletions (indels) from 7 different parent/child trios as\ produced by the\ International\ Genome Sample Resource (IGSR), from sequence data generated by the\ 1000 Genomes Project\ in its Phase 3 sequencing of 2,504 genomes from 16 populations worldwide.
\\ Variants were called on the autosomes (chromosomes 1 through 22) and on the\ Pseudo-Autosomal Regions (PARs) of chromosome X.\ Therefore this track has no annotations on alternate haplotype sequences, fix patches,\ chromosome Y, or the non-PAR portion (the majority) of chromosome X.\
\\ The variant genotypes have been phased (i.e., the two alleles of each diploid genotype\ have been assigned to two\ haplotypes,\ one inherited from each parent). This information allows us to illustrate which\ haplotypes in the child have been inherited from which parent.\
\ \Trios from six different populations are available, including:\
\ This track illustrates the vcfPhasedTrio track type, where two lines, one for each chromosome\ in the diploid genome, is drawn per sample in the underlying VCF. Variants in the window\ are then drawn on the haplotype line corresponding to which haplotype they belong to, such that\ variants on the same line were likely inherited together. The sorting routine is the same as\ what is used to draw the haplotype sorted display in the non-trio 1000 Genomes track, and is\ described here.\
\ \\ The child haplotypes are drawn in the center of each group, flanked above and below by\ parent haplotypes, and variants are sorted to show the transmitted alleles:\
\ parent 1 untransmitted haploytpe \ parent 1 transmitted haplotype\ child haplotype inherited from parent 1\ child haplotype inherited from parent 2\ parent 2 transmitted haplotype\ parent 2 untransmitted haploytpe \\ \
\ Track configuration options include:\
\ Allele coloring options include:\
\ From the subtrack configure menu, there is the option to manually rearrange \ the family order for each trio by dragging haplotypes. \
\ \\ Clicking on a variant takes one to a details page with the standard VCF details, including\ INFO column annotations, the REF and ALT alleles, and the genotypes from all three samples.\
\ \\ The genomes of 2,504 individuals were sequenced using both whole-genome sequencing\ (mean depth = 7.4x) and targeted exome sequencing (mean depth = 65.7x).\ Sequence reads were aligned to the reference genome using alt-aware BWA-MEM\ (Zheng-Bradley et al.).\ Variant discovery and quality control were performed as described in\ Lowy-Gallego et al.
\\ See also:\
\ \ \\ Trio samples were extracted out of both the main 1000 Genomes set, and the\ related samples using the pedigree information from 1000\ Genomes. Variants that were homozygous reference across all three samples were removed.\
\ \\ Trio VCFs are available for download from\ our download server.\
\ \\ Thanks to the\ International Genome Sample\ Resource (IGSR)\ for making these variant calls freely available.\
\ \\ Zheng-Bradley X, Streeter I, Fairley S, Richardson D, Clarke L, Flicek P, 1000 Genomes Project\ Consortium.\ \ Alignment of 1000 Genomes Project reads to reference assembly GRCh38.\ Gigascience. 2017 Jul 1;6(7):1-8.\ PMID: 28531267; PMC: PMC5522380\
\ \\ Fairley S, Lowy-Gallego E, Perry E, Flicek P.\ \ The International Genome Sample Resource (IGSR) collection of open human genomic variation\ resources.\ Nucleic Acids Res. 2019 Oct 4.\ PMID: 31584097\
\ \\ Lowy-Gallego E, Fairley S, Zheng-Bradley X, Ruffier M, Clarke L, Flicek P,\ 1000 Genomes Project Consortium.\ \ Variant calling on the GRCh38 assembly with the data from phase three of the 1000 Genomes Project [version 1; peer review: 2 not approved].\ Wellcome Open Research. 2019 Mar. 11.\
\ \\ 1000 Genomes Project Consortium, Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, Korbel JO,\ Marchini JL, McCarthy S, McVean GA et al.\ \ A global reference for human genetic variation.\ Nature. 2015 Oct 1;526(7571):68-74.\ PMID: 26432245\
\ varRep 0 chromosomes chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX\ compositeTrack on\ geneTrack ncbiRefSeqCurated\ html tgpTrios\ longLabel Thousand Genomes Project Family VCF Trios\ maxWindowToDraw 5000000\ parent tgpArchive\ shortLabel 1000 Genomes Trios\ track tgpTrios\ type vcfPhasedTrio\ vcfDoFilter off\ vcfDoMaf off\ vcfDoQual off\ vcfUseAltSampleNames on\ visibility pack\ tgpPhase3 1000G Ph3 Vars vcfTabix 1000 Genomes Phase 3 Integrated Variant Calls from IGSR: SNVs and Indels 0 100 0 0 0 127 127 127 0 0 23 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX,\ This track shows approximately 73 million single nucleotide variants (SNVs) and\ 5 million short insertions/deletions (indels)\ produced by the\ International\ Genome Sample Resource (IGSR) from sequence data generated by the\ 1000 Genomes Project\ in its Phase 3 sequencing of 2,504 genomes from 16 populations worldwide.
\\ Variants were called on the autosomes (chromosomes 1 through 22) and on the\ Pseudo-Autosomal Regions (PARs) of chromosome X.\ Therefore this track has no annotations on alternate haplotype sequences, fix patches,\ chromosome Y, or the non-PAR portion (the majority) of chromosome X.\
\\ The variant genotypes have been phased\ (i.e., the two alleles of each diploid genotype have been assigned to two\ haplotypes,\ one inherited from each parent).\ This extra information enables a clustering of independent haplotypes\ by local similarity for display.\
\ \\ \ \ \ In "dense" mode, a vertical line is drawn at the position of each\ variant.\ In "pack" mode, since these variants have been phased, the\ display shows a clustering of haplotypes in the viewed range, sorted\ by similarity of alleles weighted by proximity to a central variant.\ The clustering view can highlight local patterns of linkage.
\\ In the clustering display, each sample's phased diploid genotype is split\ into two independent haplotypes.\ Each haplotype is placed in a horizontal row of pixels; when the number of\ haplotypes exceeds the number of vertical pixels for the track, multiple\ haplotypes fall in the same pixel row and pixels are averaged across haplotypes.
\\ Each variant is a vertical bar with white (invisible) representing the reference allele\ and black representing the non-reference allele(s).\ Tick marks are drawn at the top and bottom of each variant's vertical bar\ to make the bar more visible when most alleles are reference alleles.\ The vertical bar for the central variant used in clustering is outlined in purple.\ In order to avoid long compute times, the range of alleles used in clustering\ may be limited; alleles used in clustering have purple tick marks at the\ top and bottom.
\\ The clustering tree is displayed to the left of the main image.\ It does not represent relatedness of individuals; it simply shows the arrangement\ of local haplotypes by similarity. When a rightmost branch is purple, it means\ that all haplotypes in that branch are identical, at least within the range of\ variants used in clustering.\
\ \\ The genomes of 2,504 individuals were sequenced using both whole-genome sequencing\ (mean depth = 7.4x) and targeted exome sequencing (mean depth = 65.7x).\ Sequence reads were aligned to the reference genome using alt-aware BWA-MEM\ (Zheng-Bradley et al.).\ Variant discovery and quality control were performed as described in\ (Lowy-Gallego et al.).\ \ \ See also:\
\ \ \\ VCF files were downloaded from\ EBI\ and are also available for download from\ UCSC.\
\ \\ Thanks to the\ International Genome Sample\ Resource (IGSR)\ for making these variant calls freely available.\
\ \\ Zheng-Bradley X, Streeter I, Fairley S, Richardson D, Clarke L, Flicek P, 1000 Genomes Project\ Consortium.\ \ Alignment of 1000 Genomes Project reads to reference assembly GRCh38.\ Gigascience. 2017 Jul 1;6(7):1-8.\ PMID: 28531267; PMC: PMC5522380\
\ \\ Fairley S, Lowy-Gallego E, Perry E, Flicek P.\ \ The International Genome Sample Resource (IGSR) collection of open human genomic variation\ resources.\ Nucleic Acids Res. 2019 Oct 4.\ PMID: 31584097\
\ \\ Lowy-Gallego E, Fairley S, Zheng-Bradley X, Ruffier M, Clarke L, Flicek P,\ 1000 Genomes Project Consortium.\ \ Variant calling on the GRCh38 assembly with the data from phase three of the 1000 Genomes Project [version 1; peer review: 2 not approved].\ Wellcome Open Research. 2019 Mar. 11.\
\ \\ 1000 Genomes Project Consortium, Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, Korbel JO,\ Marchini JL, McCarthy S, McVean GA et al.\ \ A global reference for human genetic variation.\ Nature. 2015 Oct 1;526(7571):68-74.\ PMID: 26432245\
\ varRep 1 chromosomes chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX\ geneTrack ncbiRefSeqCurated\ html tgpPhase3\ longLabel 1000 Genomes Phase 3 Integrated Variant Calls from IGSR: SNVs and Indels\ maxWindowToDraw 5000000\ parent tgpArchive\ shortLabel 1000G Ph3 Vars\ showHardyWeinberg on\ track tgpPhase3\ type vcfTabix\ visibility hide\ viennaVntr 1KG Vienna ONT VNTR bigBed 9 + 1000 Genomes Vienna ONT VNTR Allele Statistics (VAMOS, 1,019 samples, long-read) 1 100 0 0 0 127 127 127 0 0 0\ This track shows allele statistics for 361,362 variable number tandem repeat (VNTR)\ loci genotyped from Oxford Nanopore long-read whole-genome sequencing of 1,019 samples\ from the\ 1000 Genomes\ ONT Vienna project. VNTR genotyping was performed with\ VAMOS,\ a tool that determines the motif composition of VNTR alleles from long reads.\ This is version 1.1 of the dataset.\
\ \\ Unlike the other STR tracks in this collection which are based on short-read sequencing\ and limited to short tandem repeats (motifs of 1-6 bp), this track is derived from\ long-read sequencing data, which can span much longer repeat regions. The VNTR loci\ in this track have average motif lengths ranging from a few base pairs to over 100 bp,\ and allele lengths up to several kilobases.\
\ \\ For each locus, the track shows the average repeat unit length, the number of unique\ alleles observed, the range and median of repeat unit counts, and the range and median\ of allele lengths in base pairs. The 1000 Genomes Vienna ONT project also produced\ structural variant calls available in the\ Long-Read Structural Variants track.\
\ \\ Items are colored by expected heterozygosity, computed as\ het = 1 − ∑pi2 from allele frequencies\ across the 1,019 samples:
\\ The 1000 Genomes Vienna ONT project sequenced 1,019 samples from the 1000 Genomes\ collection using Oxford Nanopore Technologies long-read sequencing. VNTR genotyping\ was performed using\ VAMOS,\ which determines the motif composition of VNTR alleles by aligning long reads to\ a catalog of known VNTR sites.\ The analysis pipeline is available at\ GitHub.\
\\ At UCSC, the summary statistics file (vamos-summary.tsv) was converted\ to bigBed format using a\ custom Python script.\ Loci with coordinates exceeding chromosome boundaries were excluded.\
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ The data can be accessed from scripts through our\ API, the track name is viennaVntr.\
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed\ file that can be downloaded from\ our download server.\ The file for this track is called viennaVntr.bb.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a\ precompiled binary for your system. Instructions for downloading source code and\ binaries can be found\ here.\ The tool can also be used to obtain features within a given range, e.g.\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/strVar/viennaVntr.bb\ -chrom=chr21 -start=0 -end=100000000 stdout\
\ \\ The original data (multisample VCF and summary statistics) can be downloaded from\ the 1000 Genomes FTP server.\ The VNTR site list used for genotyping is available from\ Zenodo.\
\ \\ Thanks to the 1000 Genomes ONT Vienna consortium and the Marschall Lab at\ Heinrich Heine University Düsseldorf for making this data publicly available.\
\ \\ De Coster W, Condon DE, De Baets G, Tsui A, Saeed F, Harerimana J, Amiraghdam F, Yaari R,\ De Vos L, Mahfouz A et al.\ \ Sequencing and variant calling of 1019 samples from the\ 1000 Genomes Project using Oxford Nanopore Technology.\ bioRxiv. 2024 Dec 23;.\
\ varRep 1 bigDataUrl /gbdb/hg38/strVar/viennaVntr.bb\ dataVersion v1.1\ filter.het 0:1\ filterByRange.het on\ filterLimits.het 0:1\ itemRgb on\ longLabel 1000 Genomes Vienna ONT VNTR Allele Statistics (VAMOS, 1,019 samples, long-read)\ mouseOver Avg motif: $ruLenAvg bp\ AbSplice is a method that predicts aberrant splicing across human tissues, as described in Wagner, \ Çelik et al., 2023. This track displays precomputed AbSplice scores for all possible\ single-nucleotide variants genome-wide. The scores represent the probability that a given variant\ causes aberrant splicing in a given tissue.\ AbSplice scores\ can be computed from VCF files and are based on quantitative tissue-specific splice site annotations\ (SpliceMaps).\ While SpliceMaps can be generated for any tissue of interest from a cohort of RNA-seq samples, this \ track includes 49 tissues available from the \ Genotype-Tissue\ Expression (GTEx) dataset. \
\ \\ The AbSplice score is a probability estimate of how likely aberrant splicing of some sort takes \ place in a given tissue. The authors suggest three cutoffs which are represented by color in the track.\
\ \\ Mouseover on items shows the gene name, maximum score, and tissues that had this score. Clicking on\ any item brings up a table with scores for all 49 GTEX tissues.\
\ \\ The raw data can be explored interactively with the\ Table Browser, or the\ Data Integrator. \ For automated analysis, the data may be queried from our\ REST API.\ Please refer to our\ mailing list archives \ for questions, or our\ Data Access FAQ \ for more information.\
Precomputed AbSplice-DNA scores in all 49 GTEx tissues are available at\ \ Zenodo. \ \
\ Data was converted from the files (AbSplice_DNA_ hg38 _snvs_high_scores.zip) provided by the authors\ at zenodo.org. Files in the\ score_cutoff=0.01 directory were concatenated. To convert the data to bigBed format, scores and\ their tissues were selected from the AbSplice_DNA fields and maximum scores, and then calculated\ using a custom Python script, which can be found in the\ \ makeDoc from our GitHub repository.
\ \\ Thanks to Nils Wagner for helpful comments and suggestions.
\ \\ Wagner N, Çelik MH, Hölzlwimmer FR, Mertes C, Prokisch H, Yépez VA, Gagneur J.\ \ Aberrant splicing prediction across human tissues.\ Nat Genet. 2023 May;55(5):861-870.\ PMID: 37142848\
\ phenDis 1 bigDataUrl /gbdb/hg38/abSplice/AbSplice.bb\ dataVersion Feb 2024\ filter.spliceABscore 0.01\ filterLabel.maxScore Tissues\ filterLabel.spliceABscore Filter by minimum AbSplice score\ filterLimits.spliceABscore 0.01:1\ filterText.maxScore *\ group phenDis\ html abSplice\ itemRgb on\ longLabel Aberrant Splicing Prediction Scores\ mouseOver change: $name\ This supertrack is a collection of Affymetrix tracks showing the location of the consensus and\ exemplar sequences used for the selection of probes on the Affymetrix chips.\
\\ Thanks to\ Affymetrix for the data underlying these tracks.\
\ expression 1 cartVersion 2\ group expression\ html ../affyArchive\ longLabel Affymetrix Archive\ shortLabel Affy Archive\ superTrack on\ track affyArchive\ type psl .\ visibility hide\ affyGnf1h Affy GNF1H psl . Alignments of Affymetrix Consensus/Exemplars from GNF1H 3 100 0 0 0 127 127 127 0 0 0This track shows the location of the sequences used for the selection of\ probes on the Affymetrix GNF1H chips. This contains 11406 predicted genes that do not overlap with\ the Affy U133A chip.
\ \The sequences were mapped to the genome using blat followed by pslReps with the\ parameters:
-minCover=0.3 -minAli=0.95 -nearTop=0.005\ \
Thanks to the Genomics\ Institute of the Novartis Research Foundation (GNF) for the data underlying this track.
\ \\ Su AI, Wiltshire T, Batalov S, Lapp H, Ching KA, Block D, Zhang J, Soden R, Hayakawa M, Kreiman G\ et al.\ \ A gene atlas of the mouse and human protein-encoding transcriptomes.\ Proc Natl Acad Sci U S A. 2004 Apr 20;101(16):6062-7.\ PMID: 15075390; PMC: PMC395923\
\ expression 1 group expression\ longLabel Alignments of Affymetrix Consensus/Exemplars from GNF1H\ parent affyArchive\ shortLabel Affy GNF1H\ track affyGnf1h\ type psl .\ visibility pack\ affyU133 Affy U133 psl . Alignments of Affymetrix Consensus/Exemplars from HG-U133 3 100 0 0 0 127 127 127 0 0 0\ This track shows the location of the consensus and exemplar sequences used \ for the selection of probes on the Affymetrix HG-U133A and HG-U133B chips.
\ \\ Consensus and exemplar sequences were downloaded from the\ Affymetrix Product Support\ and mapped to the genome using blat followed by pslReps with the \ parameters:
-minCover=0.5 -minAli=0.97 -nearTop=0.005\\ \
\ Thanks to Affymetrix for the data underlying this track.
\ expression 1 group expression\ longLabel Alignments of Affymetrix Consensus/Exemplars from HG-U133\ parent affyArchive\ shortLabel Affy U133\ track affyU133\ type psl .\ visibility pack\ affyU95 Affy U95 psl . Alignments of Affymetrix Consensus/Exemplars from HG-U95 3 100 0 0 0 127 127 127 0 0 0\ This track shows the location of the consensus and exemplar sequences used \ for the selection of probes on the Affymetrix HG-U95Av2 chip. For this chip, \ probes are predominantly designed from consensus sequences.
\ \\ Consensus and exemplar sequences were downloaded from the\ Affymetrix Product Support\ and mapped to the genome using blat followed by pslReps with the \ parameters:
-minCover=0.3 -minAli=0.95 -nearTop=0.005\\ \
\ Thanks to Affymetrix for the data underlying this track.
\ expression 1 group expression\ longLabel Alignments of Affymetrix Consensus/Exemplars from HG-U95\ parent affyArchive\ shortLabel Affy U95\ track affyU95\ type psl .\ visibility pack\ lrSvAll All LR SVs merged bigBed 9 + All long-read SVs merged across subtracks by exact position, with per-database AC 3 100 0 0 0 127 127 127 0 0 0\ This track combines the structural-variant (SV) callsets from the individual\ subtracks of the Long-read SVs supertrack into a\ single, position-merged overview. Each item is an SV locus seen in one or more\ of the contributing long-read databases. For every merged locus the track\ records which databases report it, the summed allele count across those\ databases, and the range of allele frequencies observed, making it useful for\ quickly seeing how widely an SV has been reported across cohorts.\
\\ This is a summary view. For cohort-specific genotypes, per-population allele\ frequencies, and dataset-specific annotations, use the individual subtracks of\ the supertrack. The merge includes the released long-read callsets only;\ preliminary or unpublished subtracks (e.g. the Kim PD brain, 1000 Genomes\ linear, and HPRC Jasmine sets) are not part of this merged track.\
\ \\ Items are colored by SV type, matching the individual subtracks:\
\ The mouseover shows the variant name, SV type, reference and insertion lengths,\ the list of contributing source databases, the allele-frequency range across\ those databases, and the total allele count. Filters are available for the\ source database, SV type, SV length, insertion\ length, total allele count, minimum and maximum allele\ frequency, and the number of source databases reporting each locus.\ The detail page lists the per-database allele counts.\
\ \\ The merged track is built by the lrSvMergeAll.py script, which reads\ the bigBed of each contributing subtrack (configured in\ databases.tsv) and groups records that share an identical\ (chromosome, start, end) position and SV type. For each merged locus\ the script records the set of contributing databases (sources), the\ number of those databases (sourceCount), the sum of their allele counts\ (AC), and the minimum and maximum allele frequency across databases\ that report one (minAF, maxAF). The per-database allele counts\ are carried as additional columns.\
\\ The step-by-step build commands are recorded in the UCSC makeDoc for this track\ collection:\ \ doc/hg38/lrSv.txt. The merge script and autoSql schema live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSvAll.ra.\
\ \\ The data can be explored interactively with the\ Table Browser or the\ Data Integrator, and accessed\ programmatically through our API,\ track=lrSvAll.\
\\ The bigBed is available from\ our\ download server as lrSvAll.bb. Example:\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/lrSvAll.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\ \\ This merged view is derived entirely from the contributing long-read SV\ callsets; please see the individual subtrack description pages for the data\ producers and citations for each cohort.\
\ varRep 1 bigDataUrl /gbdb/hg38/lrSv/lrSvAll.bb\ filter.AC 0:30000\ filter.insLen 0:600000\ filter.maxAF 0:1\ filter.minAF 0:1\ filter.sourceCount 1:14\ filter.svLen 0:30000000\ filterByRange.AC on\ filterByRange.insLen on\ filterByRange.maxAF on\ filterByRange.minAF on\ filterByRange.sourceCount on\ filterByRange.svLen on\ filterLabel.AC Total AC (across DBs)\ filterLabel.insLen Insertion Length (bp)\ filterLabel.maxAF Max Allele Frequency (across DBs)\ filterLabel.minAF Min Allele Frequency (across DBs)\ filterLabel.sourceCount Number of Source Databases\ filterLabel.sources Source Database\ filterLabel.svLen SV Length (bp)\ filterLabel.svType SV Type\ filterLimits.maxAF 0:1\ filterLimits.minAF 0:1\ filterType.sources multipleListOr\ filterType.svType multipleListOr\ filterValues.sources CoLoRSdb|CoLoRSdb 1427 (PacBio),1000G-ONT-Vienna|1KG ONT Vienna 1019,1000G-ONT|1KG ONT 100 (Gustafson),AoU1K|All of Us 1027 (PacBio),Han945|Han Chinese 945,TommoJapan|ToMMo 333 (Japanese),GA4K|GA4K 502 (rare disease),deCODE|deCODE 3622 (Icelandic),HPRCv2.1|HPRC v2.1 233,HGSVC2|HGSVC2 32,HGSVC3|HGSVC3 65,ArabUAE53|Arab APR 53,China58|CPC 58 (Chinese),Svatalog101|SVatalog 101\ filterValues.svType DEL,INS,DUP,INV,CPX,MIXED,INSDEL,TRA\ itemRgb on\ longLabel All long-read SVs merged across subtracks by exact position, with per-database AC\ mouseOver Var: $name ($svType)\ This track shows AlphaMissense predictions for all possible single amino acid substitutions in \ the human proteome.\
\\ AlphaMissense is a deep learning method for predicting the pathogenicity of missense variants\ in human proteins. It classifies 32% of all missense variants as likely pathogenic and 57% \ as likely benign using a cutoff yielding 90% precision on the ClinVar dataset.\
\ \ \There are four lettered subtracks, one for every nucleotide, showing\ scores for mutation from the reference to that\ nucleotide. All subtracks show the AlphaMissense score on mouseover. Across the exome, \ there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing\ the reference allele, e.g. A to A, is always set to zero, "0.0". AlphaMissense only\ takes into account amino acid changes, so a nucleotide change that results in no\ amino acid change (synonymous) is not scored. These are shown in the tracks\ with score "0.0". \ \
\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and no score will be shown on the mouseover tooltip.\
\ \Track colors
\\ This track is colored according to the am_class column in the AlphaMissense_hg38.tsv file.\ \
| Range | \Classification | \
|---|---|
| ≥ .564 | \Likely Pathogenic | \
| .565 - .340 | \Likely Neutral | \
| ≤ .340 | \Likely Benign | \
\ AlphaMissense scores are available at the \ \ AlphaMissense cloud storage site. \ The site provides precomputed AlphaMissense scores for all possible human missense variants \ to facilitate the identification of pathogenic variants among the large number of \ rare variants discovered in sequencing studies.\
\ \\ The AlphaMissense data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\
\ \\ For automated download and analysis, the genome annotation is stored at UCSC in\ bigWig format that can be downloaded from\ our download server.\ The files for this track are called a.bw, c.bw, g.bw, t.bw. Individual\ regions or the whole genome annotation can be obtained using our tool bigWigToWig\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ For example, to extract only annotations in a given region, you could use the following command:\
\ \\ bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/alphaMissense/a.bw stdout\
\ \\ Data were converted from the files provided on\ the AlphaMissense Downloads website. As with all other tracks,\ a full log of all commands used for the conversion is available in our \ source\ repository, for hg19\ and hg38.\ The release used for each assembly is shown on the track description page.\
\ \\ Thanks to \
\ \\ Cheng J, Novati G, Pan J, Bycroft C, Žemgulytė A, Applebaum T, Pritzel A, Wong LH,\ Zielinski M, Sargeant T et al.\ Accurate proteome-wide missense variant effect prediction with\ AlphaMissense.\ Science. 2023 Sep 22;381(6664):eadg7492.\ PMID: 37733863\
\ \ phenDis 0 color 100,130,160\ compositeTrack on\ group phenDis\ html alphaMissense.html\ longLabel AlphaMissense Score for all possible single-basepair mutations (zoom in for scores)\ maxWindowToDraw 10000000\ mouseOverFunction noAverage\ shortLabel AlphaMissense\ track alphaMissense\ type bigWig\ visibility hide\ altSeqLiftOverPsl Alt Haplotypes psl Reference Assembly Alternate Haplotype Sequence Alignments 3 100 0 0 100 127 127 177 0 0 0\ This track shows alignments of alternate locus (also known as "alternate haplotype")\ reference sequences to main chromosome sequences in the reference genome assembly.\ Some loci in the genome are highly variable, with sets of variants that tend\ to segregate into distinct haplotypes.\ Only one haplotype can be included in a reference assembly chromosome sequence.\ Instead of providing a separate complete chromosome sequence for each haplotype,\ which could cause confusion with divergent chromosome coordinates and\ ambiguity about which sequence is the official reference, the\ Genome Reference Consortium\ (GRC) adds alternate locus sequences, ranging from tens of thousands of bases\ up to low millions of bases in size, to represent the distinct haplotypes. \
\ \\ This track follows the display conventions for\ \ PSL alignment tracks.\ Mismatching bases are highlighted in red.\ Several types of alignment gap may also be colored;\ for more information, see\ \ Alignment Insertion/Deletion Display Options.\ \
\ \\ The alignments were provided by NCBI as GFF files and translated into the PSL\ representation for browser display by UCSC.\
\ map 1 baseColorDefault diffBases\ baseColorUseSequence db\ color 0,0,100\ group map\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Reference Assembly Alternate Haplotype Sequence Alignments\ parent patchesPsl\ pennantIcon p14 black https://genome-blog.gi.ucsc.edu/blog/patches/ "Includes annotations on GRCh38.p14 patch sequences"\ shortLabel Alt Haplotypes\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ track altSeqLiftOverPsl\ type psl\ visibility pack\ ancient Ancient Hominids bed 12 Ancient Hominid DNA Variants 0 100 0 0 0 127 127 127 0 0 0\ This container track contains genome variants from ancient hominids, from DNA\ samples extracted from Denisovan and Neanderthal, provided by the database Arcseqhub (Lian et al, Gen Biol 2025).\
\ \\ Variants that differ from human are highlighted. Click onto a variant to see more details.\
\ \\ The data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be\ accessed from scripts through our API, the track name is\ "denisovan" and "neanderthal".
\ \\ For automated download and analysis, the genome annotation is stored in a tabix-indexed VCF file that\ can be downloaded from\ our download server.\ The files for this track are called denisovan.hg38.filt.vcf.gz and neanderthal.hg38.filt.vcf.gz. \ Various command line tools exist for working with VCF files. Users without command line experience can use the Galaxy website, by exporting the data directly from our table browser to Galaxy.\
\ \\ Liang et al (see below) realigned the original sequencing reads to the hg38 and\ T2T CHM13 assemblies. UCSC removed positions from the VCF without an alternate\ allele to show only variants that are present in the ancient genomes and loaded the VCFs.\
\ \\ We thank the Arcseqhub authors for making the data available.\
\ \\ Liang SA, Ren T, Zhang J, He J, Wang X, Jiang X, He Y, McCoy RC, Fu Q, Akey JM et al.\ \ A refined analysis of Neanderthal-introgressed sequences in modern humans with a complete reference\ genome.\ Genome Biol. 2025 Feb 17;26(1):32.\ PMID: 39962554; PMC: PMC11834205\
\ \ varRep 1 group varRep\ longLabel Ancient Hominid DNA Variants\ shortLabel Ancient Hominids\ superTrack on\ track ancient\ type bed 12\ visibility hide\ AorticSmoothMuscleCellResponseToIL1b06hrBiolRep2LK59_CNhs13378_ctss_rev AorticSmsToIL1b_06hrBr2- bigWig Aortic smooth muscle cell response to IL1b, 06hr, biol_rep2 (LK59)_CNhs13378_12759-136B5_reverse 0 100 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12759-136B5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20IL1b%2c%2006hr%2c%20biol_rep2%20%28LK59%29.CNhs13378.12759-136B5.hg38.ctss.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to IL1b, 06hr, biol_rep2 (LK59)_CNhs13378_12759-136B5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12759-136B5 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToIL1b_06hrBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_IL1b strand=reverse\ track AorticSmoothMuscleCellResponseToIL1b06hrBiolRep2LK59_CNhs13378_ctss_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12759-136B5\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToIL1b06hrBiolRep2LK59_CNhs13378_tpm_rev AorticSmsToIL1b_06hrBr2- bigWig Aortic smooth muscle cell response to IL1b, 06hr, biol_rep2 (LK59)_CNhs13378_12759-136B5_reverse 1 100 0 0 255 127 127 255 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12759-136B5 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20IL1b%2c%2006hr%2c%20biol_rep2%20%28LK59%29.CNhs13378.12759-136B5.hg38.tpm.rev.bw\ color 0,0,255\ longLabel Aortic smooth muscle cell response to IL1b, 06hr, biol_rep2 (LK59)_CNhs13378_12759-136B5_reverse\ maxHeightPixels 100:8:8\ metadata ontology_id=12759-136B5 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToIL1b_06hrBr2-\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_IL1b strand=reverse\ track AorticSmoothMuscleCellResponseToIL1b06hrBiolRep2LK59_CNhs13378_tpm_rev\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12759-136B5\ urlLabel FANTOM5 Details:\ aou1kSv AoU 1027 SVs bigBed 9 + Structural Variants from 1,027 AoU Individuals (PacBio HiFi Long-read) 0 100 0 0 0 127 127 127 0 0 0\ This track shows structural variants (SVs) identified by PacBio HiFi long-read\ sequencing of 1,027 individuals from the All of Us (AoU) Research Program.\ Participants self-identified as Black or African American and were sequenced\ to ~8x coverage. The track contains 540,155 SVs (443,630 insertions and\ 96,525 deletions) on autosomes, after removing byte-identical duplicate records\ from the 541,049-row release.\
\\ SVs are annotated with population-specific allele frequencies across five\ ancestry groups (African, Admixed American, East Asian, European, South Asian),\ gene intersections from curated disease gene lists (OMIM, ACMG, cancer genes),\ regulatory element overlaps, and associations with eQTLs, GWAS loci, and\ clinical phenotypes from the AoU electronic health records.\
\ \\ Items are colored by SV type:\
\ Filters are available for SV type, SV length, and population-specific allele\ frequencies. For insertions, the item is placed at the insertion site with a\ width of 1 bp; for deletions, the item spans the deleted region.\
\\ The detail page shows the following annotations when available:\
\ Garimella et al. 2025 performed PacBio HiFi long-read sequencing on 1,027\ All of Us participants self-identifying as Black or African American, to\ ~8x per-sample coverage at HudsonAlpha Discovery. SVs (≥50 bp) were\ called per sample with an ensemble of three methods: two alignment-based\ callers, \ PBSV v2.6.0 (with Tandem Repeat Finder context) and\ Sniffles2\ v2.0.6, plus the assembly-based PAV v1.2.1 (hifiasm haplotype-resolved contigs aligned\ to GRCh38 with minimap2 -x asm20). Per-caller VCFs were normalized,\ merged within and across samples and filtered into stringent and lenient\ tiers, and the callset was re-genotyped across the cohort to produce the\ final release: 541,049 autosomal SVs (444,524 insertions, 96,525 deletions)\ with per-ancestry allele frequencies (AFR, AMR, EAS, EUR, SAS) and gene,\ regulatory, eQTL, GWAS and EHR-phenotype annotations.\
\\ This track was built from the supplementary media-2 table of the AoU\ long-read sequencing preprint\ (\ doi:10.1101/2025.10.02.25336942). Access to the underlying AoU\ long-read data requires registration through the\ All of Us\ Research Hub.\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ This track was built from supplementary data (media-2) of the AoU long-read\ sequencing preprint. Access to the full AoU dataset requires registration\ through the All of\ Us Research Hub.\
\ \\ Thanks to Garimella et al. and the All of Us Research Program for making their\ structural variant annotations publicly available.\
\ \\ Garimella KV, Li Q, Wertz J, Lee SK, Cunial F, Huang Y, Mostovoy Y, Lorig-Roach R, English A, Su H\ et al.\ \ Population-scale Long-read Sequencing in the All of Us Research Program.\ medRxiv. 2025 Oct 5;.\ PMID: 41256123; PMC: PMC12622093\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/aou1k.bb\ filter.AC 0:2054\ filter.insLen 0:9998\ filter.svLen 0:9905\ filterByRange.AC on\ filterByRange.afAfr on\ filterByRange.afEas on\ filterByRange.afEur on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Allele Count (approx)\ filterLabel.afAfr AF African\ filterLabel.afEas AF East Asian\ filterLabel.afEur AF European\ filterLabel.insLen Insertion Length\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.afAfr 0:1\ filterLimits.afEas 0:1\ filterLimits.afEur 0:1\ filterType.svType multipleListOr\ filterValues.svType DEL,INS\ itemRgb on\ longLabel Structural Variants from 1,027 AoU Individuals (PacBio HiFi Long-read)\ mouseOver Var: $name ($svType)\ This track displays structural variants (SVs), at least 50 bp long\ (deletions, insertions, and complex substitutions), from the Arab Pangenome\ Reference (APR), a pangenome graph built from 53 UAE-resident Arab\ individuals drawn from eight countries (UAE, Saudi Arabia, Oman, Jordan,\ Egypt, Morocco, Syria, Yemen). Each bubble in the graph that contains an\ SV-sized alternative allele is shown as a single variant site, with allele\ counts aggregated across the 53 samples (the GRCh38 reference haplotype,\ present as an extra sample column in the source VCF, is excluded from the\ aggregation).
\ \\ The APR pangenome was built on the T2T-CHM13v2 reference. Variants are\ shown natively on the hs1 browser and lifted to hg38 using\ the UCSC hs1ToHg38.over.chain.gz chain; variants that do not lift\ cleanly (often in T2T-added euchromatic sequence) are omitted from the\ hg38 version of the track.
\ \Items are colored by SV type:
\Each item spans from the start of REF to its end on the reference.\ The name field is the graph snarl ID (e.g. <951452<1012008),\ which identifies the variant site in the APR pangenome graph.
\ \\ The source VCF is multi-allelic: a single graph snarl appears as one row\ with a comma-separated ALT list. For this track, each ALT is classified\ individually using the 50 bp threshold, and the row is emitted as a single\ bed item with:
\Rows whose alts are all smaller than 50 bp are not shown.
\ \\ Nassir et al. 2025 built the Arab Pangenome Reference (APR) from 53\ UAE-resident Arab individuals drawn from eight countries, sequenced with\ ~35x PacBio HiFi on Sequel IIe/Revio (30-h movies), ~54x Oxford Nanopore\ ultralong reads on R10.4.1 PromethION flow cells (96-h runs), and ~65x\ Hi-C (Illumina NovaSeq 6000). Haplotype-phased de novo assemblies were\ produced with hifiasm v0.19.5 (primary) and Verkko v1.3.1 (for\ comparison), with a median N50 of 124 Mb. The pangenome graph was built\ with Minigraph-Cactus seeded on T2T-CHM13v2 and augmented with GRCh38,\ and SVs were extracted by graph deconstruction. The released decomposed\ VCF (apr_review_v1_2902_chm13.vcf.gz) contains ~21 million\ variants on CHM13v2 contigs; after filtering to alt alleles with ≥50 bp\ length difference and collapsing the alts of each snarl into a single\ site, the APR SV track is obtained. Variants are shown natively on hs1\ and lifted to hg38 with the UCSC hs1ToHg38.over.chain.gz chain\ (variants not lifting cleanly are omitted from the hg38 version).
\ \\ The source APR VCF was downloaded from the Mohammed Bin Rashid\ University SharePoint page,\ \ mbru.ac.ae/the-arab-pangenome-reference; the accompanying project\ source code is at\ \ github.com/muddinmbru/arab_pangenome_reference.
\ \\ The step-by-step build commands (download, graph-VCF conversion, liftOver,\ bigBed build) are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \The data can be explored interactively with the\ Table Browser or\ Data Integrator, and accessed from\ scripts via our API\ (track=aprSv).
\ \For automated download, the bigBed files are at\ \ http://hgdownload.soe.ucsc.edu/gbdb/hs1/lrSv/apr.bb (native) and\ \ http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/apr.bb (lifted).
\ \\ The original APR pangenome VCF and assemblies can be downloaded from\ \ https://www.mbru.ac.ae/the-arab-pangenome-reference/,\ and the project source code is at\ \ https://github.com/muddinmbru/arab_pangenome_reference.
\ \Thanks to the Arab Pangenome Reference team at Mohammed Bin Rashid\ University (Dubai), led by Mohammed Uddin, for producing and releasing\ the pangenome and its variant calls.
\ \\ Nassir N, Almarri MA, Kumail M, Mohamed N, Balan B, Hanif S, AlObathani M, Jamalalail B, Elsokary H,\ Kondaramage D et al.\ \ A draft UAE-based Arab pangenome reference.\ Nat Commun. 2025 Jul 24;16(1):6747.\ PMID: 40707445; PMC: PMC12290100\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/apr.bb\ filter.AC 0:107\ filter.insLen 0:584016\ filter.svLen 0:99885\ filterByRange.AC on\ filterByRange.alleleFreq on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Allele Count\ filterLabel.alleleFreq Allele Frequency\ filterLabel.insLen Insertion Length\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.alleleFreq 0:1\ filterType.svType multipleListOr\ filterValues.svType INS,DEL,CPX,MIXED\ itemRgb on\ longLabel Structural Variants from the Arab Pangenome Reference (53 UAE-resident Arab samples)\ mouseOver Var: $name ($svType)\ This container track contains genome variants from ancient hominids, from DNA\ samples extracted from Denisovan and Neanderthal, provided by the database Arcseqhub (Lian et al, Gen Biol 2025).\
\ \\ Variants that differ from human are highlighted. Click onto a variant to see more details.\
\ \\ The data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be\ accessed from scripts through our API, the track name is\ "denisovan" and "neanderthal".
\ \\ For automated download and analysis, the genome annotation is stored in a tabix-indexed VCF file that\ can be downloaded from\ our download server.\ The files for this track are called denisovan.hg38.filt.vcf.gz and neanderthal.hg38.filt.vcf.gz. \ Various command line tools exist for working with VCF files. Users without command line experience can use the Galaxy website, by exporting the data directly from our table browser to Galaxy.\
\ \\ Liang et al (see below) realigned the original sequencing reads to the hg38 and\ T2T CHM13 assemblies. UCSC removed positions from the VCF without an alternate\ allele to show only variants that are present in the ancient genomes and loaded the VCFs.\
\ \\ We thank the Arcseqhub authors for making the data available.\
\ \\ Liang SA, Ren T, Zhang J, He J, Wang X, Jiang X, He Y, McCoy RC, Fu Q, Akey JM et al.\ \ A refined analysis of Neanderthal-introgressed sequences in modern humans with a complete reference\ genome.\ Genome Biol. 2025 Feb 17;26(1):32.\ PMID: 39962554; PMC: PMC11834205\
\ \ varRep 1 bigDataUrl /gbdb/hg38/ancient/denisovan.hg38.filt.vcf.gz\ hapClusterEnabled true\ html ancient\ longLabel Ancient Hominids: Arcseqhub Denisovan VCF Variants\ parent ancient on\ shortLabel Arcseqhub Denisovan\ track denisovan\ type vcfTabix\ visibility pack\ neanderthal Arcseqhub Neanderthal vcfTabix Ancient Hominids: Arcseqhub Neanderthal VCF Variants 3 100 0 0 0 127 127 127 0 0 0\ This container track contains genome variants from ancient hominids, from DNA\ samples extracted from Denisovan and Neanderthal, provided by the database Arcseqhub (Lian et al, Gen Biol 2025).\
\ \\ Variants that differ from human are highlighted. Click onto a variant to see more details.\
\ \\ The data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be\ accessed from scripts through our API, the track name is\ "denisovan" and "neanderthal".
\ \\ For automated download and analysis, the genome annotation is stored in a tabix-indexed VCF file that\ can be downloaded from\ our download server.\ The files for this track are called denisovan.hg38.filt.vcf.gz and neanderthal.hg38.filt.vcf.gz. \ Various command line tools exist for working with VCF files. Users without command line experience can use the Galaxy website, by exporting the data directly from our table browser to Galaxy.\
\ \\ Liang et al (see below) realigned the original sequencing reads to the hg38 and\ T2T CHM13 assemblies. UCSC removed positions from the VCF without an alternate\ allele to show only variants that are present in the ancient genomes and loaded the VCFs.\
\ \\ We thank the Arcseqhub authors for making the data available.\
\ \\ Liang SA, Ren T, Zhang J, He J, Wang X, Jiang X, He Y, McCoy RC, Fu Q, Akey JM et al.\ \ A refined analysis of Neanderthal-introgressed sequences in modern humans with a complete reference\ genome.\ Genome Biol. 2025 Feb 17;26(1):32.\ PMID: 39962554; PMC: PMC11834205\
\ \ varRep 1 bigDataUrl /gbdb/hg38/ancient/neanderthal.hg38.filt.vcf.gz\ hapClusterEnabled true\ html ancient\ longLabel Ancient Hominids: Arcseqhub Neanderthal VCF Variants\ parent ancient on\ shortLabel Arcseqhub Neanderthal\ track neanderthal\ type vcfTabix\ visibility pack\ genotypeArrays Array Probesets bigBed 4 Microarray Probesets and OGM sites 0 100 0 0 0 127 127 127 0 0 0\ The arrays listed in this track are probes from the\ Agilent Catalog Oligonucleotide Microarrays.\
\Please note that more microarray tracks are available on the hg19 genome assembly. \ To view those tracks, please \ click this link for hg19 microarrays.\ Microarrays that are not listed can be added as Custom Tracks with data from the companies.\
\ \\ Agilent's oligonucleotide CGH (Comparative Genomic Hybridization) platform enables the\ study of genome-wide DNA copy number changes at a high resolution. The CGH probes on Agilent\ CGH microarrays are 60-mer oligonucleotides synthesized in situ using Agilent's inkjet\ SurePrint technology. The probes represented on the Agilent CGH microarrays have been\ selected using algorithms developed specifically for the CGH application, assuring optimal\ performance of these probes in detecting DNA copy number changes.\
\ \\ With the Infinium MethylationEPIC BeadChip Kit, researchers can interrogate over 850,000\ methylation sites quantitatively across the genome at single-nucleotide resolution. Multiple\ samples, including FFPE, can be analyzed in parallel to deliver high-throughput power while\ minimizing the cost per sample. These tracks show positions being measured on the Illumina 450k and\ 850k (EPIC) microarray tracks, not the probe locations themselves. Contact us\ or Illumina if you need the probe locations directly. More information about\ the arrays can be found on the\ Infinium MethylationEPIC Kit website.\
\ Note: The 450k track on hg38 contains 128,989 regions representing the target regions, not the probes\ themselves.
\ \\ The Infinium CytoSNP-850K v1.2 BeadChip provides comprehensive coverage of\ cytogenetically relevant genes on a proven platform, helping researchers find valuable information\ that may be missed by other technologies. It contains approximately 850,000 empirically selected\ single nucleotide polymorphisms (SNPs) spanning the entire genome with enriched coverage for 3,262\ genes of known cytogenetics relevance in both constitutional and cancer applications. \
\ \\ The CytoScan HD Array, which is included in the\ CytoScan HD Suite, provides the broadest coverage and highest performance for\ detecting chromosomal aberrations. CytoScan HD Suite has greater than 99% sensitivity and can\ reliably detect 25-50kb copy number changes across the genome at high specificity with\ single-nucleotide polymorphism (SNP) allelic corroboration. With more than 2.6 million copy number\ markers, CytoScan HD Suite covers all OMIM and RefSeq genes.\
\ \\ Bionano Laboratories provides access to Optical Genome Mapping (OGM) data for projects across a variety of\ applications for researchers, clinicians, and pharmaceutical companies.
\This track shows the CTTAAG sites used by the \ Bionano Optical Genome Mapping system,\ an assay to detect structural variants.\
\ \\ Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \ \\ The Agilent arrays were downloaded from their \ Agilent SureDesign website tool on March 2022.
\\ The Illumina 450k and 850k (EPIC) tracks were created using a few columns from the\ Infinium MethylationEPIC v1.0 B5 Manifest File (CSV Format)\ and was then converted into a bigBed.
\\ The Illumina CytoSNP-850K track was created by downloading the\ CytoSNP-850K v1.2 Manifest File (CSV Format) (GRCh38) file and then converted\ into a bigBed file.\
\\ The Affymetrix Cytoscan HD GeneChip Array track was created by converting the \ CytoScanHD_Accel_Array.na36.bed.zip\ into a bigBed file.\
\\ The Bionano track was created by receiving the BED files from\ \ apang@bionano.\ com\ \ and converted to bigBed files using the bedToBigBed tool.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API \ or downloaded from our \ Downloads site. Please refer to our\ \ mailing list archives for questions, or our\ \ Data Access FAQ for more information.\
\ \\ Thanks to the Agilent and Illumina support teams for sharing the data and the UCSC Genome Browser\ engineers for configuring the data.
\\ Thanks to Andy Pang from Bionano Genomics for providing the BED data file.
\ varRep 1 compositeTrack on\ group varRep\ longLabel Microarray Probesets and OGM sites\ shortLabel Array Probesets\ track genotypeArrays\ type bigBed 4\ visibility hide\ gnomADPextArtery_Aorta Artery-Aorta bigWig 0 1 gnomAD pext Artery-Aorta 0 100 255 85 85 255 170 170 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Artery_Aorta.bw\ color 255,85,85\ longLabel gnomAD pext Artery-Aorta\ parent gnomadPext off\ shortLabel Artery-Aorta\ track gnomADPextArtery_Aorta\ visibility hide\ gnomADPextArtery_Coronary Artery-Coronary bigWig 0 1 gnomAD pext Artery-Coronary 0 100 255 170 153 255 212 204 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Artery_Coronary.bw\ color 255,170,153\ longLabel gnomAD pext Artery-Coronary\ parent gnomadPext off\ shortLabel Artery-Coronary\ track gnomADPextArtery_Coronary\ visibility hide\ gnomADPextArtery_Tibial Artery-Tibial bigWig 0 1 gnomAD pext Artery-Tibial 0 100 255 0 0 255 127 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Artery_Tibial.bw\ color 255,0,0\ longLabel gnomAD pext Artery-Tibial\ parent gnomadPext off\ shortLabel Artery-Tibial\ track gnomADPextArtery_Tibial\ visibility hide\ gold Assembly bed 3 + Assembly from Fragments 0 100 150 100 30 230 170 40 0 0 0\ This track shows the contigs used to construct the GRCh38 (hg38) genome assembly, as defined in the\ AGP file delivered with the sequence. \ For information on the AGP file format, see the NCBI \ AGP Specification. The NCBI website also provides an \ overview of genome assembly procedures, as well as \ specific information about the hg38 assembly.\
\\ In dense mode, this track depicts the contigs that make up the \ currently viewed scaffold. \ Contig boundaries are distinguished by the use of alternating gold and brown \ coloration. Where gaps\ exist between contigs, spaces are shown between the gold and brown\ blocks. The relative order and orientation of the contigs\ within a scaffold is always known; therefore, a line is drawn in the graphical\ display to bridge the blocks.
\\ Component types found in this track (with counts of that type in parenthesis):\
\ In addition to the standard nucleotide codes, the raw sequence files from NCBI also include\ IUPAC ambiguity codes for bases that could not be positively identified as A, C, G or T (see\ Wikipedia's IUPAC notation article for more information). As part of the UCSC\ assembly creation process, all IUPAC ambiguity characters are converted to Ns. The FASTA files\ available for download from UCSC reflect this. The raw data files containing the original IUPAC\ characters can be downloaded from the NCBI\ FTP site.\
\ \\ The following table lists the counts by chromosome of the various IUPAC ambiguity characters\ in the original NCBI data files:\
\ \\
| chromosome | \ | |||||||||||||||||
| \ | \ | 1 | \2 | \3 | \6 | \7 | \9 | \10 | \12 | \13 | \16 | \17 | \21 | \22 | \X | \Y | \\ | Total | \
| code | \ | |||||||||||||||||
| B | \\ | \ | \ | 1 | \\ | \ | \ | 1 | \\ | \ | \ | \ | \ | \ | \ | \ | \ | 2 | \
| K | \\ | \ | 1 | \\ | \ | \ | \ | 4 | \\ | 1 | \\ | 2 | \\ | \ | \ | \ | \ | 8 | \
| M | \\ | 1 | \1 | \\ | \ | \ | \ | 3 | \1 | \\ | \ | \ | 2 | \\ | \ | \ | \ | 8 | \
| R | \\ | 1 | \1 | \1 | \\ | 1 | \1 | \13 | \\ | \ | 1 | \3 | \1 | \2 | \1 | \1 | \\ | 27 | \
| S | \\ | \ | \ | \ | \ | 1 | \\ | 1 | \\ | \ | \ | 1 | \\ | \ | 1 | \1 | \\ | 5 | \
| W | \\ | \ | 2 | \2 | \\ | \ | \ | 6 | \\ | \ | \ | 1 | \\ | 1 | \1 | \1 | \\ | 14 | \
| Y | \\ | \ | 4 | \3 | \1 | \2 | \2 | \8 | \2 | \2 | \\ | 5 | \\ | 2 | \2 | \2 | \\ | 35 | \
| \ | ||||||||||||||||||
| Total | \\ | 2 | \9 | \7 | \1 | \4 | \3 | \36 | \3 | \3 | \1 | \12 | \3 | \5 | \5 | \5 | \\ | 99 | \
\ This is a container track for data related to the genome assembly. \ It contains tracks about the assembly identifiers, certain clones, and STS markers. \ Click into any of the sub-tracks to see information\ details on the specific annotations.
\ map 0 cartVersion 4\ group map\ longLabel Assembly identifiers, clones, and markers\ shortLabel Assembly Tracks\ superTrack on\ track assemblyContainer\ augustusGene AUGUSTUS genePred AUGUSTUS ab initio gene predictions v3.1 3 100 180 0 0 217 127 127 0 0 0\ This track shows ab initio predictions from the program\ AUGUSTUS (version 3.1).\ The predictions are based on the genome sequence alone.\
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ Statistical signal models were built for splice sites, branch-point\ patterns, translation start sites, and the poly-A signal.\ Furthermore, models were built for the sequence content of\ protein-coding and non-coding regions as well as for the length distributions\ of different exon and intron types. Detailed descriptions of most of these different models\ can be found in Mario Stanke's\ dissertation.\ This track shows the most likely gene structure according to a\ Semi-Markov Conditional Random Field model.\ Alternative splicing transcripts were obtained with\ a sampling algorithm (--alternatives-from-sampling=true --sample=100 --minexonintronprob=0.2\ --minmeanexonintronprob=0.5 --maxtracks=3 --temperature=2).\
\ \\ The different models used by Augustus were trained on a number of different species-specific\ gene sets, which included 1000-2000 training gene structures. The --species option allows\ one to choose the species used for training the models. Different training species were used\ for the --species option when generating these predictions for different groups of\ assemblies.\
| Assembly Group | \ \ \Training Species | \ \
| Fish | \ \ \zebrafish\ \ |
| Birds | \ \ \chicken\ \ |
| Human and all other vertebrates | \ \ \human\ \ |
| Nematodes | \ \ \caenorhabditis | \ \
| Drosophila | \ \ \fly | \ \
| A. mellifera | \ \ \honeybee1 | \ \
| A. gambiae | \ \ \culex | \ \
| S. cerevisiae | \ \ \saccharomyces | \ \
\ This table describes which training species was used for a particular group of assemblies.\ When available, the closest related training species was used.\
\ \\ Stanke M, Diekhans M, Baertsch R, Haussler D.\ \ Using native and syntenically mapped cDNA alignments to improve de novo gene finding.\ Bioinformatics. 2008 Mar 1;24(5):637-44.\ PMID: 18218656\
\ \\ Stanke M, Waack S.\ \ Gene prediction with a hidden Markov model and a new intron submodel.\ Bioinformatics. 2003 Oct;19 Suppl 2:ii215-25.\ PMID: 14534192\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 180,0,0\ group genes\ html ../../augustusGene\ longLabel AUGUSTUS ab initio gene predictions v3.1\ parent genePredArchive\ shortLabel AUGUSTUS\ track augustusGene\ type genePred\ visibility pack\ avada Avada Variants bigBed 9 + Avada Variants extracted from full text publications 1 100 0 0 0 127 127 127 0 0 0The tracks that are listed here contain genetic variants and links to scientific publications that \ mention them.
\\ For additional information please click on the hyperlink of the respective track above.\
\ By default, each variant is labeled with the nucleotide change. Hover over the\ feature to see more information, explained on the track details page of the particular track\ or when clicking onto the feature.
\\ For data provenance, access and descriptions, please click the documentation via the link above.\
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/avada.bb\ dataVersion release 1\ exonNumbers off\ html varsInPubs.html\ longLabel Avada Variants extracted from full text publications\ mouseOver Variant: $variant\ These tracks indicate regions with uniquely mappable reads of particular lengths before and after\ bisulfite conversion. Both Umap and Bismap tracks contain single-read mappability and multi-read\ mappability tracks for four different read lengths: 24 bp, 36 bp, 50 bp, and 100 bp.
\\ You can use these tracks for many purposes, including filtering unreliable signal from\ sequencing assays. The Bismap track can help filter unreliable signal from sequencing assays\ involving bisulfite conversion, such as whole-genome bisulfite sequencing or reduced representation\ bisulfite sequencing.
\ \ \These tracks mark any region of the bisulfite-converted genome that is uniquely mappable by\ at least one k-mer on the specified strand. Mappability of the forward strand was\ generated by converting all instances of cytosine to thymine. Similarly, mappability of the\ reverse strand was generated by converting all instances of guanine to adenine.
\To calculate the single-read mappability, you must find the overlap of a given region with\ the region that is uniquely mappable on both strands. Regions not uniquely mappable on both\ strands or have a low multi-read mappability might bias the downstream analysis.
These tracks represent the probability that a randomly selected k-mer which overlaps\ with a given position is uniquely mappable. Multi-read mappability track is calculated for\ k-mers that are uniquely mappable on both strands, and thus there is no strand\ specification.
These tracks mark any region of the genome that is uniquely mappable by at least one\ k-mer. To calculate the single-read mappability, you must find the overlap of a given\ region with this track.
These tracks represent the probability that a randomly selected k-mer which overlaps\ with a given position is uniquely mappable.
For greater detail and explanatory diagrams, see the\ preprint, the\ Umap and Bismap project website, or the\ Umap and Bismap software\ documentation.\ \
\ The raw data can be explored interactively with the Table Browser, or the Data Integrator. For automated analysis, genome annotation is stored in a bigBed\ or bigWig file that can be downloaded from the\ download\ server. Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed or bigWigToWig, which can be compiled from the source code or\ downloaded as a precompiled binary for your system. Instructions for downloading source code and\ binaries can be found here.\ The tool can also be used to obtain only features within a given range, for example:
\ bigBedToBed -chrom=chr6 -start=0 -end=1000000\ http://hgdownload.soe.ucsc.edu/gbdb/hg38/hoffmanMappability/k24.Unique.Mappability.bb stdout\\ Please refer to our mailing list archives for questions, or our\ Data Access FAQ for more\ information.
\ \\ Anshul Kundaje (Stanford\ University) created the original Umap software in MATLAB. The original Umap repository is available\ here.\ Mehran Karimzadeh (Michael Hoffman\ lab, Princess Margaret Cancer Centre) implemented the Python version of Umap and added features,\ including Bismap.
\ \\ Karimzadeh M, Ernst C, Kundaje A, Hoffman MM.,\ Umap and Bismap:\ quantifying genome and methylome mappability\ bioRxiv bioRxiv, p. 095463, 2016.; doi: https://doi.org/10.1101/095463.
\ map 0 compositeTrack on\ group map\ html mappability\ longLabel Single-read and multi-read mappability after bisulfite conversion\ noInherit on\ parent mappability\ shortLabel Bismap\ subGroup1 view Views SR=Single-read MR=Multi-read\ track bismap\ type bigWig\ visibility full\ gnomADPextBladder Bladder bigWig 0 1 gnomAD pext Bladder 0 100 170 0 0 212 127 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Bladder.bw\ color 170,0,0\ longLabel gnomAD pext Bladder\ parent gnomadPext off\ shortLabel Bladder\ track gnomADPextBladder\ visibility hide\ bloodHao Blood (PBMC) Hao Peripheral blood mononuclear cells (PBMC) from Hao et al 2020 0 100 0 0 0 127 127 127 0 0 0\ This track displays data from Integrated analysis of\ multimodal single-cell data. Human peripheral blood mononuclear cells\ (PBMCs) taken from pre-vaccinated and post-vaccinated individuals were profiled\ using both CITE-seq and ECCITE-seq. A total of 57 cell type clusters were\ identified and each cluster included cells from all 24 samples with rare\ exceptions. This dataset contains three annotations for cell clustering: Level\ 1 (8 cell types), Level 2 (30 cell types), Level 3 (57 cell types).
\ \\ This track collection contains six bar chart tracks of RNA expression in PBMCs\ where cells are grouped by cell type level 1 \ (Blood PBMC Cells), cell type level 2 \ (Blood PBMC Cells 2), \ cell type level 3 (Blood PBMC Cells 3), donor \ (Blood PBMC Donor), phase of cell cycle \ (Blood PBMC Phase), or time into experiment \ (Blood PBMC Time). The default track displayed \ is Blood PBMC Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ PBMC samples were taken from 8 volunteers ages 20-49 enrolled in an HIV\ vaccine trial (NCT01578889). A total of 24 blood samples were collected at 3\ time points: day 0 (the day before), day 3, and day 7 after the administration\ of a VSV-vectored HIV vaccine. Samples were collected at these different time\ points to minimize batch effects. Cells were then divided into separate\ aliquots for modified versions of the 3' CITE-seq and 5' ECCITE-seq staining\ protocols. In the 3' CITE-seq staining protocol, the samples are simultaneously\ stained with the antibody and unique hashtag. Whereas, 5' ECCITE-seq samples\ are stained first with a unique hashtag. 3' libraries were loaded into 8 lanes\ of a 10x Genomics Chip B using the 10x Genomics 3' v3 kit. 5' libraries\ were loaded into 2 lanes of a 10x Genomics Chip A using the 10x Genomics V(D)J\ kit (v1). Both 3' and 5' libraries were pooled together and sequenced on an\ Illumina Novaseq S4 flowcell. In total, 210,911 cells were profiled after \ quality control and doublet filtration.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell \ Browser. The UCSC command line utility matrixClusterColumns, matrixToBarChart, \ and bedToBigBed were used to transform these into a bar chart format bigBed file \ that can be visualized. The coloring was done by defining colors for the broad \ level cell classes and then using another UCSC utility, hcaColorCells, to interpolate \ the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yuhan Hao, Stephanie Hao, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ singleCell 0 group singleCell\ longLabel Peripheral blood mononuclear cells (PBMC) from Hao et al 2020\ shortLabel Blood (PBMC) Hao\ superTrack on\ track bloodHao\ visibility hide\ adult_wblood_models Blood models bigBed 12 + Adult Blood transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-WBlood.bb\ longLabel Adult Blood transcript models\ parent sample_models_view on\ shortLabel Blood models\ subGroups view=sample_models_view sample=adult_wblood type=models\ track adult_wblood_models\ type bigBed 12 +\ visibility squish\ adult_wblood_ont_post_models Blood ONT post models bigBed 12 + Adult Blood ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_WBlood01Rep1.bb\ itemRgb on\ longLabel Adult Blood ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Blood ONT post models\ subGroups view=per_expr_models_view sample=adult_wblood type=post_capture_ont_models\ track adult_wblood_ont_post_models\ type bigBed 12 +\ visibility hide\ adult_wblood_ont_post_reads Blood ONT post reads bam Adult Blood ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_WBlood01Rep1.bam\ longLabel Adult Blood ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Blood ONT post reads\ subGroups view=per_expr_reads_view sample=adult_wblood type=post_capture_ont_reads\ track adult_wblood_ont_post_reads\ type bam\ visibility hide\ adult_wblood_ont_pre_models Blood ONT pre models bigBed 12 + Adult Blood ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_WBlood01Rep1.bb\ itemRgb on\ longLabel Adult Blood ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Blood ONT pre models\ subGroups view=per_expr_models_view sample=adult_wblood type=pre_capture_ont_models\ track adult_wblood_ont_pre_models\ type bigBed 12 +\ visibility hide\ adult_wblood_ont_pre_reads Blood ONT pre reads bam Adult Blood ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_WBlood01Rep1.bam\ longLabel Adult Blood ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Blood ONT pre reads\ subGroups view=per_expr_reads_view sample=adult_wblood type=pre_capture_ont_reads\ track adult_wblood_ont_pre_reads\ type bam\ visibility hide\ adult_wblood_pacbio_post_models Blood PB post models bigBed 12 + Adult Blood PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_WBlood01Rep1.bb\ itemRgb on\ longLabel Adult Blood PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Blood PB post models\ subGroups view=per_expr_models_view sample=adult_wblood type=post_capture_pacbio_models\ track adult_wblood_pacbio_post_models\ type bigBed 12 +\ visibility hide\ adult_wblood_pacbio_post_reads Blood PB post reads bam Adult Blood PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_WBlood01Rep1.bam\ longLabel Adult Blood PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Blood PB post reads\ subGroups view=per_expr_reads_view sample=adult_wblood type=post_capture_pacbio_reads\ track adult_wblood_pacbio_post_reads\ type bam\ visibility hide\ adult_wblood_pacbio_pre_models Blood PB pre models bigBed 12 + Adult Blood PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_WBlood01Rep1.bb\ itemRgb on\ longLabel Adult Blood PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Blood PB pre models\ subGroups view=per_expr_models_view sample=adult_wblood type=pre_capture_pacbio_models\ track adult_wblood_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ adult_wblood_pacbio_pre_reads Blood PB pre reads bam Adult Blood PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_WBlood01Rep1.bam\ longLabel Adult Blood PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Blood PB pre reads\ subGroups view=per_expr_reads_view sample=adult_wblood type=pre_capture_pacbio_reads\ track adult_wblood_pacbio_pre_reads\ type bam\ visibility hide\ bloodHaoCellType Blood PBMC Cells bigBarChart Blood (PBMCs) binned by cell type (level 1) from Hao et al 2020 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ This track displays data from Integrated analysis of\ multimodal single-cell data. Human peripheral blood mononuclear cells\ (PBMCs) taken from pre-vaccinated and post-vaccinated individuals were profiled\ using both CITE-seq and ECCITE-seq. A total of 57 cell type clusters were\ identified and each cluster included cells from all 24 samples with rare\ exceptions. This dataset contains three annotations for cell clustering: Level\ 1 (8 cell types), Level 2 (30 cell types), Level 3 (57 cell types).
\ \\ This track collection contains six bar chart tracks of RNA expression in PBMCs\ where cells are grouped by cell type level 1 \ (Blood PBMC Cells), cell type level 2 \ (Blood PBMC Cells 2), \ cell type level 3 (Blood PBMC Cells 3), donor \ (Blood PBMC Donor), phase of cell cycle \ (Blood PBMC Phase), or time into experiment \ (Blood PBMC Time). The default track displayed \ is Blood PBMC Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ PBMC samples were taken from 8 volunteers ages 20-49 enrolled in an HIV\ vaccine trial (NCT01578889). A total of 24 blood samples were collected at 3\ time points: day 0 (the day before), day 3, and day 7 after the administration\ of a VSV-vectored HIV vaccine. Samples were collected at these different time\ points to minimize batch effects. Cells were then divided into separate\ aliquots for modified versions of the 3' CITE-seq and 5' ECCITE-seq staining\ protocols. In the 3' CITE-seq staining protocol, the samples are simultaneously\ stained with the antibody and unique hashtag. Whereas, 5' ECCITE-seq samples\ are stained first with a unique hashtag. 3' libraries were loaded into 8 lanes\ of a 10x Genomics Chip B using the 10x Genomics 3' v3 kit. 5' libraries\ were loaded into 2 lanes of a 10x Genomics Chip A using the 10x Genomics V(D)J\ kit (v1). Both 3' and 5' libraries were pooled together and sequenced on an\ Illumina Novaseq S4 flowcell. In total, 210,911 cells were profiled after \ quality control and doublet filtration.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell \ Browser. The UCSC command line utility matrixClusterColumns, matrixToBarChart, \ and bedToBigBed were used to transform these into a bar chart format bigBed file \ that can be visualized. The coloring was done by defining colors for the broad \ level cell classes and then using another UCSC utility, hcaColorCells, to interpolate \ the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yuhan Hao, Stephanie Hao, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ singleCell 1 barChartBars B_cell T_cell_CD4+ T_cell_CD8+ dendritic_cell_(DC) monocyte natural_killer_cell_(NK) other T_cell_other\ barChartColors #fe3247 #fe3248 #fe3248 #e92812 #e02900 #fb2e3e #f01111 #fe3247\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/bloodHao/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/bloodHao/cell_type.bb\ defaultLabelFields name\ html bloodHao\ longLabel Blood (PBMCs) binned by cell type (level 1) from Hao et al 2020\ parent bloodHao\ shortLabel Blood PBMC Cells\ track bloodHaoCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ bloodHaoL2 Blood PBMC Cells 2 bigBarChart Blood PBMCs binned by cell type (level 2) from Hao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ This track displays data from Integrated analysis of\ multimodal single-cell data. Human peripheral blood mononuclear cells\ (PBMCs) taken from pre-vaccinated and post-vaccinated individuals were profiled\ using both CITE-seq and ECCITE-seq. A total of 57 cell type clusters were\ identified and each cluster included cells from all 24 samples with rare\ exceptions. This dataset contains three annotations for cell clustering: Level\ 1 (8 cell types), Level 2 (30 cell types), Level 3 (57 cell types).
\ \\ This track collection contains six bar chart tracks of RNA expression in PBMCs\ where cells are grouped by cell type level 1 \ (Blood PBMC Cells), cell type level 2 \ (Blood PBMC Cells 2), \ cell type level 3 (Blood PBMC Cells 3), donor \ (Blood PBMC Donor), phase of cell cycle \ (Blood PBMC Phase), or time into experiment \ (Blood PBMC Time). The default track displayed \ is Blood PBMC Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ PBMC samples were taken from 8 volunteers ages 20-49 enrolled in an HIV\ vaccine trial (NCT01578889). A total of 24 blood samples were collected at 3\ time points: day 0 (the day before), day 3, and day 7 after the administration\ of a VSV-vectored HIV vaccine. Samples were collected at these different time\ points to minimize batch effects. Cells were then divided into separate\ aliquots for modified versions of the 3' CITE-seq and 5' ECCITE-seq staining\ protocols. In the 3' CITE-seq staining protocol, the samples are simultaneously\ stained with the antibody and unique hashtag. Whereas, 5' ECCITE-seq samples\ are stained first with a unique hashtag. 3' libraries were loaded into 8 lanes\ of a 10x Genomics Chip B using the 10x Genomics 3' v3 kit. 5' libraries\ were loaded into 2 lanes of a 10x Genomics Chip A using the 10x Genomics V(D)J\ kit (v1). Both 3' and 5' libraries were pooled together and sequenced on an\ Illumina Novaseq S4 flowcell. In total, 210,911 cells were profiled after \ quality control and doublet filtration.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell \ Browser. The UCSC command line utility matrixClusterColumns, matrixToBarChart, \ and bedToBigBed were used to transform these into a bar chart format bigBed file \ that can be visualized. The coloring was done by defining colors for the broad \ level cell classes and then using another UCSC utility, hcaColorCells, to interpolate \ the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yuhan Hao, Stephanie Hao, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ singleCell 1 barChartBars ASDC B_intermediate B_memory B_naive CD14_Mono CD16_Mono CD4_CTL CD4_Naive CD4_Proliferating CD4_TCM CD4_TEM CD8_Naive CD8_Proliferating CD8_TCM CD8_TEM Doublet Eryth HSPC ILC MAIT NK NK_Proliferating NK_CD56bright Plasmablast Platelet Treg cDC1 cDC2 dnT gdT pDC\ barChartColors #f77170 #fe3246 #fe3246 #fe3246 #e02901 #e22803 #fd3145 #fe3248 #fb737b #fe3248 #fe3248 #fe3248 #fc737c #fe3248 #fd3145 #e22803 #fa9fa1 #fd7580 #fe7683 #fe3246 #fb2e3e #f82b36 #fd3043 #fc747d #f01212 #fe3248 #f77071 #e5270a #fe7685 #fe3145 #f72c34\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/bloodHao/celltype.l2.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/bloodHao/celltype.l2.bb\ defaultLabelFields name\ html bloodHao\ labelFields name,name2\ longLabel Blood PBMCs binned by cell type (level 2) from Hao et al 2020\ parent bloodHao\ shortLabel Blood PBMC Cells 2\ track bloodHaoL2\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ bloodHaoL3 Blood PBMC Cells 3 bigBarChart Blood PBMCs binned by cell type (level 3) from Hao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ This track displays data from Integrated analysis of\ multimodal single-cell data. Human peripheral blood mononuclear cells\ (PBMCs) taken from pre-vaccinated and post-vaccinated individuals were profiled\ using both CITE-seq and ECCITE-seq. A total of 57 cell type clusters were\ identified and each cluster included cells from all 24 samples with rare\ exceptions. This dataset contains three annotations for cell clustering: Level\ 1 (8 cell types), Level 2 (30 cell types), Level 3 (57 cell types).
\ \\ This track collection contains six bar chart tracks of RNA expression in PBMCs\ where cells are grouped by cell type level 1 \ (Blood PBMC Cells), cell type level 2 \ (Blood PBMC Cells 2), \ cell type level 3 (Blood PBMC Cells 3), donor \ (Blood PBMC Donor), phase of cell cycle \ (Blood PBMC Phase), or time into experiment \ (Blood PBMC Time). The default track displayed \ is Blood PBMC Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ PBMC samples were taken from 8 volunteers ages 20-49 enrolled in an HIV\ vaccine trial (NCT01578889). A total of 24 blood samples were collected at 3\ time points: day 0 (the day before), day 3, and day 7 after the administration\ of a VSV-vectored HIV vaccine. Samples were collected at these different time\ points to minimize batch effects. Cells were then divided into separate\ aliquots for modified versions of the 3' CITE-seq and 5' ECCITE-seq staining\ protocols. In the 3' CITE-seq staining protocol, the samples are simultaneously\ stained with the antibody and unique hashtag. Whereas, 5' ECCITE-seq samples\ are stained first with a unique hashtag. 3' libraries were loaded into 8 lanes\ of a 10x Genomics Chip B using the 10x Genomics 3' v3 kit. 5' libraries\ were loaded into 2 lanes of a 10x Genomics Chip A using the 10x Genomics V(D)J\ kit (v1). Both 3' and 5' libraries were pooled together and sequenced on an\ Illumina Novaseq S4 flowcell. In total, 210,911 cells were profiled after \ quality control and doublet filtration.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell \ Browser. The UCSC command line utility matrixClusterColumns, matrixToBarChart, \ and bedToBigBed were used to transform these into a bar chart format bigBed file \ that can be visualized. The coloring was done by defining colors for the broad \ level cell classes and then using another UCSC utility, hcaColorCells, to interpolate \ the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yuhan Hao, Stephanie Hao, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ singleCell 1 barChartBars ASDC_mDC ASDC_pDC B_intermediate_kappa B_intermediate_lambda B_memory_kappa B_memory_lambda B_naive_kappa B_naive_lambda CD14_Mono CD16_Mono CD4_CTL CD4_Naive CD4_Proliferating CD4_TCM_1 CD4_TCM_2 CD4_TCM_3 CD4_TEM_1 CD4_TEM_2 CD4_TEM_3 CD4_TEM_4 CD8_Naive CD8_Naive_2 CD8_Proliferating CD8_TCM_1 CD8_TCM_2 CD8_TCM_3 CD8_TEM_1 CD8_TEM_2 CD8_TEM_3 CD8_TEM_4 CD8_TEM_5 CD8_TEM_6 Doublet Eryth HSPC ILC MAIT NK_Proliferating NK_1 NK_2 NK_3 NK_4 NK_CD56bright Plasma Plasmablast Platelet Treg_Memory Treg_Naive cDC1 cDC2_1 cDC2_2 dnT_1 dnT_2 gdT_1 gdT_2 gdT_3 gdT_4 pDC\ barChartColors #fabfbc #fcc0c1 #fe3246 #fe3146 #fe3246 #fe3246 #fe3246 #fe3246 #e02901 #e22803 #fd3145 #fe3248 #fb737b #fe3248 #fd3145 #fe3248 #fe3247 #fe7785 #fe3248 #ffa4ad #fe3248 #fe7684 #fc737c #fe3248 #fe3248 #fe3248 #fe3247 #fd3144 #fe7684 #fc3042 #fc2f41 #fe3246 #e22803 #fa9fa1 #fd7580 #fe7683 #fe3246 #f82b36 #fb2e3e #fa2d3c #fc2f40 #fc2f41 #fd3043 #fc747d #fdc1c4 #f01212 #fe3248 #fe3248 #f77071 #e22804 #e8270f #ff7785 #fd7581 #fd3145 #fb2e3f #fe3248 #fc3042 #f72c34\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/bloodHao/celltype.l3.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/bloodHao/celltype.l3.bb\ defaultLabelFields name\ html bloodHao\ labelFields name,name2\ longLabel Blood PBMCs binned by cell type (level 3) from Hao et al 2020\ parent bloodHao\ shortLabel Blood PBMC Cells 3\ track bloodHaoL3\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ bloodHaoDonor Blood PBMC Donor bigBarChart Blood PBMCs binned by blood donor from Hao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ This track displays data from Integrated analysis of\ multimodal single-cell data. Human peripheral blood mononuclear cells\ (PBMCs) taken from pre-vaccinated and post-vaccinated individuals were profiled\ using both CITE-seq and ECCITE-seq. A total of 57 cell type clusters were\ identified and each cluster included cells from all 24 samples with rare\ exceptions. This dataset contains three annotations for cell clustering: Level\ 1 (8 cell types), Level 2 (30 cell types), Level 3 (57 cell types).
\ \\ This track collection contains six bar chart tracks of RNA expression in PBMCs\ where cells are grouped by cell type level 1 \ (Blood PBMC Cells), cell type level 2 \ (Blood PBMC Cells 2), \ cell type level 3 (Blood PBMC Cells 3), donor \ (Blood PBMC Donor), phase of cell cycle \ (Blood PBMC Phase), or time into experiment \ (Blood PBMC Time). The default track displayed \ is Blood PBMC Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ PBMC samples were taken from 8 volunteers ages 20-49 enrolled in an HIV\ vaccine trial (NCT01578889). A total of 24 blood samples were collected at 3\ time points: day 0 (the day before), day 3, and day 7 after the administration\ of a VSV-vectored HIV vaccine. Samples were collected at these different time\ points to minimize batch effects. Cells were then divided into separate\ aliquots for modified versions of the 3' CITE-seq and 5' ECCITE-seq staining\ protocols. In the 3' CITE-seq staining protocol, the samples are simultaneously\ stained with the antibody and unique hashtag. Whereas, 5' ECCITE-seq samples\ are stained first with a unique hashtag. 3' libraries were loaded into 8 lanes\ of a 10x Genomics Chip B using the 10x Genomics 3' v3 kit. 5' libraries\ were loaded into 2 lanes of a 10x Genomics Chip A using the 10x Genomics V(D)J\ kit (v1). Both 3' and 5' libraries were pooled together and sequenced on an\ Illumina Novaseq S4 flowcell. In total, 210,911 cells were profiled after \ quality control and doublet filtration.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell \ Browser. The UCSC command line utility matrixClusterColumns, matrixToBarChart, \ and bedToBigBed were used to transform these into a bar chart format bigBed file \ that can be visualized. The coloring was done by defining colors for the broad \ level cell classes and then using another UCSC utility, hcaColorCells, to interpolate \ the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yuhan Hao, Stephanie Hao, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ singleCell 1 barChartBars P1 P2 P3 P4 P5 P6 P7 P8\ barChartColors #fd3144 #fe3247 #fd3144 #fd3144 #f32b2b #f92e3a #f52c30 #fa2f3c\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/bloodHao/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/bloodHao/donor.bb\ defaultLabelFields name\ html bloodHao\ labelFields name,name2\ longLabel Blood PBMCs binned by blood donor from Hao et al 2020\ parent bloodHao\ shortLabel Blood PBMC Donor\ track bloodHaoDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ bloodHaoPhase Blood PBMC Phase bigBarChart Blood PBMCs binned by phase of cell cycle from Hao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ This track displays data from Integrated analysis of\ multimodal single-cell data. Human peripheral blood mononuclear cells\ (PBMCs) taken from pre-vaccinated and post-vaccinated individuals were profiled\ using both CITE-seq and ECCITE-seq. A total of 57 cell type clusters were\ identified and each cluster included cells from all 24 samples with rare\ exceptions. This dataset contains three annotations for cell clustering: Level\ 1 (8 cell types), Level 2 (30 cell types), Level 3 (57 cell types).
\ \\ This track collection contains six bar chart tracks of RNA expression in PBMCs\ where cells are grouped by cell type level 1 \ (Blood PBMC Cells), cell type level 2 \ (Blood PBMC Cells 2), \ cell type level 3 (Blood PBMC Cells 3), donor \ (Blood PBMC Donor), phase of cell cycle \ (Blood PBMC Phase), or time into experiment \ (Blood PBMC Time). The default track displayed \ is Blood PBMC Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ PBMC samples were taken from 8 volunteers ages 20-49 enrolled in an HIV\ vaccine trial (NCT01578889). A total of 24 blood samples were collected at 3\ time points: day 0 (the day before), day 3, and day 7 after the administration\ of a VSV-vectored HIV vaccine. Samples were collected at these different time\ points to minimize batch effects. Cells were then divided into separate\ aliquots for modified versions of the 3' CITE-seq and 5' ECCITE-seq staining\ protocols. In the 3' CITE-seq staining protocol, the samples are simultaneously\ stained with the antibody and unique hashtag. Whereas, 5' ECCITE-seq samples\ are stained first with a unique hashtag. 3' libraries were loaded into 8 lanes\ of a 10x Genomics Chip B using the 10x Genomics 3' v3 kit. 5' libraries\ were loaded into 2 lanes of a 10x Genomics Chip A using the 10x Genomics V(D)J\ kit (v1). Both 3' and 5' libraries were pooled together and sequenced on an\ Illumina Novaseq S4 flowcell. In total, 210,911 cells were profiled after \ quality control and doublet filtration.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell \ Browser. The UCSC command line utility matrixClusterColumns, matrixToBarChart, \ and bedToBigBed were used to transform these into a bar chart format bigBed file \ that can be visualized. The coloring was done by defining colors for the broad \ level cell classes and then using another UCSC utility, hcaColorCells, to interpolate \ the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yuhan Hao, Stephanie Hao, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ singleCell 1 barChartBars G1 G2M S\ barChartColors #e92913 #fd3144 #fe3247\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/bloodHao/Phase.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/bloodHao/Phase.bb\ defaultLabelFields name\ html bloodHao\ labelFields name,name2\ longLabel Blood PBMCs binned by phase of cell cycle from Hao et al 2020\ parent bloodHao\ shortLabel Blood PBMC Phase\ track bloodHaoPhase\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ bloodHaoTime Blood PBMC Time bigBarChart Blood PBMCs binned by time into experiment from Hao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ This track displays data from Integrated analysis of\ multimodal single-cell data. Human peripheral blood mononuclear cells\ (PBMCs) taken from pre-vaccinated and post-vaccinated individuals were profiled\ using both CITE-seq and ECCITE-seq. A total of 57 cell type clusters were\ identified and each cluster included cells from all 24 samples with rare\ exceptions. This dataset contains three annotations for cell clustering: Level\ 1 (8 cell types), Level 2 (30 cell types), Level 3 (57 cell types).
\ \\ This track collection contains six bar chart tracks of RNA expression in PBMCs\ where cells are grouped by cell type level 1 \ (Blood PBMC Cells), cell type level 2 \ (Blood PBMC Cells 2), \ cell type level 3 (Blood PBMC Cells 3), donor \ (Blood PBMC Donor), phase of cell cycle \ (Blood PBMC Phase), or time into experiment \ (Blood PBMC Time). The default track displayed \ is Blood PBMC Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ PBMC samples were taken from 8 volunteers ages 20-49 enrolled in an HIV\ vaccine trial (NCT01578889). A total of 24 blood samples were collected at 3\ time points: day 0 (the day before), day 3, and day 7 after the administration\ of a VSV-vectored HIV vaccine. Samples were collected at these different time\ points to minimize batch effects. Cells were then divided into separate\ aliquots for modified versions of the 3' CITE-seq and 5' ECCITE-seq staining\ protocols. In the 3' CITE-seq staining protocol, the samples are simultaneously\ stained with the antibody and unique hashtag. Whereas, 5' ECCITE-seq samples\ are stained first with a unique hashtag. 3' libraries were loaded into 8 lanes\ of a 10x Genomics Chip B using the 10x Genomics 3' v3 kit. 5' libraries\ were loaded into 2 lanes of a 10x Genomics Chip A using the 10x Genomics V(D)J\ kit (v1). Both 3' and 5' libraries were pooled together and sequenced on an\ Illumina Novaseq S4 flowcell. In total, 210,911 cells were profiled after \ quality control and doublet filtration.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell \ Browser. The UCSC command line utility matrixClusterColumns, matrixToBarChart, \ and bedToBigBed were used to transform these into a bar chart format bigBed file \ that can be visualized. The coloring was done by defining colors for the broad \ level cell classes and then using another UCSC utility, hcaColorCells, to interpolate \ the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yuhan Hao, Stephanie Hao, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ singleCell 1 barChartBars 0 3 7\ barChartColors #f92e3b #fc3043 #fc3041\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/bloodHao/time.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/bloodHao/time.bb\ defaultLabelFields name\ html bloodHao\ labelFields name,name2\ longLabel Blood PBMCs binned by time into experiment from Hao et al 2020\ parent bloodHao\ shortLabel Blood PBMC Time\ track bloodHaoTime\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=multimodal-pbmc+sct&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ adult_brain_models Brain models bigBed 12 + Adult Brain transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-Brain.bb\ longLabel Adult Brain transcript models\ parent sample_models_view on\ shortLabel Brain models\ subGroups view=sample_models_view sample=adult_brain type=models\ track adult_brain_models\ type bigBed 12 +\ visibility squish\ adult_brain_ont_post_models Brain ONT post models bigBed 12 + Adult Brain ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Brain03Rep1.bb\ itemRgb on\ longLabel Adult Brain ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Brain ONT post models\ subGroups view=per_expr_models_view sample=adult_brain type=post_capture_ont_models\ track adult_brain_ont_post_models\ type bigBed 12 +\ visibility hide\ adult_brain_ont_post_reads Brain ONT post reads bam Adult Brain ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Brain03Rep1.bam\ longLabel Adult Brain ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Brain ONT post reads\ subGroups view=per_expr_reads_view sample=adult_brain type=post_capture_ont_reads\ track adult_brain_ont_post_reads\ type bam\ visibility hide\ adult_brain_ont_pre_models Brain ONT pre models bigBed 12 + Adult Brain ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Brain03Rep1.bb\ itemRgb on\ longLabel Adult Brain ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Brain ONT pre models\ subGroups view=per_expr_models_view sample=adult_brain type=pre_capture_ont_models\ track adult_brain_ont_pre_models\ type bigBed 12 +\ visibility hide\ adult_brain_ont_pre_reads Brain ONT pre reads bam Adult Brain ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Brain03Rep1.bam\ longLabel Adult Brain ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Brain ONT pre reads\ subGroups view=per_expr_reads_view sample=adult_brain type=pre_capture_ont_reads\ track adult_brain_ont_pre_reads\ type bam\ visibility hide\ adult_brain_pacbio_post_models Brain PB post models bigBed 12 + Adult Brain PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Brain03Rep1.bb\ itemRgb on\ longLabel Adult Brain PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Brain PB post models\ subGroups view=per_expr_models_view sample=adult_brain type=post_capture_pacbio_models\ track adult_brain_pacbio_post_models\ type bigBed 12 +\ visibility hide\ adult_brain_pacbio_post_reads Brain PB post reads bam Adult Brain PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Brain03Rep1.bam\ longLabel Adult Brain PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Brain PB post reads\ subGroups view=per_expr_reads_view sample=adult_brain type=post_capture_pacbio_reads\ track adult_brain_pacbio_post_reads\ type bam\ visibility hide\ adult_brain_pacbio_pre_models Brain PB pre models bigBed 12 + Adult Brain PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Brain03Rep1.bb\ itemRgb on\ longLabel Adult Brain PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Brain PB pre models\ subGroups view=per_expr_models_view sample=adult_brain type=pre_capture_pacbio_models\ track adult_brain_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ adult_brain_pacbio_pre_reads Brain PB pre reads bam Adult Brain PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Brain03Rep1.bam\ longLabel Adult Brain PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Brain PB pre reads\ subGroups view=per_expr_reads_view sample=adult_brain type=pre_capture_pacbio_reads\ track adult_brain_pacbio_pre_reads\ type bam\ visibility hide\ gnomADPextBrain_Amygdala Brain-Amygdala bigWig 0 1 gnomAD pext Brain-Amygdala 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Amygdala.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Amygdala\ parent gnomadPext off\ shortLabel Brain-Amygdala\ track gnomADPextBrain_Amygdala\ visibility hide\ gnomADPextBrain_Anteriorcingulatecortex_BA24 Brain-Anterior Cingulate Cortex (BA24) bigWig 0 1 gnomAD pext Brain-Anterior Cingulate Cortex (BA24) 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Anteriorcingulatecortex_BA24.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Anterior Cingulate Cortex (BA24)\ parent gnomadPext off\ shortLabel Brain-Anterior Cingulate Cortex (BA24)\ track gnomADPextBrain_Anteriorcingulatecortex_BA24\ visibility hide\ gnomADPextBrain_Caudate_basalganglia Brain-Caudate (basal ganglia) bigWig 0 1 gnomAD pext Brain-Caudate (basal ganglia) 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Caudate_basalganglia.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Caudate (basal ganglia)\ parent gnomadPext off\ shortLabel Brain-Caudate (basal ganglia)\ track gnomADPextBrain_Caudate_basalganglia\ visibility hide\ gnomADPextBrain_CerebellarHemisphere Brain-Cerebellar Hemisphere bigWig 0 1 gnomAD pext Brain-Cerebellar Hemisphere 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_CerebellarHemisphere.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Cerebellar Hemisphere\ parent gnomadPext off\ shortLabel Brain-Cerebellar Hemisphere\ track gnomADPextBrain_CerebellarHemisphere\ visibility hide\ gnomADPextBrain_Cerebellum Brain-Cerebellum bigWig 0 1 gnomAD pext Brain-Cerebellum 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Cerebellum.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Cerebellum\ parent gnomadPext off\ shortLabel Brain-Cerebellum\ track gnomADPextBrain_Cerebellum\ visibility hide\ gnomADPextBrain_Cortex Brain-Cortex bigWig 0 1 gnomAD pext Brain-Cortex 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Cortex.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Cortex\ parent gnomadPext off\ shortLabel Brain-Cortex\ track gnomADPextBrain_Cortex\ visibility hide\ gnomADPextBrain_FrontalCortex_BA9 Brain-Frontal Cortex (BA9) bigWig 0 1 gnomAD pext Brain-Frontal Cortex (BA9) 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_FrontalCortex_BA9.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Frontal Cortex (BA9)\ parent gnomadPext off\ shortLabel Brain-Frontal Cortex (BA9)\ track gnomADPextBrain_FrontalCortex_BA9\ visibility hide\ gnomADPextBrain_Hippocampus Brain-Hippocampus bigWig 0 1 gnomAD pext Brain-Hippocampus 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Hippocampus.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Hippocampus\ parent gnomadPext off\ shortLabel Brain-Hippocampus\ track gnomADPextBrain_Hippocampus\ visibility hide\ gnomADPextBrain_Hypothalamus Brain-Hypothalamus bigWig 0 1 gnomAD pext Brain-Hypothalamus 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Hypothalamus.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Hypothalamus\ parent gnomadPext off\ shortLabel Brain-Hypothalamus\ track gnomADPextBrain_Hypothalamus\ visibility hide\ gnomADPextBrain_Nucleusaccumbens_basalganglia Brain-Nucleus Accumbens (basal ganglia) bigWig 0 1 gnomAD pext Brain-Nucleus Accumbens (basal ganglia) 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Nucleusaccumbens_basalganglia.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Nucleus Accumbens (basal ganglia)\ parent gnomadPext off\ shortLabel Brain-Nucleus Accumbens (basal ganglia)\ track gnomADPextBrain_Nucleusaccumbens_basalganglia\ visibility hide\ gnomADPextBrain_Putamen_basalganglia Brain-Putamen (basal ganglia) bigWig 0 1 gnomAD pext Brain-Putamen (basal ganglia) 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Putamen_basalganglia.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Putamen (basal ganglia)\ parent gnomadPext off\ shortLabel Brain-Putamen (basal ganglia)\ track gnomADPextBrain_Putamen_basalganglia\ visibility hide\ gnomADPextBrain_Spinalcord_cervicalc_1 Brain-Spinal Cord (cervicalc 1) bigWig 0 1 gnomAD pext Brain-Spinal Cord (cervicalc 1) 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Spinalcord_cervicalc_1.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Spinal Cord (cervicalc 1)\ parent gnomadPext off\ shortLabel Brain-Spinal Cord (cervicalc 1)\ track gnomADPextBrain_Spinalcord_cervicalc_1\ visibility hide\ gnomADPextBrain_Substantianigra Brain-Substantia Nigra bigWig 0 1 gnomAD pext Brain-Substantia Nigra 0 100 238 238 0 246 246 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Brain_Substantianigra.bw\ color 238,238,0\ longLabel gnomAD pext Brain-Substantia Nigra\ parent gnomadPext off\ shortLabel Brain-Substantia Nigra\ track gnomADPextBrain_Substantianigra\ visibility hide\ gnomADPextBreast_MammaryTissue Breast-Mammary Tissue bigWig 0 1 gnomAD pext Breast-Mammary Tissue 0 100 51 204 204 153 229 229 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Breast_MammaryTissue.bw\ color 51,204,204\ longLabel gnomAD pext Breast-Mammary Tissue\ parent gnomadPext off\ shortLabel Breast-Mammary Tissue\ track gnomADPextBreast_MammaryTissue\ visibility hide\ cactus447way Cactus 447-way bigMaf Cactus alignment on 447 mammal species, including Zoonomia genomes and 233 primates 3 100 0 10 100 0 90 10 0 0 0\ This track shows a multiple alignment of 447 mammalian genomes made with Cactus and constraint scores derived from it.\ To build this track, the Zoonomia 241 alignment was used as a starting point, all primates and a few outdated\ assemblies were removed and an alignment between 233 newly sequenced primates was added. See the Methods section below for details, and \ also the publications by Kuderna et al. 2023 in the Reference section.\ All alignments and operations on them were performed using the Cactus toolkit.\
\ \\ This track shows four phyloP conservation score subtracks computed from the\ 447-way Cactus alignment (and a primates subset of it):\
\ The SSREV substitution model is strand-symmetric, which avoids\ strand-dependent bias in single-base conservation scores (Pollard\ et al. 2010, supplementary section 2.4) -- relevant when analyzing\ transcript-related nucleotides such as splice sites, miRNA seed regions, or\ other strand-specific sequence features. The REV model is the standard\ phyloP model and is appropriate for general genome-wide conservation\ analysis. The primates subset tracks restrict scoring to the 233 primate\ genomes included in the alignment, useful when conservation across\ non-primate mammals would dilute primate-specific signal.\
\ \\ Downloads for data in this track are available from the directory:\
\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ size of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ Note that excluding species from the pairwise display does not alter the\ conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. Missing sequence in any\ assembly is highlighted in the track display by regions of yellow when zoomed\ out and by Ns when displayed at base level. The following conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment. The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\\ Codon translation is available in base-level display mode if the\ displayed region is identified as a coding segment. To display this annotation,\ select the species for translation from the pull-down menu in the Codon\ Translation configuration section at the top of the page. Then, select one of\ the following modes:\
\ Codon translation uses the following gene tracks as the basis for translation:\
\ \\
\ Table 2. Gene tracks used for codon translation.\\ Gene Track Species \ RefSeq Genes Bos mutus, Canis lupus familiaris, Carlito syrichta, Cercocebus atys, Chinchilla lanigera, Colobus angolensis, Condylura cristata, Dipodomys ordii, Elephantulus edwardii, Eptesicus fuscus, Felis catus, Felis catus fca126, Fukomys damarensis, Homo sapiens, Ictidomys tridecemlineatus, Macaca mulatta, Macaca nemestrina, Marmota marmota, Microtus ochrogaster, Miniopterus natalensis, Mus musculus, Mus pahari, Myotis brandtii, Myotis davidii, Myotis lucifugus, Odobenus rosmarus, Orcinus orca, Otolemur garnettii, Peromyscus maniculatus, Piliocolobus tephrosceles, Propithecus coquerelli, Pteropus alecto, Pteropus vampyrus, Rattus norvegicus, Rhinopithecus roxellana, Saimiri boliviensis, Sorex araneus, Sus scrofa, Theropithecus gelada, Tupaia chinensis \ Ensembl Genes Cavia aperea \ Augustus Genes Eidolon helvum, Pteronotus parnellii \ no annotation Acinonyx jubatus, Acomys cahirinus, Ailuropoda melanoleuca, Ailurus fulgens, Allactaga bullata, Allenopithecus nigroviridis, Allochrocebus lhoesti, Allochrocebus preussi, Allochrocebus solatus, Alouatta belzebul, Alouatta caraya, Alouatta discolor, Alouatta juara, Alouatta macconnelli, Alouatta nigerrima, Alouatta palliata, Alouatta puruensis, Alouatta seniculus, Ammotragus lervia, Anoura caudifer, Antilocapra americana, Aotus azarae, Aotus griseimembra, Aotus nancymaae, Aotus trivirgatus, Aotus vociferans, Aplodontia rufa, Arctocebus calabarensis, Artibeus jamaicensis, Ateles geoffroyi_a, Ateles geoffroyi_b, Ateles belzebuth, Ateles chamek, Ateles marginatus, Ateles paniscus, Avahi laniger, Avahi peyrierasi, Balaenoptera acutorostrata, Balaenoptera bonaerensis, Beatragus hunteri, Bison bison, Bos indicus, Bos taurus, Bubalus bubalis, Cacajao ayresi, Cacajao calvus, Cacajao hosomi, Cacajao melanocephalus, Callibella humilis, Callimico goeldii, Callithrix geoffroyi, Callithrix jacchus, Callithrix kuhlii, Camelus bactrianus, Camelus dromedarius, Camelus ferus, Canis lupus VD, Canis lupus dingo, Canis lupus orion, Capra aegagrus, Capra hircus, Capromys pilorides, Carollia perspicillata, Castor canadensis, Catagonus wagneri, Cavia porcellus, Cavia tschudii, Cebuella niveiventris, Cebuella pygmaea, Cebus albifrons, Cebus olivaceus, Cebus unicolor, Cephalopachus bancanus, Ceratotherium simum, Ceratotherium simum cottoni, Cercocebus chrysogaster, Cercocebus lunulatus, Cercocebus torquatus, Cercopithecus ascanius, Cercopithecus cephus, Cercopithecus diana, Cercopithecus hamlyni, Cercopithecus lowei, Cercopithecus albogularis, Cercopithecus mona, Cercopithecus neglectus, Cercopithecus nictitans, Cercopithecus petaurista, Cercopithecus pogonias, Cercopithecus roloway, Chaetophractus vellerosus, Cheirogaleus major, Cheirogaleus medius, Cheracebus lucifer, Cheracebus lugens, Cheracebus regulus, Cheracebus torquatus, Chiropotes albinasus, Chiropotes israelita, Chiropotes sagulatus, Chlorocebus aethiops, Chlorocebus pygerythrus, Chlorocebus sabaeus, Choloepus didactylus, Choloepus hoffmanni, Chrysochloris asiatica, Colobus guereza, Colobus polykomos, Craseonycteris thonglongyai, Cricetomys gambianus, Cricetulus griseus, Crocidura indochinensis, Cryptoprocta ferox, Ctenodactylus gundi, Ctenomys sociabilis, Cuniculus paca, Dasyprocta punctata, Dasypus novemcinctus, Daubentonia madagascariensis, Delphinapterus leucas, Desmodus rotundus, Dicerorhinus sumatrensis, Diceros bicornis, Dinomys branickii, Dipodomys stephensi, Dolichotis patagonum, Echinops telfairi, Elaphurus davidianus, Ellobius lutescens, Ellobius talpinus, Enhydra lutris, Equus asinus, Equus caballus, Equus przewalskii, Erinaceus europaeus, Erythrocebus patas, Eschrichtius robustus, Eubalaena japonica, Eulemur albifrons, Eulemur collaris, Eulemur coronatus, Eulemur flavifrons, Eulemur fulvus, Eulemur macaco, Eulemur mongoz, Eulemur rubriventer, Eulemur rufus, Eulemur sanfordi, Felis nigripes, Galago moholi, Galago senegalensis, Galagoides demidoff, Galeopterus variegatus, Giraffa tippelskirchi, Glis glis, Gorilla beringei, Gorilla gorilla, Graphiurus murinus, Hapalemur alaotrensis, Hapalemur gilberti, Hapalemur griseus, Hapalemur meridionalis, Hapalemur occidentalis, Helogale parvula, Hemitragus hylocrius, Heterocephalus glaber, Heterohyrax brucei, Hippopotamus amphibius, Hipposideros armiger, Hipposideros galeritus, Hoolock leuconedys, Hyaena hyaena, Hydrochoerus hydrochaeris, Hylobates abbotti, Hylobates agilis, Hylobates klossii, Hylobates pileatus, Hylobates muelleri, Hylobates pileatus, Hystrix cristata, Indri indri, Inia geoffrensis, Jaculus jaculus, Kogia breviceps, Lagothrix lagothricha, Lasiurus borealis, Lemur catta, Leontocebus fuscicollis, Leontocebus illigeri, Leontocebus nigricollis, Leontopithecus chrysomelas, Leontopithecus rosalia, Lepilemur ankaranensis, Lepilemur dorsalis, Lepilemur ruficaudatus, Lepilemur septentrionalis, Leptonychotes weddellii, Lepus americanus, Lipotes vexillifer, Lophocebus aterrimus, Loris lydekkerianus, Loris tardigradus, Loxodonta africana, Lycaon pictus, Macaca arctoides, Macaca assamensis, Macaca cyclopis, Macaca fascicularis, Macaca fuscata, Macaca leonina, Macaca maura, Macaca nigra, Macaca radiata, Macaca siberu, Macaca silenus, Macaca thibetana, Macaca tonkeana, Macroglossus sobrinus, Mandrillus leucophaeus, Mandrillus sphinx, Manis javanica, Manis pentadactyla, Megaderma lyra, Mellivora capensis, Meriones unguiculatus, Mesocricetus auratus, Mesoplodon bidens, Mico argentatus, Mico humeralifer, Mico schneideri, Microcebus murinus, Microgale talazaci, Micronycteris hirsuta, Miniopterus schreibersii, Miopithecus ogouensis, Mirounga angustirostris, Mirza zaza, Monodon monoceros, Mormoops blainvillei, Moschus moschiferus, Mungos mungo, Murina feae, Mus caroli, Mus spretus, Muscardinus avellanarius, Mustela putorius, Myocastor coypus, Myotis myotis, Myrmecophaga tridactyla, Nannospalax galili, Nasalis larvatus, Neomonachus schauinslandi, Neophocaena asiaeorientalis, Noctilio leporinus, Nomascus annamensis, Nomascus concolor, Nomascus gabriellae, Nomascus siki_a, Nomascus siki_b, Nyctereutes procyonoides, Nycticebus bengalensis, Nycticebus coucang, Nycticebus pygmaeus, Ochotona princeps, Octodon degus, Odocoileus virginianus, Okapia johnstoni, Ondatra zibethicus, Onychomys torridus, Orycteropus afer, Oryctolagus cuniculus, Otocyon megalotis, Otolemur crassicaudatus, Ovis aries, Ovis canadensis, Pan paniscus, Pan troglodytes, Panthera onca, Panthera pardus, Panthera tigris, Pantholops hodgsonii, Papio anubis, Papio cynocephalus, Papio hamadryas, Papio kindae, Papio papio, Papio ursinus, Paradoxurus hermaphroditus, Perodicticus ibeanus, Perodicticus potto, Perognathus longimembris, Petromus typicus, Phocoena phocoena, Piliocolobus badius, Piliocolobus gordonorum, Piliocolobus kirkii, Pipistrellus pipistrellus, Pithecia albicans, Pithecia chrysocephala, Pithecia hirsuta, Pithecia mittermeieri, Pithecia pissinattii, Pithecia pithecia, Pithecia vanzolinii, Platanista gangetica, Plecturocebus bernhardi, Plecturocebus brunneus, Plecturocebus caligatus, Plecturocebus cinerascens, Plecturocebus cupreus, Plecturocebus dubius, Plecturocebus grovesi, Plecturocebus hoffmannsi, Plecturocebus miltoni, Plecturocebus moloch, Pongo abelii, Pongo pygmaeus, Presbytis comata, Presbytis mitrata, Procavia capensis, Prolemur simus, Propithecus coronatus, Propithecus diadema, Propithecus edwardsi, Propithecus perrieri, Propithecus tattersalli, Propithecus verreauxi, Psammomys obesus, Pteronura brasiliensis, Puma concolor, Pygathrix cinerea, Pygathrix nigripes, Pygathrix nigripes, Rangifer tarandus, Rhinolophus sinicus, Rhinopithecus bieti, Rhinopithecus strykeri, Rousettus aegyptiacus, Saguinus bicolor, Saguinus geoffroyi, Saguinus imperator, Saguinus inustus, Saguinus labiatus, Saguinus midas, Saguinus mystax, Saguinus oedipus, Saiga tatarica, Saimiri cassiquiarensis, Saimiri macrodon, Saimiri oerstedii, Saimiri sciureus, Saimiri ustus, Sapajus apella, Sapajus macrocephalus, Scalopus aquaticus, Semnopithecus entellus, Semnopithecus hypoleucos, Semnopithecus johnii, Semnopithecus priam, Semnopithecus schistaceus, Semnopithecus vetulus, Sigmodon hispidus, Solenodon paradoxus, Spermophilus dauricus, Spilogale gracilis, Suricata suricatta, Symphalangus syndactylus, Tadarida brasiliensis, Tamandua tetradactyla, Tapirus indicus, Tapirus terrestris, Tarsius lariang, Tarsius wallacei, Thryonomys swinderianus, Tolypeutes matacus, Tonatia saurophila, Trachypithecus auratus, Trachypithecus crepusculus, Trachypithecus cristatus, Trachypithecus francoisi, Trachypithecus geei, Trachypithecus germaini, Trachypithecus hatinhensis, Trachypithecus laotum, Trachypithecus leucocephalus, Trachypithecus melamera, Trachypithecus obscurus, Trachypithecus phayrei, Trachypithecus pileatus, Tragulus javanicus, Trichechus manatus, Tupaia tana, Tursiops truncatus, Uropsilus gracilis, Ursus maritimus, Varecia rubra, Varecia variegata, Vicugna pacos, Vulpes lagopus, Xerus inauris, Zalophus californianus, Zapus hudsonius, Ziphius cavirostris\
\ This alignment was created by making three edits (using Cactus) to the\ 241-way mammalian Zoonomia Cactus alignment\ (\ https://cglgenomics.ucsc.edu/data/cactus/).\
\
phyloP scores were computed from the Cactus 447-way alignment using the\
phyloP program from the\
PHAST package.\
Per-base scores were produced with options\
--method LRT --mode CONACC --wig-scores; positive scores\
indicate conservation under purifying selection, negative scores indicate\
acceleration relative to neutral evolution.\
\
For the all-species tracks, base-composition and substitution-rate\
parameters were estimated from 4-fold degenerate sites using\
phyloFit (PHAST, EM algorithm, medium precision) under either the\
REV or strand-symmetric reversible (SSREV) substitution model. Background\
base frequencies were adjusted with modFreqs so that\
complementary bases (A/T and C/G) appear at equal expected frequencies,\
which is required for strand-symmetric scoring.\
\
For the primates-subset tracks, the alignment was restricted to the 233\
primate species and an independent phyloFit / phyloP run was performed on\
that sub-alignment using the SSREV model. All scores were encoded into\
wiggle format and loaded as either bigWig files (REV all-species,\
primates LRT) or wig SQL tables backed by .wib data files\
(SSREV all-species, SSREV primates).\
\ The phylogenic tree was established by the research described\ in A global catalog of whole-genome diversity from 233 primate\ species.\ \
\
\\ \ \\
\ count \common \
nameclade \scientific name \
(link to browser when existing)taxon id \
link to NCBI\ 001 human primates catarrhini Homo sapiens/hg38
reference species9606 \ 002 western gorilla primates catarrhini Gorilla gorilla
GCA_900006655.3_Susie39593 \ 003 Sumatran orangutan primates catarrhini Pongo abelii
GCA_002880775.3_Susie_PABv29601 \ 004 Eastern Gorilla primates catarrhini Gorilla beringei 499232 \ 005 chimpanzee primates catarrhini Pan troglodytes
GCA_002880755.3_Clint_PTRv29598 \ 006 Bornean orangutan primates catarrhini Pongo pygmaeus 9600 \ 007 Rhesus monkey primates catarrhini Macaca mulatta
rheMac109544 \ 008 gelada primates catarrhini Theropithecus gelada
GCF_003255815.1_Tgel_1.09565 \ 009 stump-tailed macaque primates catarrhini Macaca arctoides 9540 \ 010 Northern Talapoin Monkey primates catarrhini Miopithecus ogouensis 100488 \ 011 crab-eating macaque primates catarrhini Macaca fascicularis 9541 \ 012 Allen's swamp monkey primates catarrhini Allenopithecus nigroviridis 54135 \ 013 siamang primates catarrhini Symphalangus syndactylus 9590 \ 014 black crested mangabey primates catarrhini Lophocebus aterrimus 75566 \ 015 drill primates catarrhini Mandrillus leucophaeus 9568 \ 016 Bonnet Macaque primates catarrhini Macaca radiata 9548 \ 017 Red-capped Mangabey primates catarrhini Cercocebus torquatus 9530 \ 018 Golden-bellied Mangabey primates catarrhini Cercocebus chrysogaster 75569 \ 019 Owl-faced Monkey primates catarrhini Cercopithecus hamlyni 9536 \ 020 Siberut Macaque primates catarrhini Macaca siberu 244255 \ 021 pig-tailed macaque primates catarrhini Macaca nemestrina
GCF_000956065.1_Mnem_1.09545 \ 022 White-naped Mangabey primates catarrhini Cercocebus lunulatus (Cercocebus atys lunulatus) 75570 \ 023 Tonkean Macaque primates catarrhini Macaca tonkeana 40843 \ 024 Diana Monkey primates catarrhini Cercopithecus diana 36224 \ 025 red guenon primates catarrhini Erythrocebus patas 9538 \ 026 Northern Pig-tailed Macaque primates catarrhini Macaca leonina 90387 \ 027 Moor Macaque primates catarrhini Macaca maura 90383 \ 028 Guinea Baboon primates catarrhini Papio papio 100937 \ 029 hamadryas baboon primates catarrhini Papio hamadryas 9557 \ 030 liontail macaque primates catarrhini Macaca silenus 54601 \ 031 olive baboon primates catarrhini Papio anubis
GCA_000264685.2_Panu_3.09555 \ 032 Roloway Monkey primates catarrhini Cercopithecus roloway 1137049 \ 033 Kinda Baboon primates catarrhini Papio kindae 208091 \ 034 Chacma Baboon primates catarrhini Papio ursinus 36229 \ 035 Sun-tailed Monkey primates catarrhini Allochrocebus solatus 147650 \ 036 golden snub-nosed monkey primates catarrhini Rhinopithecus roxellana
GCF_007565055.1_ASM756505v161622 \ 037 Vervet Monkey primates catarrhini Chlorocebus pygerythrus 60710 \ 038 sooty mangabey primates catarrhini Cercocebus atys
GCF_000955945.1_Caty_1.09531 \ 039 green monkey primates catarrhini Chlorocebus sabaeus
GCA_000409795.2_Chlorocebus_sabeus_1.160711 \ 040 De Brazza's monkey primates catarrhini Cercopithecus neglectus 36227 \ 041 Yellow Baboon primates catarrhini Papio cynocephalus 9556 \ 042 Celebes crested macaque primates catarrhini Macaca nigra 54600 \ 043 proboscis monkey primates catarrhini Nasalis larvatus 43780 \ 044 Preuss's Monkey primates catarrhini Allochrocebus preussi 147649 \ 045 Putty-nosed Monkey primates catarrhini Cercopithecus nictitans 36228 \ 046 Javan Surili primates catarrhini Presbytis comata 78452 \ 047 Sykes' Monkey primates catarrhini Cercopithecus albogularis 36225 \ 048 LHoests Monkey primates catarrhini Allochrocebus lhoesti 100224 \ 049 Crowned Monkey primates catarrhini Cercopithecus pogonias 102108 \ 050 Southern Mitered Langur primates catarrhini Presbytis mitrata (Presbytis melalophos mitrata) 272115 \ 051 Grey-shanked Douc Langur primates catarrhini Pygathrix cinerea 693712 \ 052 Mona monkey primates catarrhini Cercopithecus mona 36226 \ 053 Spot-nosed Monkey primates catarrhini Cercopithecus petaurista 100487 \ 054 grivet primates catarrhini Chlorocebus aethiops 9534 \ 055 Lowes Monkey primates catarrhini Cercopithecus lowei 304410 \ 056 Northern Yellow-cheeked Crested Gibbon primates catarrhini Nomascus annamensis 1616038 \ 057 Red-cheeked Gibbon primates catarrhini Nomascus gabriellae 61852 \ 058 Japanese macaque primates catarrhini Macaca fuscata 9542 \ 059 Western Red Colobus primates catarrhini Piliocolobus badius 164648 \ 060 southern white-cheeked gibbon primates catarrhini Nomascus siki_a 9586 \ 061 Taiwan macaque primates catarrhini Macaca cyclopis 78449 \ 062 black-shanked douc langur primates catarrhini Pygathrix nigripes 310352 \ 063 King Colobus primates catarrhini Colobus polykomos 9572 \ 064 Black Crested Gibbon primates catarrhini Nomascus concolor 29089 \ 065 Udzungwa Red Colobus primates catarrhini Piliocolobus gordonorum 591933 \ 066 Gee's Golden Langur primates catarrhini Trachypithecus geei 164650 \ 067 Kloss's Gibbon primates catarrhini Hylobates klossii 9587 \ 068 Spectacled Leaf Monkey primates catarrhini Trachypithecus obscurus 54181 \ 069 Zanzibar Red Colobus primates catarrhini Piliocolobus kirkii 591937 \ 070 Indochinese Silvered Langur primates catarrhini Trachypithecus germaini 271260 \ 071 Hatinh Langur primates catarrhini Trachypithecus hatinhensis 867383 \ 072 Moustached Monkey primates catarrhini Cercopithecus cephus 9535 \ 073 Laotian Langur primates catarrhini Trachypithecus laotum 465718 \ 074 Francois's langur primates catarrhini Trachypithecus francoisi 54180 \ 075 Purple-faced Langur primates catarrhini Semnopithecus vetulus (Trachypithecus vetulus) 54137 \ 076 Capped Langur primates catarrhini Trachypithecus pileatus 164651 \ 077 Ugandan red Colobus primates catarrhini Piliocolobus tephrosceles
GCF_002776525.2_ASM277652v2591936 \ 078 Spangled Ebony Langur primates catarrhini Trachypithecus auratus 222416 \ 079 Red-tailed Monkey primates catarrhini Cercopithecus ascanius 36223 \ 080 Silvery Lutung primates catarrhini Trachypithecus cristatus 122765 \ 081 Nilgiri Langur primates catarrhini Semnopithecus johnii (Trachypithecus johnii) 66063 \ 082 Indochinese grey langur primates catarrhini Trachypithecus crepusculus (Trachypithecus phayrei crepuscula) 272121 \ 083 White-headed langur primates catarrhini Trachypithecus leucocephalus (Trachypithecus poliocephalus) 465719 \ 084 pygmy chimpanzee primates catarrhini Pan paniscus
GCA_000258655.2_panpan1.19597 \ 085 northern white-cheeked gibbon primates catarrhini Nomascus siki_b 9586 \ 086 Agile Gibbon primates catarrhini Hylobates agilis 9579 \ 087 Phayre's Leaf-monkey primates catarrhini Trachypithecus melamera n/a \ 088 Nepal Gray Langur primates catarrhini Semnopithecus schistaceus 2804203 \ 089 Abbott's Gray Gibbon primates catarrhini Hylobates abbotti (Hylobates muelleri abbotti) 716694 \ 090 Bornean Gibbon primates catarrhini Hylobates muelleri 9588 \ 091 Tufted Gray Langur primates catarrhini Semnopithecus priam 1208733 \ 092 Black-footed Gray Langur primates catarrhini Semnopithecus hypoleucos 1208734 \ 093 mantled guereza primates catarrhini Colobus guereza 33548 \ 094 Hanuman langur primates catarrhini Semnopithecus entellus 88029 \ 095 pileated gibbon primates catarrhini Hylobates pileatus 9589 \ 096 black snub-nosed monkey primates catarrhini Rhinopithecus bieti 61621 \ 097 Burmese snub-nosed monkey primates catarrhini Rhinopithecus strykeri 1194336 \ 098 Angolan colobus primates catarrhini Colobus angolensis
colAng154131 \ 099 Pileated Gibbon primates catarrhini Hylobates pileatus 9589 \ 100 black-shanked douc langur primates catarrhini Pygathrix nigripes 310352 \ 101 Milne-edwards' Macaque primates catarrhini Macaca thibetana 54602 \ 102 Phayre's Leaf-monkey primates catarrhini Trachypithecus phayrei 61618 \ 103 Assam macaque primates catarrhini Macaca assamensis 9551 \ 104 Eastern hoolock gibbon primates catarrhini Hoolock leuconedys 61851 \ 105 mandrill primates catarrhini Mandrillus sphinx 9561 \ 106 White-faced Saki primates platyrrhini Pithecia chrysocephala 2946515 \ 107 Monk Saki primates platyrrhini Pithecia hirsuta 2946516 \ 108 white-faced saki primates platyrrhini Pithecia pithecia 43777 \ 109 Mittermeier's Tapajós saki primates platyrrhini Pithecia mittermeieri 2946517 \ 110 Buffy Saki primates platyrrhini Pithecia albicans 2946514 \ 111 Pissinatti's saki primates platyrrhini Pithecia pissinattii (Pithecia pissinatti) 2946518 \ 112 Vanzolini's Bald-faced Saki primates platyrrhini Pithecia vanzolinii 2946519 \ 113 Bald-headed Uacari primates platyrrhini Cacajao calvus 30596 \ 114 Ayres Black Uakari primates platyrrhini Cacajao ayresi 535896 \ 115 Black-headed Uacari primates platyrrhini Cacajao melanocephalus 70825 \ 116 Black-headed Uacari primates platyrrhini Cacajao hosomi 535897 \ 117 Reddish-brown bearded saki primates platyrrhini Chiropotes sagulatus (Chiropotes chiropotes) 658221 \ 118 brown-backed bearded saki primates platyrrhini Chiropotes israelita 280163 \ 119 Collared Titi Monkey primates platyrrhini Cheracebus lugens 210166 \ 120 Brown Titi Monkey primates platyrrhini Plecturocebus brunneus 1812042 \ 121 Hoffmanns's titi monkey primates platyrrhini Plecturocebus hoffmannsi 78255 \ 122 Milton's Titi Monkey primates platyrrhini Plecturocebus miltoni 1812038 \ 123 Widow Monkey primates platyrrhini Cheracebus torquatus 30592 \ 124 Ashy Black Titi Monkey primates platyrrhini Plecturocebus cinerascens 1812037 \ 125 Prince Bernhard's Titi Monkey primates platyrrhini Plecturocebus bernhardi 1812036 \ 126 Yellow-handed Titi Monkey primates platyrrhini Cheracebus lucifer 2487712 \ 127 Coppery Titi Monkey primates platyrrhini Plecturocebus cupreus 202457 \ 128 Chestnut-bellied Titi primates platyrrhini Plecturocebus caligatus 867332 \ 129 Hershkovitzs Titi primates platyrrhini Plecturocebus dubius 2946520 \ 130 Red-bellied Titi Monkey primates platyrrhini Plecturocebus moloch 9523 \ 131 Groves' Titi primates platyrrhini Plecturocebus grovesi 2488670 \ 132 black-handed spider monkey primates platyrrhini Ateles geoffroyi_a 9509 \ 133 Widow Monkey primates platyrrhini Cheracebus regulus 1812110 \ 134 Guiana Spider Monkey primates platyrrhini Ateles paniscus 9510 \ 135 Black-faced Black Spider Monkey primates platyrrhini Ateles chamek 118643 \ 136 White-cheeked Spider Monkey primates platyrrhini Ateles marginatus 1529884 \ 137 White-bellied Spider Monkey primates platyrrhini Ateles belzebuth 9507 \ 138 Common Woolly Monkey primates platyrrhini Lagothrix lagothricha (Lagothrix lagotricha) 9519 \ 139 large-headed capuchin primates platyrrhini Sapajus macrocephalus (Sapajus apella macrocephalus) 1547595 \ 140 Spixs White-fronted Capuchin primates platyrrhini Cebus unicolor 1985288 \ 141 Central American spider monkey primates platyrrhini Ateles geoffroyi_b 9509 \ 142 Guinan Weeper Capuchin primates platyrrhini Cebus olivaceus 37295 \ 143 mantled howler monkey primates platyrrhini Alouatta palliata 30589 \ 144 white-fronted capuchin primates platyrrhini Cebus albifrons 9514 \ 145 Northern Night Monkey primates platyrrhini Aotus trivirgatus 9505 \ 146 Grey-handed Night Monkey primates platyrrhini Aotus griseimembra 292213 \ 147 Black-and-gold Howler Monkey primates platyrrhini Alouatta caraya 9502 \ 148 Spixs Night Monkey primates platyrrhini Aotus vociferans 57176 \ 149 Red-handed Howler Monkey primates platyrrhini Alouatta belzebul 30590 \ 150 Red-handed Howler Monkey primates platyrrhini Alouatta discolor 2905217 \ 151 Azara's Night Monkey primates platyrrhini Aotus azarae (Aotus azarai) 30591 \ 152 Purús Red Howler Monkey primates platyrrhini Alouatta puruensis (Alouatta seniculus puruensis) 1347729 \ 153 Black Howler Monkey primates platyrrhini Alouatta nigerrima (Alouatta belzebul) 30590 \ 154 Guianan Red Howler Monkey primates platyrrhini Alouatta macconnelli 198115 \ 155 Colombian Red Howler Monkey primates platyrrhini Alouatta juara 2946512 \ 156 Colombian Red Howler Monkey primates platyrrhini Alouatta seniculus 9503 \ 157 tufted capuchin primates platyrrhini Sapajus apella 9515 \ 158 Ma's night monkey primates platyrrhini Aotus nancymaae
GCA_000952055.2_Anan_2.037293 \ 159 Bolivian squirrel monkey primates platyrrhini Saimiri boliviensis
GCF_016699345.1_BCM_Sbol_2.027679 \ 160 White-nosed Saki primates platyrrhini Chiropotes albinasus 198627 \ 161 Black Mantle Tamarin primates platyrrhini Leontocebus nigricollis 9489 \ 162 brown-mantled tamarin primates platyrrhini Leontocebus fuscicollis 9487 \ 163 Illiger's saddle-back tamarin primates platyrrhini Leontocebus illigeri (Leontocebus fuscicollis illigeri) 881947 \ 164 Cotton-headed Tamarin primates platyrrhini Saguinus oedipus 9490 \ 165 Pied Tamarin primates platyrrhini Saguinus bicolor 37588 \ 166 Geoffroy's Tamarin primates platyrrhini Saguinus geoffroyi 43778 \ 167 White-fronted Titi Monkey primates platyrrhini Saguinus inustus 1079039 \ 168 Moustached Tamarin primates platyrrhini Saguinus mystax 9488 \ 169 tamarin primates platyrrhini Saguinus imperator 9491 \ 170 Guianan Squirrel Monkey primates platyrrhini Saimiri sciureus 9521 \ 171 Red-chested Mustached Tamarin primates platyrrhini Saguinus labiatus 78454 \ 172 Goeldi's Monkey primates platyrrhini Callimico goeldii 9495 \ 173 Black-crowned Central American Squirrel Monkey primates platyrrhini Saimiri oerstedii 70928 \ 174 Golden-headed Lion Tamarin primates platyrrhini Leontopithecus chrysomelas 57374 \ 175 golden lion tamarin primates platyrrhini Leontopithecus rosalia 30588 \ 176 Humboldt's Squirrel Monkey primates platyrrhini Saimiri cassiquiarensis 2946521 \ 177 bare-eared squirrel monkey primates platyrrhini Saimiri ustus 66265 \ 178 Ecuadorian squirrel monkey primates platyrrhini Saimiri macrodon 2946522 \ 179 white-tufted-ear marmoset primates platyrrhini Callithrix jacchus 9483 \ 180 Eastern Pygmy Marmoset primates platyrrhini Cebuella niveiventris 2826950 \ 181 Western Pygmy Marmoset primates platyrrhini Cebuella pygmaea 9493 \ 182 Black And White Tassel-ear Marmoset primates platyrrhini Mico humeralifer 52232 \ 183 Black-crowned Dwarf Marmoset primates platyrrhini Callibella humilis (Mico humilis) 666519 \ 184 Mico schneideri primates platyrrhini Mico schneideri n/a \ 185 Silvery Marmoset primates platyrrhini Mico argentatus 9482 \ 186 Midas tamarin primates platyrrhini Saguinus midas 30586 \ 187 Wieds Marmoset primates platyrrhini Callithrix kuhlii 867363 \ 188 Geoffroy's Tufted-ear Marmoset primates platyrrhini Callithrix geoffroyi 52231 \ 189 Horsfield's tarsier primates tarsiidae Cephalopachus bancanus 9477 \ 190 Philippine tarsier primates tarsiidae Carlito syrichta
tarSyr21868482 \ 191 Lariang Tarsier primates tarsiidae Tarsius lariang 630277 \ 192 Wallace's Tarsier primates tarsiidae Tarsius wallacei 981131 \ 193 aye-aye primates strepsirrhini Daubentonia madagascariensis 31869 \ 194 Crowned Sifaka primates strepsirrhini Propithecus coronatus (Propithecus deckenii coronatus) 475619 \ 195 Perrier's Sifaka primates strepsirrhini Propithecus perrieri 989338 \ 196 ruffed lemur primates strepsirrhini Varecia variegata 9455 \ 197 Diademed Sifaka primates strepsirrhini Propithecus diadema 83281 \ 198 Milne-Edwards Sifaka primates strepsirrhini Propithecus edwardsi 543559 \ 199 babakoto primates strepsirrhini Indri indri 34827 \ 200 Golden-crowned Sifaka primates strepsirrhini Propithecus tattersalli 30601 \ 201 Eastern Woolly Lemur primates strepsirrhini Avahi laniger 122246 \ 202 Verreauxs Sifaka primates strepsirrhini Propithecus verreauxi 34825 \ 203 Peyrieras Woolly Lemur primates strepsirrhini Avahi peyrierasi 1313323 \ 204 Red Ruffed Lemur primates strepsirrhini Varecia rubra 554167 \ 205 greater bamboo lemur primates strepsirrhini Prolemur simus 1328070 \ 206 Red-bellied Lemur primates strepsirrhini Eulemur rubriventer 34829 \ 207 mongoose lemur primates strepsirrhini Eulemur mongoz 34828 \ 208 Geoffroys Dwarf Lemur primates strepsirrhini Cheirogaleus major 47177 \ 209 Crowned Lemur primates strepsirrhini Eulemur coronatus 13514 \ 210 black lemur primates strepsirrhini Eulemur macaco 30602 \ 211 lesser dwarf lemur primates strepsirrhini Cheirogaleus medius 9460 \ 212 Sclater's lemur primates strepsirrhini Eulemur flavifrons 87288 \ 213 Coquerel's sifaka primates strepsirrhini Propithecus coquerelli (Propithecus coquereli)
proCoq1379532 \ 214 Collared Brown Lemur primates strepsirrhini Eulemur collaris (Eulemur fulvus collaris) 47178 \ 215 Red-tailed Sportive Lemur primates strepsirrhini Lepilemur ruficaudatus 78866 \ 216 Red Brown Lemur primates strepsirrhini Eulemur rufus 859983 \ 217 Sanfords Brown Lemur primates strepsirrhini Eulemur sanfordi 122225 \ 218 White-fronted Lemur primates strepsirrhini Eulemur albifrons 1215604 \ 219 Gray's Sportive Lemur primates strepsirrhini Lepilemur dorsalis 78583 \ 220 brown lemur primates strepsirrhini Eulemur fulvus 13515 \ 221 Sahafary Sportive Lemur primates strepsirrhini Lepilemur septentrionalis 78584 \ 222 Sambirano Lesser Bamboo Lemur primates strepsirrhini Hapalemur occidentalis 867377 \ 223 Alaotra Reed Lemur primates strepsirrhini Hapalemur alaotrensis (Hapalemur griseus alaotrensis) 122220 \ 224 Eastern Lesser Bamboo Lemur primates strepsirrhini Hapalemur griseus 13557 \ 225 Ankarana Sportive Lemur primates strepsirrhini Lepilemur ankaranensis 342401 \ 226 ring-tailed lemur primates strepsirrhini Lemur catta 9447 \ 227 gray bamboo lemur primates strepsirrhini Hapalemur gilberti 3043110 \ 228 Rusty-gray Lesser Bamboo Lemur primates strepsirrhini Hapalemur meridionalis 3043112 \ 229 Demidoffs Dwarf Galago primates strepsirrhini Galagoides demidoff 89672 \ 230 northern giant mouse lemur primates strepsirrhini Mirza zaza 339999 \ 231 gray mouse lemur primates strepsirrhini Microcebus murinus
GCA_000165445.3_Mmur_3.030608 \ 232 small-eared galago primates strepsirrhini Otolemur garnettii
otoGar330611 \ 233 Northern Lesser Galago primates strepsirrhini Galago senegalensis 9465 \ 234 Thick-tailed Greater Galago primates strepsirrhini Otolemur crassicaudatus 9463 \ 235 Grey Slender Loris primates strepsirrhini Loris lydekkerianus 300163 \ 236 slender loris primates strepsirrhini Loris tardigradus 9468 \ 237 West African Potto primates strepsirrhini Perodicticus potto 9472 \ 238 East African Potto primates strepsirrhini Perodicticus ibeanus (Perodicticus potto ibeanus) 261737 \ 239 Moholi bushbaby primates strepsirrhini Galago moholi 30609 \ 240 Pygmy Slow Loris primates strepsirrhini Nycticebus pygmaeus (Xanthonycticebus pygmaeus) 101278 \ 241 Bengal slow loris primates strepsirrhini Nycticebus bengalensis 261741 \ 242 Calabar Angwantibo primates strepsirrhini Arctocebus calabarensis 261739 \ 243 slow loris primates strepsirrhini Nycticebus coucang 9470 \ 244 jaguar carnivora Panthera onca
GCA_004023805.1_PanOnc_v1_BIUU9690 \ 245 leopard carnivora Panthera pardus
GCA_001857705.1_PanPar1.09691 \ 246 giant panda carnivora Ailuropoda melanoleuca
GCA_002007445.1_ASM200744v19646 \ 247 Hawaiian monk seal carnivora Neomonachus schauinslandi
GCA_002201575.1_ASM220157v129088 \ 248 California sea lion carnivora Zalophus californianus
GCA_004024565.1_ZalCal_v1_BIUU9704 \ 249 Greenland wolf carnivora Canis lupus orion
GCA_905319855.2_mCanLor1.22605939 \ 250 Pacific walrus carnivora Odobenus rosmarus
odoRosDiv19707 \ 251 domestic cat (Fca126) carnivora Felis catus fca126 (Felis catus)
GCF_018350175.1_F.catus_Fca126_mat1.09685 \ 252 northern elephant seal carnivora Mirounga angustirostris
GCA_004023865.1_MirAng_v1_BIUU9716 \ 253 domestic cat carnivora Felis catus
felCat89685 \ 254 domestic dog (BS72/Village Dog) carnivora Canis lupus familiaris
GCA_004027395.1_CanFam_VD_v1_BIUU\ 255 German Shepherd dog (Mischka) carnivora Canis lupus familiaris (CanFam4) (Canis lupus familiaris)
canFam4\ 256 dingo carnivora Canis lupus dingo 286419 \ 257 raccoon dog carnivora Nyctereutes procyonoides 34880 \ 258 fossa carnivora Cryptoprocta ferox 94188 \ 259 polar bear carnivora Ursus maritimus
GCA_000687225.1_UrsMar_1.029073 \ 260 Asian palm civet carnivora Paradoxurus hermaphroditus
GCA_004024585.1_ParHer_v1_BIUU71117 \ 261 African hunting dog carnivora Lycaon pictus
GCA_001887905.1_LycPicSAfr1.09622 \ 262 Arctic fox carnivora Vulpes lagopus
GCA_004023825.1_VulLag_v1_BIUU494514 \ 263 dog carnivora Canis lupus familiaris
GCF_000002285.3_CanFam3.19615 \ 264 striped hyena carnivora Hyaena hyaena
GCA_004023945.1_HyaHya_v1_BIUU95912 \ 265 n/a carnivora Acinonyx jubatus
GCA_001443585.1_aciJub132536 \ 266 tiger carnivora Panthera tigris
GCA_000464555.1_PanTig1.09694 \ 267 Sea otter carnivora Enhydra lutris
GCA_002288905.2_ASM228890v234882 \ 268 giant otter carnivora Pteronura brasiliensis 9672 \ 269 bat-eared fox carnivora Otocyon megalotis 9624 \ 270 Weddell seal carnivora Leptonychotes weddellii
GCA_000349705.1_LepWed1.09713 \ 271 Lesser panda carnivora Ailurus fulgens
GCA_002007465.1_ASM200746v19649 \ 272 ratel carnivora Mellivora capensis
GCA_004024625.1_MelCap_v1_BIUU9664 \ 273 banded mongoose carnivora Mungos mungo
GCA_004023785.1_MunMun_v1_BIUU210652 \ 274 dwarf mongoose carnivora Helogale parvula
GCA_004023845.1_HelPar_v1_BIUU210647 \ 275 meerkat carnivora Suricata suricatta
GCA_004023905.1_SurSur_v1_BIUU37032 \ 276 puma carnivora Puma concolor
GCA_003327715.1_PumCon1.09696 \ 277 black-footed cat carnivora Felis nigripes
GCA_004023925.1_FelNig_v1_BIUU61379 \ 278 European polecat carnivora Mustela putorius
GCA_000239315.1_MusPutFurMale1.09668 \ 279 western spotted skunk carnivora Spilogale gracilis
GCA_004023965.1_SpiGra_v1_BIUU30551 \ 280 Sumatran rhinoceros laurasiatheria Dicerorhinus sumatrensis
GCA_002844835.1_ASM284483v189632 \ 281 black rhinoceros laurasiatheria Diceros bicornis
GCA_004027315.1_DicBicMic_v1_BIUU9805 \ 282 Asiatic tapir laurasiatheria Tapirus indicus
GCA_004024905.1_TapInd_v1_BIUU9802 \ 283 Brazilian tapir laurasiatheria Tapirus terrestris
GCA_004025025.1_TapTer_v1_BIUU9801 \ 284 northern white rhinoceros laurasiatheria Ceratotherium simum cottoni 310713 \ 285 ass laurasiatheria Equus asinus
GCA_001305755.1_ASM130575v19793 \ 286 Southern white rhinoceros laurasiatheria Ceratotherium simum
GCA_000283155.1_CerSimSim1.09807 \ 287 Przewalski's horse laurasiatheria Equus przewalskii
GCA_000696695.1_Burgud9798 \ 288 horse laurasiatheria Equus caballus
GCA_000002305.1_EquCab2.09796 \ 289 Malayan pangolin laurasiatheria Manis javanica
GCA_001685135.1_ManJav1.09974 \ 290 Chinese pangolin laurasiatheria Manis pentadactyla
GCA_000738955.1_M_pentadactyla-1.1.1143292 \ 291 Hispaniolan solenodon laurasiatheria Solenodon paradoxus 79805 \ 292 eastern mole laurasiatheria Scalopus aquaticus
GCA_004024925.1_ScaAqu_v1_BIUU71119 \ 293 gracile shrew mole laurasiatheria Uropsilus gracilis
GCA_004024945.1_UroGra_v1_BIUU182669 \ 294 star-nosed mole laurasiatheria Condylura cristata
GCF_000260355.1_ConCri1.0143302 \ 295 western European hedgehog laurasiatheria Erinaceus europaeus
GCA_000296755.1_EriEur2.09365 \ 296 European shrew laurasiatheria Sorex araneus
sorAra242254 \ 297 Indochinese shrew laurasiatheria Crocidura indochinensis
GCA_004027635.1_CroInd_v1_BIUU876679 \ 298 Hoffmann's two-fingered sloth xenarthra Choloepus hoffmanni
GCA_000164785.2_C_hoffmanni-2.0.19358 \ 299 nine-banded armadillo xenarthra Dasypus novemcinctus
GCA_000208655.2_Dasnov3.09361 \ 300 giant anteater xenarthra Myrmecophaga tridactyla
GCA_004026745.1_MyrTri_v1_BIUU71006 \ 301 southern tamandua xenarthra Tamandua tetradactyla
GCA_004025105.1_TamTet_v1_BIUU48850 \ 302 placentals xenarthra Tolypeutes matacus 183749 \ 303 southern two-toed sloth xenarthra Choloepus didactylus
GCA_004027855.1_ChoDid_v1_BIUU27675 \ 304 screaming hairy armadillo xenarthra Chaetophractus vellerosus
GCA_004027955.1_ChaVel_v1_BIUU340076 \ 305 North Pacific right whale artiodactyla Eubalaena japonica 302098 \ 306 grey whale artiodactyla Eschrichtius robustus 9764 \ 307 hippopotamus artiodactyla Hippopotamus amphibius
GCA_004027065.1_HipAmp_v1_BIUU9833 \ 308 Minke whale artiodactyla Balaenoptera acutorostrata
GCA_000493695.1_BalAcu1.09767 \ 309 beluga whale artiodactyla Delphinapterus leucas
GCA_002288925.2_ASM228892v29749 \ 310 Antarctic minke whale artiodactyla Balaenoptera bonaerensis
GCA_000978805.1_ASM97880v133556 \ 311 boutu artiodactyla Inia geoffrensis 9725 \ 312 harbor porpoise artiodactyla Phocoena phocoena 9742 \ 313 narwhal artiodactyla Monodon monoceros
GCA_004026685.1_MonMon_M_v1_BIUU40151 \ 314 Yangtze River dolphin artiodactyla Lipotes vexillifer
GCA_000442215.1_Lipotes_vexillifer_v1118797 \ 315 killer whale artiodactyla Orcinus orca
orcOrc19733 \ 316 Ganges River dolphin artiodactyla Platanista gangetica 118798 \ 317 Yangtze finless porpoise artiodactyla Neophocaena asiaeorientalis
GCA_003031525.1_Neophocaena_asiaeorientalis_V1189058 \ 318 Sowerby's beaked whale artiodactyla Mesoplodon bidens 48745 \ 319 alpaca artiodactyla Vicugna pacos
GCA_000767525.1_Vi_pacos_V1.030538 \ 320 Cuvier's beaked whale" artiodactyla Ziphius cavirostris 9760 \ 321 Bactrian camel artiodactyla Camelus bactrianus
GCA_000767855.1_Ca_bactrianus_MBC_1.09837 \ 322 Arabian camel artiodactyla Camelus dromedarius
GCA_000767585.1_PRJNA234474_Ca_dromedarius_V1.09838 \ 323 wild Bactrian camel artiodactyla Camelus ferus
GCA_000311805.2_CB1419612 \ 324 pygmy sperm whale artiodactyla Kogia breviceps 27615 \ 325 Chacoan peccary artiodactyla Catagonus wagneri
GCA_004024745.1_CatWag_v1_BIUU51154 \ 326 reindeer artiodactyla Rangifer tarandus
GCA_004026565.1_RanTarSib_v1_BIUU9870 \ 327 Pere David's deer artiodactyla Elaphurus davidianus
GCA_002443075.1_Milu1.043332 \ 328 okapi artiodactyla Okapia johnstoni
GCA_001660835.1_ASM166083v186973 \ 329 Masai giraffe artiodactyla Giraffa tippelskirchi
GCA_001651235.1_ASM165123v1439328 \ 330 Siberian musk deer artiodactyla Moschus moschiferus
GCA_004024705.1_MosMos_v1_BIUU68415 \ 331 water buffalo artiodactyla Bubalus bubalis
GCA_000471725.1_UMD_CASPUR_WB_2.089462 \ 332 cow artiodactyla Bos taurus
GCA_000003205.6_Btau_5.0.19913 \ 333 pronghorn artiodactyla Antilocapra americana
GCA_004027515.1_AntAmePen_v1_BIUU9891 \ 334 white-tailed deer artiodactyla Odocoileus virginianus
GCA_002102435.1_Ovir.te_1.09874 \ 335 aoudad artiodactyla Ammotragus lervia
GCA_002201775.1_ALER1.09899 \ 336 bighorn sheep artiodactyla Ovis canadensis
GCA_004026945.1_OviCan_v1_BIUU37174 \ 337 goat artiodactyla Capra hircus
GCA_001704415.1_ARS19925 \ 338 Nilgiri tahr artiodactyla Hemitragus hylocrius
GCA_004026825.1_HemHyl_v1_BIUU330464 \ 339 hirola artiodactyla Beatragus hunteri
GCA_004027495.1_BeaHun_v1_BIUU59527 \ 340 wild yak artiodactyla Bos mutus
bosMut172004 \ 341 American bison artiodactyla Bison bison
GCA_000754665.1_Bison_UMD1.09901 \ 342 sheep artiodactyla Ovis aries
GCA_000298735.2_Oar_v4.09940 \ 343 chiru artiodactyla Pantholops hodgsonii
GCA_000400835.1_PHO1.059538 \ 344 wild goat artiodactyla Capra aegagrus
GCA_000978405.1_CapAeg_1.09923 \ 345 Java mouse-deer artiodactyla Tragulus javanicus
GCA_004024965.1_TraJav_v1_BIUU9849 \ 346 pig artiodactyla Sus scrofa
susScr39823 \ 347 zebu cattle artiodactyla Bos indicus
GCA_000247795.2_Bos_indicus_1.09915 \ 348 common bottlenose dolphin artiodactyla Tursiops truncatus
GCA_001922835.1_NIST_Tur_tru_v19739 \ 349 Saiga antelope artiodactyla Saiga tatarica
GCA_004024985.1_SaiTat_v1_BIUU34875 \ 350 Chinese rufous horseshoe bat chiroptera Rhinolophus sinicus
GCA_001888835.1_ASM188883v189399 \ 351 black flying fox chiroptera Pteropus alecto
pteAle19402 \ 352 Cantor's roundleaf bat chiroptera Hipposideros galeritus 58069 \ 353 Egyptian rousette chiroptera Rousettus aegyptiacus
GCA_004024865.1_RouAeg_v1_BIUU9407 \ 354 long-tongued fruit bat chiroptera Macroglossus sobrinus 326083 \ 355 large flying fox chiroptera Pteropus vampyrus
GCF_000151845.1_Pvam_2.0132908 \ 356 Brazilian free-tailed bat chiroptera Tadarida brasiliensis
GCA_004025005.1_TadBra_v1_BIUU9438 \ 357 great roundleaf bat chiroptera Hipposideros armiger
GCA_001890085.1_ASM189008v1186990 \ 358 straw-colored fruit bat chiroptera Eidolon helvum
eidHel177214 \ 359 Antillean ghost-faced bat chiroptera Mormoops blainvillei
GCA_004026545.1_MorMeg_v1_BIUU118852 \ 360 tailed tailless bat chiroptera Anoura caudifer
GCA_004027475.1_AnoCau_v1_BIUU27642 \ 361 common vampire bat chiroptera Desmodus rotundus
GCA_002940915.2_ASM294091v29430 \ 362 hairy big-eared bat chiroptera Micronycteris hirsuta
GCA_004026765.1_MicHir_v1_BIUU148065 \ 363 stripe-headed round-eared bat chiroptera Tonatia saurophila
GCA_004024845.1_TonSau_v1_BIUU171122 \ 364 Seba's short-tailed bat chiroptera Carollia perspicillata
GCA_004027735.1_CarPer_v1_BIUU40233 \ 365 Jamaican fruit-eating bat chiroptera Artibeus jamaicensis
GCA_004027435.1_ArtJam_v1_BIUU9417 \ 366 Indian false vampire chiroptera Megaderma lyra
GCA_004026885.1_MegLyr_v1_BIUU9413 \ 367 Schreibers' long-fingered bat chiroptera Miniopterus schreibersii
GCA_004026525.1_MinSch_v1_BIUU9433 \ 368 greater bulldog bat chiroptera Noctilio leporinus
GCA_004026585.1_NocLep_v1_BIUU94963 \ 369 Natal long-fingered bat chiroptera Miniopterus natalensis
GCF_001595765.1_Mnat.v1291302 \ 370 hog-nosed bat chiroptera Craseonycteris thonglongyai
GCA_004027555.1_CraTho_v1_BIUU208972 \ 371 Parnell's mustached bat chiroptera Pteronotus parnellii
ptePar159476 \ 372 greater mouse-eared bat chiroptera Myotis myotis
GCA_004026985.1_MyoMyo_v1_BIUU51298 \ 373 Ashy-gray tube-nosed bat chiroptera Murina feae (Murina aurata feae)
GCA_004026665.1_MurFea_v1_BIUU1453894 \ 374 David's myotis chiroptera Myotis davidii
myoDav1225400 \ 375 Brandt's bat chiroptera Myotis brandtii
myoBra1109478 \ 376 big brown bat chiroptera Eptesicus fuscus
GCF_000308155.1_EptFus1.029078 \ 377 red bat chiroptera Lasiurus borealis
GCA_004026805.1_LasBor_v1_BIUU258930 \ 378 little brown bat chiroptera Myotis lucifugus
myoLuc259463 \ 379 common pipistrelle chiroptera Pipistrellus pipistrellus
GCA_004026625.1_PipPip_v1_BIUU59474 \ 380 African savanna elephant afrotheria Loxodonta africana
GCA_000001905.1_Loxafr3.09785 \ 381 Florida manatee afrotheria Trichechus manatus
GCA_000243295.1_TriManLat1.09778 \ 382 yellow-spotted hyrax afrotheria Heterohyrax brucei
GCA_004026845.1_HetBruBak_v1_BIUU77598 \ 383 Cape rock hyrax afrotheria Procavia capensis
GCA_004026925.1_ProCapCap_v1_BIUU9813 \ 384 aardvark afrotheria Orycteropus afer 9818 \ 385 Cape golden mole afrotheria Chrysochloris asiatica
GCA_004027935.1_ChrAsi_v1_BIUU185453 \ 386 Cape elephant shrew afrotheria Elephantulus edwardii
eleEdw128737 \ 387 Talazac's shrew tenrec afrotheria Microgale talazaci (Nesogale talazaci)
GCA_004026705.1_MicTal_v1_BIUU2583312 \ 388 small Madagascar hedgehog afrotheria Echinops telfairi
GCA_000313985.1_EchTel2.09371 \ 389 Sunda flying lemur euarchontoglires Galeopterus variegatus
GCA_004027255.1_GalVar_v1_BIUU482537 \ 390 Chinese tree shrew euarchontoglires Tupaia chinensis
tupChi1246437 \ 391 South African ground squirrel euarchontoglires Xerus inauris
GCA_004024805.1_XerIna_v1_BIUU234690 \ 392 large tree shrew euarchontoglires Tupaia tana 70687 \ 393 mountain beaver euarchontoglires Aplodontia rufa
GCA_004027875.1_AplRuf_v1_BIUU51342 \ 394 Alpine marmot euarchontoglires Marmota marmota
GCF_001458135.1_marMar2.19993 \ 395 Daurian ground squirrel euarchontoglires Spermophilus dauricus
GCA_002406435.1_ASM240643v199837 \ 396 crested porcupine euarchontoglires Hystrix cristata
GCA_004026905.1_HysCri_v1_BIUU10137 \ 397 thirteen-lined ground squirrel euarchontoglires Ictidomys tridecemlineatus
speTri243179 \ 398 American beaver euarchontoglires Castor canadensis
GCA_004027675.1_CasCan_v1_BIUU51338 \ 399 long-tailed chinchilla euarchontoglires Chinchilla lanigera
chiLan134839 \ 400 punctate agouti euarchontoglires Dasyprocta punctata 34846 \ 401 pacarana euarchontoglires Dinomys branickii
GCA_004027595.1_DinBra_v1_BIUU108858 \ 402 fat dormouse euarchontoglires Glis glis
GCA_004027185.1_GliGli_v1_BIUU41261 \ 403 northern gundi euarchontoglires Ctenodactylus gundi
GCA_004027205.1_CteGun_v1_BIUU10166 \ 404 naked mole-rat euarchontoglires Heterocephalus glaber
GCA_000247695.1_HetGla_female_1.010181 \ 405 Patagonian cavy euarchontoglires Dolichotis patagonum
GCA_004027295.1_DolPat_v1_BIUU29091 \ 406 capybara euarchontoglires Hydrochoerus hydrochaeris
GCA_004027455.1_HydHyd_v1_BIUU10149 \ 407 Montane guinea pig euarchontoglires Cavia tschudii
GCA_004027695.1_CavTsc_v1_BIUU143287 \ 408 domestic guinea pig euarchontoglires Cavia porcellus
GCA_000151735.1_Cavpor3.010141 \ 409 degu euarchontoglires Octodon degus
GCA_000260255.1_OctDeg1.010160 \ 410 lowland paca euarchontoglires Cuniculus paca 108852 \ 411 social tuco-tuco euarchontoglires Ctenomys sociabilis
GCA_004027165.1_CteSoc_v1_BIUU43321 \ 412 Damara mole-rat euarchontoglires Fukomys damarensis
fukDam1885580 \ 413 woodland dormouse euarchontoglires Graphiurus murinus 51346 \ 414 Desmarest's hutia euarchontoglires Capromys pilorides
GCA_004027915.1_CapPil_v1_BIUU34842 \ 415 Upper Galilee mountains blind mole rat euarchontoglires Nannospalax galili
GCA_000622305.1_S.galili_v1.01026970 \ 416 nutria euarchontoglires Myocastor coypus
GCA_004027025.1_MyoCoy_v1_BIUU10157 \ 417 hazel dormouse euarchontoglires Muscardinus avellanarius
GCA_004027005.1_MusAve_v1_BIUU39082 \ 418 dassie-rat euarchontoglires Petromus typicus
GCA_004026965.1_PetTyp_v1_BIUU10183 \ 419 greater cane rat euarchontoglires Thryonomys swinderianus
GCA_004025085.1_ThrSwi_v1_BIUU10169 \ 420 snowshoe hare euarchontoglires Lepus americanus
GCA_004026855.1_LepAme_v1_BIUU48086 \ 421 Gambian giant pouched rat euarchontoglires Cricetomys gambianus
GCA_004027575.1_CriGam_v1_BIUU10085 \ 422 Prairie deer mouse euarchontoglires Peromyscus maniculatus
GCF_000500345.1_Pman_1.010042 \ 423 southern grasshopper mouse euarchontoglires Onychomys torridus
GCA_004026725.1_OnyTor_v1_BIUU38674 \ 424 rabbit euarchontoglires Oryctolagus cuniculus
GCA_000003625.1_OryCun2.09986 \ 425 muskrat euarchontoglires Ondatra zibethicus
GCA_004026605.1_OndZib_v1_BIUU10060 \ 426 northern mole vole euarchontoglires Ellobius talpinus
GCA_001685095.1_ETalpinus_0.1329620 \ 427 Mongolian gerbil euarchontoglires Meriones unguiculatus
GCA_004026785.1_MerUng_v1_BIUU10047 \ 428 fat sand rat euarchontoglires Psammomys obesus
GCA_002215935.1_ASM221593v148139 \ 429 house mouse euarchontoglires Mus musculus
mm1010090 \ 430 Chinese hamster euarchontoglires Cricetulus griseus
GCA_900186095.1_CHOK1S_HZDv110029 \ 431 Norway rat euarchontoglires Rattus norvegicus
GCF_000001895.5_Rnor_6.010116 \ 432 western wild mouse euarchontoglires Mus spretus
GCA_001624865.1_SPRET_EiJ_v110096 \ 433 meadow jumping mouse euarchontoglires Zapus hudsonius
GCA_004024765.1_ZapHud_v1_BIUU160400 \ 434 prairie vole euarchontoglires Microtus ochrogaster
micOch179684 \ 435 Ryukyu mouse euarchontoglires Mus caroli
GCA_900094665.2_CAROLI_EIJ_v1.110089 \ 436 Egyptian spiny mouse euarchontoglires Acomys cahirinus
GCA_004027535.1_AcoCah_v1_BIUU10068 \ 437 Gobi jerboa euarchontoglires Allactaga bullata (Orientallactaga bullata)
GCA_004027895.1_AllBul_v1_BIUU1041416 \ 438 shrew mouse euarchontoglires Mus pahari
GCF_900095145.1_PAHARI_EIJ_v1.110093 \ 439 Transcaucasian mole vole euarchontoglires Ellobius lutescens
GCA_001685075.1_ASM168507v139086 \ 440 hispid cotton rat euarchontoglires Sigmodon hispidus
GCA_004025045.1_SigHis_v1_BIUU42415 \ 441 lesser Egyptian jerboa euarchontoglires Jaculus jaculus
GCA_000280705.1_JacJac1.051337 \ 442 Brazilian guinea pig euarchontoglires Cavia aperea
cavApe137548 \ 443 golden hamster euarchontoglires Mesocricetus auratus
GCA_000349665.1_MesAur1.010036 \ 444 Stephens's kangaroo rat euarchontoglires Dipodomys stephensi
GCA_004024685.1_DipSte_v1_BIUU323379 \ 445 American pika euarchontoglires Ochotona princeps
GCA_000292845.1_OchPri3.09978 \ 446 Ord's kangaroo rat euarchontoglires Dipodomys ordii
dipOrd210020 \ 447 little pocket mouse euarchontoglires Perognathus longimembris 38669
\ Table 1. Genome assemblies included in the 447-way Conservation track.\
\ Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A.\ \ Detection of nonneutral substitution rates on mammalian phylogenies.\ Genome Res. 2010 Jan;20(1):110-21.\ PMID: 19858363;\ PMC: PMC2798823\
\\ Kuderna LFK, Ulirsch JC, Rashid S, Ameen M, Sundaram L, Hickey G, Cox AJ, Gao H, Kumar A, Aguet F\ et al.\ \ Identification of constrained sequence elements across 239 primate genomes.\ Nature. 2023 Nov 29;.\ DOI: 10.1038/s41586-023-06798-8; PMID: 38030727\
\\ Kuderna LFK, Gao H, Janiak MC, Kuhlwilm M, Orkin JD, Bataillon T, Manu S, Valenzuela A, Bergman J,\ Rousselle M et al.\ \ A global catalog of whole-genome diversity from 233 primate species.\ Science. 2023 Jun 2;380(6648):906-913.\ DOI: 10.1126/science.abn7829;\ PMID: 37262161\
\\ Zoonomia Consortium.\ \ A comparative genomics multitool for scientific discovery and conservation.\ Nature. 2020 Nov;587(7833):240-245.\ DOI: 10.1038/s41586-020-2876-6; PMID: 33177664; PMC: PMC7759459\
\\ Feng S, Stiller J, Deng Y, Armstrong J, Fang Q, Reeve AH, Xie D, Chen G, Guo C, Faircloth BC et\ al.\ \ Dense sampling of bird diversity increases power of comparative genomics.\ Nature. 2020 Nov;587(7833):252-257.\ DOI: 10.1038/s41586-020-2873-9; PMID: 33177665; PMC: PMC7759463\
\\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q, Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ DOI: 10.1038/s41586-020-2871-y; PMID: 33177663; PMC: PMC7673649\
\ compGeno 1 altColor 0,90,10\ bigDataUrl https://hgdownload.soe.ucsc.edu/goldenPath/hg38/cactus447way/hg38.cactus447way.bb\ color 0, 10, 100\ frames https://hgdownload.soe.ucsc.edu/goldenPath/hg38/cactus447way/cactus447wayFrames.bb\ group compGeno\ irows on\ itemFirstCharCase noChange\ longLabel Cactus alignment on 447 mammal species, including Zoonomia genomes and 233 primates\ noInherit on\ parent cons447wayViewalign\ sGroup_Afrotheria Loxodonta_africana Trichechus_manatus Heterohyrax_brucei Procavia_capensis Orycteropus_afer Chrysochloris_asiatica Elephantulus_edwardii Microgale_talazaci Echinops_telfairi\ sGroup_Artiodactyla Eubalaena_japonica Eschrichtius_robustus Hippopotamus_amphibius Balaenoptera_acutorostrata Delphinapterus_leucas Balaenoptera_bonaerensis Inia_geoffrensis Phocoena_phocoena Monodon_monoceros Lipotes_vexillifer Orcinus_orca Platanista_gangetica Neophocaena_asiaeorientalis Mesoplodon_bidens Vicugna_pacos Ziphius_cavirostris Camelus_bactrianus Camelus_dromedarius Camelus_ferus Kogia_breviceps Catagonus_wagneri Rangifer_tarandus Elaphurus_davidianus Okapia_johnstoni Giraffa_tippelskirchi Moschus_moschiferus Bubalus_bubalis Bos_taurus Antilocapra_americana Odocoileus_virginianus Ammotragus_lervia Ovis_canadensis Capra_hircus Hemitragus_hylocrius Beatragus_hunteri Bos_mutus Bison_bison Ovis_aries Pantholops_hodgsonii Capra_aegagrus Tragulus_javanicus Sus_scrofa Bos_indicus Tursiops_truncatus Saiga_tatarica\ sGroup_Carnivora Panthera_onca Panthera_pardus Ailuropoda_melanoleuca Neomonachus_schauinslandi Zalophus_californianus Canis_lupus_orion Odobenus_rosmarus Felis_catus_fca126 Mirounga_angustirostris Felis_catus Canis_lupus_VD CanFam4 Canis_lupus_dingo Nyctereutes_procyonoides Cryptoprocta_ferox Ursus_maritimus Paradoxurus_hermaphroditus Lycaon_pictus Vulpes_lagopus Canis_lupus_familiaris Hyaena_hyaena Acinonyx_jubatus Panthera_tigris Enhydra_lutris Pteronura_brasiliensis Otocyon_megalotis Leptonychotes_weddellii Ailurus_fulgens Mellivora_capensis Mungos_mungo Helogale_parvula Suricata_suricatta Puma_concolor Felis_nigripes Mustela_putorius Spilogale_gracilis\ sGroup_Chiroptera Rhinolophus_sinicus Pteropus_alecto Hipposideros_galeritus Rousettus_aegyptiacus Macroglossus_sobrinus Pteropus_vampyrus Tadarida_brasiliensis Hipposideros_armiger Eidolon_helvum Mormoops_blainvillei Anoura_caudifer Desmodus_rotundus Micronycteris_hirsuta Tonatia_saurophila Carollia_perspicillata Artibeus_jamaicensis Megaderma_lyra Miniopterus_schreibersii Noctilio_leporinus Miniopterus_natalensis Craseonycteris_thonglongyai Pteronotus_parnellii Myotis_myotis Murina_feae Myotis_davidii Myotis_brandtii Eptesicus_fuscus Lasiurus_borealis Myotis_lucifugus Pipistrellus_pipistrellus\ sGroup_Euarchontoglires Galeopterus_variegatus Tupaia_chinensis Xerus_inauris Tupaia_tana Aplodontia_rufa Marmota_marmota Spermophilus_dauricus Hystrix_cristata Ictidomys_tridecemlineatus Castor_canadensis Chinchilla_lanigera Dasyprocta_punctata Dinomys_branickii Glis_glis Ctenodactylus_gundi Heterocephalus_glaber Dolichotis_patagonum Hydrochoerus_hydrochaeris Cavia_tschudii Cavia_porcellus Octodon_degus Cuniculus_paca Ctenomys_sociabilis Fukomys_damarensis Graphiurus_murinus Capromys_pilorides Nannospalax_galili Myocastor_coypus Muscardinus_avellanarius Petromus_typicus Thryonomys_swinderianus Lepus_americanus Cricetomys_gambianus Peromyscus_maniculatus Onychomys_torridus Oryctolagus_cuniculus Ondatra_zibethicus Ellobius_talpinus Meriones_unguiculatus Psammomys_obesus Mus_musculus Cricetulus_griseus Rattus_norvegicus Mus_spretus Zapus_hudsonius Microtus_ochrogaster Mus_caroli Acomys_cahirinus Allactaga_bullata Mus_pahari Ellobius_lutescens Sigmodon_hispidus Jaculus_jaculus Cavia_aperea Mesocricetus_auratus Dipodomys_stephensi Ochotona_princeps Dipodomys_ordii Perognathus_longimembris\ sGroup_Laurasiatheria Dicerorhinus_sumatrensis Diceros_bicornis Tapirus_indicus Tapirus_terrestris Ceratotherium_simum_cottoni Equus_asinus Ceratotherium_simum Equus_przewalskii Equus_caballus Manis_javanica Manis_pentadactyla Solenodon_paradoxus Scalopus_aquaticus Uropsilus_gracilis Condylura_cristata Erinaceus_europaeus Sorex_araneus Crocidura_indochinensis\ sGroup_Primates_catarrhini Pan_troglodytes Gorilla_gorilla Gorilla_beringei Pongo_abelii Pongo_pygmaeus Macaca_mulatta Theropithecus_gelada Macaca_arctoides Miopithecus_ogouensis Macaca_fascicularis Allenopithecus_nigroviridis Symphalangus_syndactylus Lophocebus_aterrimus Mandrillus_leucophaeus Macaca_radiata Cercocebus_torquatus Cercocebus_chrysogaster Cercopithecus_hamlyni Macaca_siberu Macaca_nemestrina Cercocebus_lunulatus Macaca_tonkeana Cercopithecus_diana Erythrocebus_patas Macaca_leonina Macaca_maura Papio_papio Papio_hamadryas Macaca_silenus Papio_anubis Cercopithecus_roloway Papio_kindae Papio_ursinus Allochrocebus_solatus Rhinopithecus_roxellana Chlorocebus_pygerythrus Cercocebus_atys Chlorocebus_sabaeus Cercopithecus_neglectus Papio_cynocephalus Macaca_nigra Nasalis_larvatus Allochrocebus_preussi Cercopithecus_nictitans Presbytis_comata Cercopithecus_albogularis Allochrocebus_lhoesti Cercopithecus_pogonias Presbytis_mitrata Pygathrix_cinerea Cercopithecus_mona Cercopithecus_petaurista Chlorocebus_aethiops Cercopithecus_lowei Nomascus_annamensis Nomascus_gabriellae Macaca_fuscata Piliocolobus_badius Nomascus_siki_a Nomascus_siki_b Macaca_cyclopis Pygathrix_nigripes_a Pygathrix_nigripes_b Colobus_polykomos Nomascus_concolor Piliocolobus_gordonorum Trachypithecus_geei Hylobates_klossii Trachypithecus_obscurus Piliocolobus_kirkii Trachypithecus_germaini Trachypithecus_hatinhensis Cercopithecus_cephus Trachypithecus_laotum Trachypithecus_francoisi Semnopithecus_vetulus Trachypithecus_pileatus Piliocolobus_tephrosceles Trachypithecus_auratus Cercopithecus_ascanius Trachypithecus_cristatus Semnopithecus_johnii Trachypithecus_crepusculus Trachypithecus_leucocephalus Pan_paniscus Hylobates_agilis Trachypithecus_melamera Semnopithecus_schistaceus Hylobates_abbotti Hylobates_muelleri Semnopithecus_priam Semnopithecus_hypoleucos Colobus_guereza Semnopithecus_entellus Hylobates_pileatus_a Hylobates_pileatus_b Rhinopithecus_bieti Rhinopithecus_strykeri Colobus_angolensis Macaca_thibetana Trachypithecus_phayrei Macaca_assamensis Hoolock_leuconedys Mandrillus_sphinx\ sGroup_Primates_platyrrhini Pithecia_chrysocephala Pithecia_hirsuta Pithecia_pithecia Pithecia_mittermeieri Pithecia_albicans Pithecia_pissinattii Pithecia_vanzolinii Cacajao_calvus Cacajao_ayresi Cacajao_melanocephalus Cacajao_hosomi Chiropotes_sagulatus Chiropotes_israelita Cheracebus_lugens Plecturocebus_brunneus Plecturocebus_hoffmannsi Plecturocebus_miltoni Cheracebus_torquatus Plecturocebus_cinerascens Plecturocebus_bernhardi Cheracebus_lucifer Plecturocebus_cupreus Plecturocebus_caligatus Plecturocebus_dubius Plecturocebus_moloch Plecturocebus_grovesi Ateles_geoffroyi_a Atele_geoffroyi_b Cheracebus_regulus Ateles_paniscus Ateles_chamek Ateles_marginatus Ateles_belzebuth Lagothrix_lagothricha Sapajus_macrocephalus Cebus_unicolor Cebus_olivaceus Alouatta_palliata Cebus_albifrons Aotus_trivirgatus Aotus_griseimembra Alouatta_caraya Aotus_vociferans Alouatta_belzebul Alouatta_discolor Aotus_azarae Alouatta_puruensis Alouatta_nigerrima Alouatta_macconnelli Alouatta_juara Alouatta_seniculus Sapajus_apella Aotus_nancymaae Saimiri_boliviensis Chiropotes_albinasus Leontocebus_nigricollis Leontocebus_fuscicollis Leontocebus_illigeri Saguinus_oedipus Saguinus_bicolor Saguinus_geoffroyi Saguinus_inustus Saguinus_mystax Saguinus_imperator Saimiri_sciureus Saguinus_labiatus Callimico_goeldii Saimiri_oerstedii Leontopithecus_chrysomelas Leontopithecus_rosalia Saimiri_cassiquiarensis Saimiri_ustus Saimiri_macrodon Callithrix_jacchus Cebuella_niveiventris Cebuella_pygmaea Mico_humeralifer Callibella_humilis Mico_spnv Mico_argentatus Saguinus_midas Callithrix_kuhlii Callithrix_geoffroyi\ sGroup_Primates_strepsirrhini Daubentonia_madagascariensis Propithecus_coronatus Propithecus_perrieri Varecia_variegata Propithecus_diadema Propithecus_edwardsi Indri_indri Propithecus_tattersalli Avahi_laniger Propithecus_verreauxi Avahi_peyrierasi Varecia_rubra Prolemur_simus Eulemur_rubriventer Eulemur_mongoz Cheirogaleus_major Eulemur_coronatus Eulemur_macaco Cheirogaleus_medius Eulemur_flavifrons Propithecus_coquerelli Eulemur_collaris Lepilemur_ruficaudatus Eulemur_rufus Eulemur_sanfordi Eulemur_albifrons Lepilemur_dorsalis Eulemur_fulvus Lepilemur_septentrionalis Hapalemur_occidentalis Hapalemur_alaotrensis Hapalemur_griseus Lepilemur_ankaranensis Lemur_catta Hapalemur_gilberti Hapalemur_meridionalis Galagoides_demidoff Mirza_zaza Microcebus_murinus Otolemur_garnettii Galago_senegalensis Otolemur_crassicaudatus Loris_lydekkerianus Loris_tardigradus Perodicticus_potto Perodicticus_ibeanus Galago_moholi Nycticebus_pygmaeus Nycticebus_bengalensis Arctocebus_calabarensis Nycticebus_coucang\ sGroup_Primates_tarsiidae Cephalopachus_bancanus Carlito_syrichta Tarsius_lariang Tarsius_wallacei\ sGroup_Xenarthra Choloepus_hoffmanni Dasypus_novemcinctus Myrmecophaga_tridactyla Tamandua_tetradactyla Tolypeutes_matacus Choloepus_didactylus Chaetophractus_vellerosus\ shortLabel Cactus 447-way\ speciesCodonDefault hg38\ speciesDefaultOff Pongo_abelii Gorilla_beringei Pan_troglodytes Pongo_pygmaeus Macaca_mulatta Theropithecus_gelada Macaca_arctoides Miopithecus_ogouensis Macaca_fascicularis Allenopithecus_nigroviridis Symphalangus_syndactylus Lophocebus_aterrimus Mandrillus_leucophaeus Macaca_radiata Cercocebus_torquatus Cercocebus_chrysogaster Cercopithecus_hamlyni Macaca_siberu Macaca_nemestrina Cercocebus_lunulatus Macaca_tonkeana Cercopithecus_diana Erythrocebus_patas Macaca_leonina Macaca_maura Papio_papio Papio_hamadryas Macaca_silenus Papio_anubis Cercopithecus_roloway Papio_kindae Papio_ursinus Allochrocebus_solatus Rhinopithecus_roxellana Chlorocebus_pygerythrus Cercocebus_atys Chlorocebus_sabaeus Cercopithecus_neglectus Papio_cynocephalus Macaca_nigra Nasalis_larvatus Allochrocebus_preussi Cercopithecus_nictitans Presbytis_comata Cercopithecus_albogularis Allochrocebus_lhoesti Cercopithecus_pogonias Presbytis_mitrata Pygathrix_cinerea Cercopithecus_mona Cercopithecus_petaurista Chlorocebus_aethiops Cercopithecus_lowei Nomascus_annamensis Nomascus_gabriellae Macaca_fuscata Piliocolobus_badius Nomascus_siki_a Nomascus_siki_b Macaca_cyclopis Pygathrix_nigripes_a Pygathrix_nigripes_b Colobus_polykomos Nomascus_concolor Piliocolobus_gordonorum Trachypithecus_geei Hylobates_klossii Trachypithecus_obscurus Piliocolobus_kirkii Trachypithecus_germaini Trachypithecus_hatinhensis Cercopithecus_cephus Trachypithecus_laotum Trachypithecus_francoisi Semnopithecus_vetulus Trachypithecus_pileatus Piliocolobus_tephrosceles Trachypithecus_auratus Cercopithecus_ascanius Trachypithecus_cristatus Semnopithecus_johnii Trachypithecus_crepusculus Trachypithecus_leucocephalus Pan_paniscus Hylobates_agilis Semnopithecus_schistaceus Hylobates_abbotti Hylobates_muelleri Trachypithecus_melamera Semnopithecus_priam Semnopithecus_hypoleucos Colobus_guereza Semnopithecus_entellus Hylobates_pileatus_a Hylobates_pileatus_b Rhinopithecus_bieti Rhinopithecus_strykeri Colobus_angolensis Macaca_thibetana Trachypithecus_phayrei Macaca_assamensis Pithecia_chrysocephala Pithecia_hirsuta Pithecia_pithecia Pithecia_mittermeieri Pithecia_albicans Hoolock_leuconedys Pithecia_pissinattii Pithecia_vanzolinii Cacajao_calvus Cacajao_ayresi Cacajao_melanocephalus Cacajao_hosomi Chiropotes_sagulatus Chiropotes_israelita Cheracebus_lugens Plecturocebus_brunneus Plecturocebus_hoffmannsi Plecturocebus_miltoni Cheracebus_torquatus Plecturocebus_cinerascens Plecturocebus_bernhardi Cheracebus_lucifer Plecturocebus_cupreus Plecturocebus_caligatus Plecturocebus_dubius Plecturocebus_moloch Plecturocebus_grovesi Ateles_geoffroyi_a Cheracebus_regulus Ateles_paniscus Ateles_chamek Ateles_marginatus Ateles_belzebuth Lagothrix_lagothricha Sapajus_macrocephalus Cebus_unicolor Cebus_olivaceus Alouatta_palliata Cebus_albifrons Aotus_trivirgatus Aotus_griseimembra Alouatta_caraya Aotus_vociferans Alouatta_belzebul Alouatta_discolor Aotus_azarae Alouatta_puruensis Alouatta_nigerrima Alouatta_macconnelli Alouatta_juara Alouatta_seniculus Sapajus_apella Aotus_nancymaae Saimiri_boliviensis Chiropotes_albinasus Leontocebus_nigricollis Leontocebus_fuscicollis Leontocebus_illigeri Saguinus_oedipus Saguinus_bicolor Saguinus_geoffroyi Saguinus_inustus Saguinus_mystax Saguinus_imperator Saimiri_sciureus Saguinus_labiatus Callimico_goeldii Saimiri_oerstedii Leontopithecus_chrysomelas Leontopithecus_rosalia Saimiri_cassiquiarensis Saimiri_ustus Saimiri_macrodon Callithrix_jacchus Cebuella_niveiventris Cebuella_pygmaea Mico_humeralifer Callibella_humilis Mico_spnv Mico_argentatus Saguinus_midas Callithrix_kuhlii Callithrix_geoffroyi Daubentonia_madagascariensis Cephalopachus_bancanus Carlito_syrichta Mandrillus_sphinx Galeopterus_variegatus Tarsius_lariang Propithecus_coronatus Propithecus_perrieri Varecia_variegata Propithecus_diadema Propithecus_edwardsi Indri_indri Propithecus_tattersalli Avahi_laniger Propithecus_verreauxi Tarsius_wallacei Avahi_peyrierasi Varecia_rubra Prolemur_simus Eulemur_rubriventer Eulemur_mongoz Cheirogaleus_major Eulemur_coronatus Eulemur_macaco Cheirogaleus_medius Eulemur_flavifrons Propithecus_coquerelli Eulemur_collaris Lepilemur_ruficaudatus Eulemur_rufus Eulemur_sanfordi Eulemur_albifrons Lepilemur_dorsalis Eulemur_fulvus Lepilemur_septentrionalis Hapalemur_occidentalis Hapalemur_alaotrensis Hapalemur_griseus Lepilemur_ankaranensis Lemur_catta Hapalemur_gilberti Hapalemur_meridionalis Galagoides_demidoff Mirza_zaza Microcebus_murinus Otolemur_garnettii Galago_senegalensis Otolemur_crassicaudatus Loris_lydekkerianus Loris_tardigradus Perodicticus_potto Perodicticus_ibeanus Galago_moholi Nycticebus_pygmaeus Nycticebus_bengalensis Arctocebus_calabarensis Nycticebus_coucang Dicerorhinus_sumatrensis Diceros_bicornis Tapirus_indicus Tapirus_terrestris Ceratotherium_simum_cottoni Equus_asinus Ceratotherium_simum Equus_przewalskii Equus_caballus Panthera_onca Panthera_pardus Ailuropoda_melanoleuca Neomonachus_schauinslandi Zalophus_californianus Canis_lupus_orion Odobenus_rosmarus Felis_catus_fca126 Mirounga_angustirostris Felis_catus Canis_lupus_VD CanFam4 Canis_lupus_dingo Nyctereutes_procyonoides Cryptoprocta_ferox Ursus_maritimus Paradoxurus_hermaphroditus Lycaon_pictus Vulpes_lagopus Canis_lupus_familiaris Hyaena_hyaena Acinonyx_jubatus Panthera_tigris Eubalaena_japonica Enhydra_lutris Eschrichtius_robustus Pteronura_brasiliensis Otocyon_megalotis Leptonychotes_weddellii Hippopotamus_amphibius Ailurus_fulgens Mellivora_capensis Rhinolophus_sinicus Pteropus_alecto Mungos_mungo Helogale_parvula Suricata_suricatta Puma_concolor Manis_javanica Balaenoptera_acutorostrata Felis_nigripes Mustela_putorius Hipposideros_galeritus Delphinapterus_leucas Rousettus_aegyptiacus Balaenoptera_bonaerensis Inia_geoffrensis Phocoena_phocoena Monodon_monoceros Lipotes_vexillifer Orcinus_orca Platanista_gangetica Macroglossus_sobrinus Neophocaena_asiaeorientalis Pteropus_vampyrus Mesoplodon_bidens Spilogale_gracilis Vicugna_pacos Ziphius_cavirostris Tupaia_chinensis Tadarida_brasiliensis Hipposideros_armiger Camelus_bactrianus Xerus_inauris Camelus_dromedarius Eidolon_helvum Choloepus_hoffmanni Camelus_ferus Kogia_breviceps Tupaia_tana Dasypus_novemcinctus Manis_pentadactyla Loxodonta_africana Trichechus_manatus Myrmecophaga_tridactyla Tamandua_tetradactyla Aplodontia_rufa Tolypeutes_matacus Choloepus_didactylus Catagonus_wagneri Marmota_marmota Spermophilus_dauricus Solenodon_paradoxus Mormoops_blainvillei Hystrix_cristata Anoura_caudifer Heterohyrax_brucei Procavia_capensis Desmodus_rotundus Micronycteris_hirsuta Orycteropus_afer Rangifer_tarandus Tonatia_saurophila Elaphurus_davidianus Okapia_johnstoni Giraffa_tippelskirchi Moschus_moschiferus Ictidomys_tridecemlineatus Bubalus_bubalis Bos_taurus Antilocapra_americana Odocoileus_virginianus Ammotragus_lervia Ovis_canadensis Castor_canadensis Capra_hircus Hemitragus_hylocrius Beatragus_hunteri Bos_mutus Carollia_perspicillata Artibeus_jamaicensis Chinchilla_lanigera Bison_bison Dasyprocta_punctata Dinomys_branickii Ovis_aries Megaderma_lyra Pantholops_hodgsonii Glis_glis Miniopterus_schreibersii Ctenodactylus_gundi Noctilio_leporinus Miniopterus_natalensis Heterocephalus_glaber Dolichotis_patagonum Capra_aegagrus Tragulus_javanicus Hydrochoerus_hydrochaeris Cavia_tschudii Cavia_porcellus Sus_scrofa Octodon_degus Craseonycteris_thonglongyai Cuniculus_paca Ctenomys_sociabilis Chaetophractus_vellerosus Fukomys_damarensis Graphiurus_murinus Capromys_pilorides Nannospalax_galili Bos_indicus Tursiops_truncatus Myocastor_coypus Muscardinus_avellanarius Saiga_tatarica Pteronotus_parnellii Petromus_typicus Myotis_myotis Thryonomys_swinderianus Murina_feae Lepus_americanus Myotis_davidii Myotis_brandtii Cricetomys_gambianus Eptesicus_fuscus Peromyscus_maniculatus Onychomys_torridus Oryctolagus_cuniculus Scalopus_aquaticus Ondatra_zibethicus Lasiurus_borealis Ellobius_talpinus Meriones_unguiculatus Psammomys_obesus Mus_musculus Cricetulus_griseus Rattus_norvegicus Mus_spretus Zapus_hudsonius Chrysochloris_asiatica Microtus_ochrogaster Mus_caroli Acomys_cahirinus Allactaga_bullata Mus_pahari Ellobius_lutescens Sigmodon_hispidus Uropsilus_gracilis Jaculus_jaculus Myotis_lucifugus Cavia_aperea Pipistrellus_pipistrellus Mesocricetus_auratus Elephantulus_edwardii Dipodomys_stephensi Ochotona_princeps Dipodomys_ordii Perognathus_longimembris Condylura_cristata Microgale_talazaci Echinops_telfairi Erinaceus_europaeus Sorex_araneus\ speciesDefaultOn Gorilla_gorilla Pongo_abelii Pithecia_chrysocephala Pithecia_hirsuta Cephalopachus_bancanus Carlito_syrichta Daubentonia_madagascariensis Propithecus_coronatus Panthera_onca Panthera_pardus Dicerorhinus_sumatrensis Diceros_bicornis Choloepus_hoffmanni Dasypus_novemcinctus Eubalaena_japonica Eschrichtius_robustus Rhinolophus_sinicus Pteropus_alecto Loxodonta_africana Trichechus_manatus Galeopterus_variegatus Tupaia_chinensis Dipodomys_ordii Perognathus_longimembris\ speciesGroups Primates_catarrhini Primates_platyrrhini Primates_tarsiidae Primates_strepsirrhini Carnivora Laurasiatheria Xenarthra Artiodactyla Chiroptera Afrotheria Euarchontoglires\ speciesLabels Acinonyx_jubatus="cheetah" Acomys_cahirinus="Egyptian spiny mouse" Ailuropoda_melanoleuca="giant panda" Ailurus_fulgens="Lesser panda" Allactaga_bullata="Gobi jerboa" Allenopithecus_nigroviridis="Allen's swamp monkey" Allochrocebus_lhoesti="L'Hoest's monkey" Allochrocebus_preussi="Preuss's monkey" Allochrocebus_solatus="Sun-tailed monkey" Alouatta_palliata="mantled howler" Alouatta_belzebul="Eastern Red-handed howler" Alouatta_caraya="black-and-gold howler" Alouatta_discolor="Spix's Red-handed howler" Alouatta_juara="Jurua red howler monkey" Alouatta_macconnelli="Guianan red howler" Alouatta_nigerrima="Amazon black howler" Alouatta_puruensis="Purús red howler monkey" Alouatta_seniculus="Colombian red howler" Ammotragus_lervia="aoudad" Anoura_caudifer="tailed tailless bat" Antilocapra_americana="pronghorn" Aotus_nancymaae="Ma's night monkey" Aotus_azarae="Azara's night monkey" Aotus_griseimembra="Gray-legged Night monkey" Aotus_trivirgatus="Humboldt's night monkey" Aotus_vociferans="Spix's night monkey" Aplodontia_rufa="mountain beaver" Arctocebus_calabarensis="Calabar Angwantibo" Artibeus_jamaicensis="Jamaican fruit-eating bat" Ateles_geoffroyi="Central American spider monkey" Ateles_belzebuth="white-bellied spider monkey" Ateles_chamek="black spider monkey" Ateles_marginatus="white-whiskered spider monkey" Ateles_paniscus="Red-faced black spider monkey" Avahi_laniger="Eastern Woolly lemur" Avahi_peyrierasi="Peyrieras's Woolly lemur" Balaenoptera_bonaerensis="Antarctic minke whale" Balaenoptera_acutorostrata="Minke whale" Beatragus_hunteri="hirola" Bison_bison="American bison" Bos_indicus="zebu cattle" Bos_mutus="wild yak" Bos_taurus="cow" Bubalus_bubalis="water buffalo" Cacajao_ayresi="Araca Uakari" Cacajao_calvus="Bald Uakari" Cacajao_hosomi="Neblina black Uakari" Cacajao_melanocephalus="Golden-brown Uakari" Callibella_humilis="black-crowned Dwarf Marmoset" Callimico_goeldii="Goeldi's monkey" Callithrix_jacchus="common marmoset" Callithrix_geoffroyi="Geoffroy's tufted-ear marmoset" Callithrix_kuhlii="Wied's Marmoset" Camelus_bactrianus="Bactrian camel" Camelus_dromedarius="Arabian camel" Camelus_ferus="wild Bactrian camel" CanFam4="German Shepherd dog (Mischka)" Canis_lupus_dingo="dingo" Canis_lupus_familiaris="dog" Canis_lupus_VD="domestic dog (BS72/Village Dog)" Canis_lupus_orion="Greenland wolf" Capra_aegagrus="wild goat" Capra_hircus="goat" Capromys_pilorides="Desmarest's hutia" Carlito_syrichta="Philippine tarsier" Carollia_perspicillata="Seba's short-tailed bat" Castor_canadensis="American beaver" Catagonus_wagneri="Chacoan peccary" Cavia_aperea="Brazilian guinea pig" Cavia_porcellus="domestic guinea pig" Cavia_tschudii="Montane guinea pig" Cebuella_niveiventris="Southern Pygmy Marmoset" Cebuella_pygmaea="Northern Pygmy Marmoset" Cebus_albifrons="white-fronted capuchin" Cebus_olivaceus="Guinan Weeper capuchin" Cebus_unicolor="Spix's white-fronted capuchin" Cephalopachus_bancanus="Western tarsier" Ceratotherium_simum_cottoni="northern white rhinoceros" Ceratotherium_simum="Southern white rhinoceros" Cercocebus_atys="sooty mangabey" Cercocebus_chrysogaster="Golden-bellied Mangabey" Cercocebus_lunulatus="white-naped Mangabey" Cercocebus_torquatus="Red-capped Mangabey" Cercopithecus_mona="Mona monkey" Cercopithecus_neglectus="De Brazza's monkey" Cercopithecus_ascanius="Red-tailed monkey" Cercopithecus_cephus="Mustached monkey" Cercopithecus_diana="Diana monkey" Cercopithecus_hamlyni="Owl-faced monkey" Cercopithecus_lowei="Lowe's monkey" Cercopithecus_albogularis="Sykes' monkey" Cercopithecus_nictitans="Putty-nosed monkey" Cercopithecus_petaurista="Spot-nosed monkey" Cercopithecus_pogonias="Crowned monkey" Cercopithecus_roloway="Roloway monkey" Chaetophractus_vellerosus="screaming hairy armadillo" Cheirogaleus_medius="fat-tailed dwarf lemur" Cheirogaleus_major="Greater Dwarf lemur" Cheracebus_lucifer="Yellow-handed Titi" Cheracebus_lugens="white-chested Titi" Cheracebus_regulus="Rio Jurua Collared Titi" Cheracebus_torquatus="white-collared Titi" Chinchilla_lanigera="long-tailed chinchilla" Chiropotes_albinasus="Red-nosed Bearded saki" Chiropotes_israelita="Spix's Bearded saki" Chiropotes_sagulatus="Guianan Bearded saki" Chlorocebus_aethiops="grivet monkey" Chlorocebus_sabaeus="green monkey" Chlorocebus_pygerythrus="Vervet monkey" Choloepus_didactylus="southern two-toed sloth" Choloepus_hoffmanni="Hoffmann's two-fingered sloth" Chrysochloris_asiatica="Cape golden mole" Colobus_guereza="guereza" Colobus_angolensis="Angolan colobus" Colobus_polykomos="King Colobus" Condylura_cristata="star-nosed mole" Craseonycteris_thonglongyai="hog-nosed bat" Cricetomys_gambianus="Gambian giant pouched rat" Cricetulus_griseus="Chinese hamster" Crocidura_indochinensis="Indochinese shrew" Cryptoprocta_ferox="fossa" Ctenodactylus_gundi="northern gundi" Ctenomys_sociabilis="social tuco-tuco" Cuniculus_paca="lowland paca" Dasyprocta_punctata="punctate agouti" Dasypus_novemcinctus="nine-banded armadillo" Daubentonia_madagascariensis="aye-aye" Delphinapterus_leucas="beluga whale" Desmodus_rotundus="common vampire bat" Dicerorhinus_sumatrensis="Sumatran rhinoceros" Diceros_bicornis="black rhinoceros" Dinomys_branickii="pacarana" Dipodomys_ordii="Ord's kangaroo rat" Dipodomys_stephensi="Stephens's kangaroo rat" Dolichotis_patagonum="Patagonian cavy" Echinops_telfairi="small Madagascar hedgehog" Eidolon_helvum="straw-colored fruit bat" Elaphurus_davidianus="Pere David's deer" Elephantulus_edwardii="Cape elephant shrew" Ellobius_lutescens="Transcaucasian mole vole" Ellobius_talpinus="northern mole vole" Enhydra_lutris="Sea otter" Eptesicus_fuscus="big brown bat" Equus_asinus="ass" Equus_caballus="horse" Equus_przewalskii="Przewalski's horse" Erinaceus_europaeus="western European hedgehog" Erythrocebus_patas="common Patas monkey" Eschrichtius_robustus="grey whale" Eubalaena_japonica="North Pacific right whale" Eulemur_flavifrons="blue-eyed black lemur" Eulemur_fulvus="brown lemur" Eulemur_macaco="black lemur" Eulemur_mongoz="mongoose lemur" Eulemur_albifrons="white-fronted brown lemur" Eulemur_collaris="Red-collared brown lemur" Eulemur_coronatus="Crowned lemur" Eulemur_rubriventer="Red-bellied lemur" Eulemur_rufus="Rufous brown lemur" Eulemur_sanfordi="Sanford's brown lemur" Felis_catus="domestic cat" Felis_nigripes="black-footed cat" Felis_catus_fca126="domestic cat (Fca126)" Fukomys_damarensis="Damara mole-rat" Galago_moholi="southern leser galago" Galago_senegalensis="Northern Lesser Galago" Galagoides_demidoff="Demidoff's Dwarf Galago" Galeopterus_variegatus="Sunda flying lemur" Giraffa_tippelskirchi="Masai giraffe" Glis_glis="fat dormouse" Gorilla_gorilla="western gorilla" Gorilla_beringei="Eastern Gorilla" Graphiurus_murinus="woodland dormouse" Hapalemur_alaotrensis="Lac Alaotra Bamboo lemur" Hapalemur_gilberti="Gilbert's Gray Bamboo lemur" Hapalemur_griseus="Common Gray Bamboo lemur" Hapalemur_meridionalis="Southern Bamboo lemur" Hapalemur_occidentalis="Northern Bamboo lemur" Helogale_parvula="dwarf mongoose" Hemitragus_hylocrius="Nilgiri tahr" Heterocephalus_glaber="naked mole-rat" Heterohyrax_brucei="yellow-spotted hyrax" Hippopotamus_amphibius="hippopotamus" Hipposideros_armiger="great roundleaf bat" Hipposideros_galeritus="Cantor's roundleaf bat" Hoolock_leuconedys="Eastern hoolock Gibbon" Hyaena_hyaena="striped hyena" Hydrochoerus_hydrochaeris="capybara" Hylobates_pileatus_a="pileated gibbon" Hylobates_pileatus_b="pileated gibbon" Hylobates_abbotti="Western gray gibbon" Hylobates_agilis="agile gibbon" Hylobates_klossii="Kloss's gibbon" Hylobates_muelleri="Southern gray gibbon" Hystrix_cristata="crested porcupine" Ictidomys_tridecemlineatus="thirteen-lined ground squirrel" Indri_indri="indri" Inia_geoffrensis="boutu" Jaculus_jaculus="lesser Egyptian jerboa" Kogia_breviceps="pygmy sperm whale" Lagothrix_lagothricha="Common Woolly monkey" Lasiurus_borealis="red bat" Lemur_catta="ring-tailed lemur" Leontocebus_fuscicollis="Spix's Saddle-back tamarin" Leontocebus_illigeri="Illiger's Saddle-back tamarin" Leontocebus_nigricollis="black-mantled tamarin" Leontopithecus_rosalia="golden lion tamarin" Leontopithecus_chrysomelas="Golden-headed Lion tamarin" Lepilemur_ankaranensis="Ankarana sportive lemur" Lepilemur_dorsalis="Gray's sportive lemur" Lepilemur_ruficaudatus="Red-tailed sportive lemur" Lepilemur_septentrionalis="Sahafary sportive lemur" Leptonychotes_weddellii="Weddell seal" Lepus_americanus="snowshoe hare" Lipotes_vexillifer="Yangtze River dolphin" Lophocebus_aterrimus="black crested mangabey" Loris_tardigradus="red slender loris" Loris_lydekkerianus="Gray Slender Loris" Loxodonta_africana="African savanna elephant" Lycaon_pictus="African hunting dog" Macaca_arctoides="stump-tailed macaque" Macaca_assamensis="Assamese macaque" Macaca_cyclopis="Taiwanexe macaque" Macaca_fascicularis="long-tailed macaque" Macaca_mulatta="Rhesus macaque" Macaca_nemestrina="southern pig-tailed macaque" Macaca_nigra="crested macaque" Macaca_silenus="lion-tailed macaque" Macaca_fuscata="Japanese macaque" Macaca_leonina="Northern Pig-tailed Macaque" Macaca_maura="Moor Macaque" Macaca_radiata="Bonnet Macaque" Macaca_siberu="Siberut Macaque" Macaca_thibetana="Tibetan Macaque" Macaca_tonkeana="Tonkean Macaque" Macroglossus_sobrinus="long-tongued fruit bat" Mandrillus_leucophaeus="drill" Mandrillus_sphinx="mandrill" Manis_javanica="Malayan pangolin" Manis_pentadactyla="Chinese pangolin" Marmota_marmota="Alpine marmot" Megaderma_lyra="Indian false vampire" Mellivora_capensis="ratel" Meriones_unguiculatus="Mongolian gerbil" Mesocricetus_auratus="golden hamster" Mesoplodon_bidens="Sowerby's beaked whale" Mico_argentatus="silvery marmoset" Mico_humeralifer="Santarem marmoset" Mico_spnv="Schneider's marmoset" Microcebus_murinus="gray mouse lemur" Microgale_talazaci="Talazac's shrew tenrec" Micronycteris_hirsuta="hairy big-eared bat" Microtus_ochrogaster="prairie vole" Miniopterus_natalensis="Natal long-fingered bat" Miniopterus_schreibersii="Schreibers' long-fingered bat" Miopithecus_ogouensis="Northern Talapoin monkey" Mirounga_angustirostris="northern elephant seal" Mirza_zaza="northern giant mouse lemur" Monodon_monoceros="narwhal" Mormoops_blainvillei="Antillean ghost-faced bat" Moschus_moschiferus="Siberian musk deer" Mungos_mungo="banded mongoose" Murina_feae="Ashy-gray tube-nosed bat" Mus_caroli="Ryukyu mouse" Mus_musculus="house mouse" Mus_pahari="shrew mouse" Mus_spretus="western wild mouse" Muscardinus_avellanarius="hazel dormouse" Mustela_putorius="European polecat" Myocastor_coypus="nutria" Myotis_brandtii="Brandt's bat" Myotis_davidii="David's myotis" Myotis_lucifugus="little brown bat" Myotis_myotis="greater mouse-eared bat" Myrmecophaga_tridactyla="giant anteater" Nannospalax_galili="Upper Galilee mountains blind mole rat" Nasalis_larvatus="proboscis monkey" Neomonachus_schauinslandi="Hawaiian monk seal" Neophocaena_asiaeorientalis="Yangtze finless porpoise" Noctilio_leporinus="greater bulldog bat" Nomascus_siki_a="southern white-cheeked crested gibbon" Nomascus_siki_b="southern white-cheeked crested gibbon" Nomascus_annamensis="Northern yellow-cheeked crested gibbon" Nomascus_concolor="Western black crested gibbon" Nomascus_gabriellae="Southern yellow-cheeked crested gibbon" Nyctereutes_procyonoides="raccoon dog" Nycticebus_bengalensis="Bengal slow loris" Nycticebus_coucang="Malaysian slow loris" Nycticebus_pygmaeus="Pygmy Slow Loris" Ochotona_princeps="American pika" Octodon_degus="degu" Odobenus_rosmarus="Pacific walrus" Odocoileus_virginianus="white-tailed deer" Okapia_johnstoni="okapi" Ondatra_zibethicus="muskrat" Onychomys_torridus="southern grasshopper mouse" Orcinus_orca="killer whale" Orycteropus_afer="aardvark" Oryctolagus_cuniculus="rabbit" Otocyon_megalotis="bat-eared fox" Otolemur_garnettii="Garnetts greater galago" Otolemur_crassicaudatus="Thick-tailed Greater Galago" Ovis_aries="sheep" Ovis_canadensis="bighorn sheep" Pan_paniscus="bonobo" Pan_troglodytes="chimpanzee" Panthera_onca="jaguar" Panthera_pardus="leopard" Panthera_tigris="tiger" Pantholops_hodgsonii="chiru" Papio_anubis="olive baboon" Papio_hamadryas="hamadryas baboon" Papio_cynocephalus="Yellow Baboon" Papio_kindae="Kinda Baboon" Papio_papio="Guinea Baboon" Papio_ursinus="Chacma Baboon" Paradoxurus_hermaphroditus="Asian palm civet" Perodicticus_ibeanus="East African Potto" Perodicticus_potto="West African Potto" Perognathus_longimembris="little pocket mouse" Peromyscus_maniculatus="Prairie deer mouse" Petromus_typicus="dassie-rat" Phocoena_phocoena="harbor porpoise" Piliocolobus_tephrosceles="Ashy red Colobus" Piliocolobus_badius="Upper Guinea red Colobus" Piliocolobus_gordonorum="Udzungwa red Colobus" Piliocolobus_kirkii="Zanzibar red Colobus" Pipistrellus_pipistrellus="common pipistrelle" Pithecia_pithecia="white-faced saki" Pithecia_albicans="Buffy saki" Pithecia_chrysocephala="Golden-faced saki" Pithecia_hirsuta="Hairy saki" Pithecia_mittermeieri="Mittermeier's saki" Pithecia_pissinattii="Pissinatti's saki" Pithecia_vanzolinii="Vanzolini's bald-faced saki" Platanista_gangetica="Ganges River dolphin" Plecturocebus_bernhardi="Prince Bernhard's Titi" Plecturocebus_brunneus="brown Titi" Plecturocebus_caligatus="Chestnut-bellied Titi" Plecturocebus_cinerascens="Ashy Titi" Plecturocebus_cupreus="Coppery Titi" Plecturocebus_dubius="Hershkovitzs Titi" Plecturocebus_grovesi="Groves's Titi" Plecturocebus_hoffmannsi="Hoffmanns's Titi" Plecturocebus_miltoni="Milton's Titi" Plecturocebus_moloch="Red-bellied Titi" Pongo_abelii="Sumatran orangutan" Pongo_pygmaeus="Bornean orangutan" Presbytis_comata="Javan langur" Presbytis_mitrata="Mitered langur" Procavia_capensis="Cape rock hyrax" Prolemur_simus="greater bamboo lemur" Propithecus_coquerelli="Coquerel's Sifaka" Propithecus_coronatus="Crowned Sifaka" Propithecus_diadema="Diademed Sifaka" Propithecus_edwardsi="Milne-Edward's Sifaka" Propithecus_perrieri="Perrier's Sifaka" Propithecus_tattersalli="Tattersall's Sifaka" Propithecus_verreauxi="Verreaux's Sifaka" Psammomys_obesus="fat sand rat" Pteronotus_parnellii="Parnell's mustached bat" Pteronura_brasiliensis="giant otter" Pteropus_alecto="black flying fox" Pteropus_vampyrus="large flying fox" Puma_concolor="puma" Pygathrix_nigripes_a="black-shanked douc" Pygathrix_nigripes_b="black-shanked douc" Pygathrix_cinerea="gray-shanked douc" Rangifer_tarandus="reindeer" Rattus_norvegicus="Norway rat" Rhinolophus_sinicus="Chinese rufous horseshoe bat" Rhinopithecus_bieti="Yunnan snub-nosed monkey" Rhinopithecus_roxellana="golden snub-nosed monkey" Rhinopithecus_strykeri="Stryker's snub-nosed monkey" Rousettus_aegyptiacus="Egyptian rousette" Saguinus_imperator="Emperor tamarin" Saguinus_midas="Midas tamarin" Saguinus_bicolor="Pied Bare-faced tamarin" Saguinus_geoffroyi="Geoffroy's tamarin" Saguinus_inustus="Mottled-face tamarin" Saguinus_labiatus="Red-bellied tamarin" Saguinus_mystax="Mustached tamarin" Saguinus_oedipus="Cotton-top tamarin" Saiga_tatarica="Saiga antelope" Saimiri_boliviensis="black-capped squirrel monkey" Saimiri_cassiquiarensis="Humboldt's squirrel monkey" Saimiri_macrodon="Ecuadorian squirrel monkey" Saimiri_oerstedii="Central American squirrel monkey" Saimiri_sciureus="Guianan squirrel monkey" Saimiri_ustus="Golden-backed squirrel monkey" Sapajus_apella="brown capuchin" Sapajus_macrocephalus="large-headed capuchin" Scalopus_aquaticus="eastern mole" Semnopithecus_entellus="Bengal sacred langur" Semnopithecus_hypoleucos="Malabar Sacred langur" Semnopithecus_johnii="Nilgiri langur" Semnopithecus_priam="Tufted Gray langur" Semnopithecus_schistaceus="Nepal Sacred langur" Semnopithecus_vetulus="Purple-faced langur" Sigmodon_hispidus="hispid cotton rat" Solenodon_paradoxus="Hispaniolan solenodon" Sorex_araneus="European shrew" Spermophilus_dauricus="Daurian ground squirrel" Spilogale_gracilis="western spotted skunk" Suricata_suricatta="meerkat" Sus_scrofa="pig" Symphalangus_syndactylus="siamang" Tadarida_brasiliensis="Brazilian free-tailed bat" Tamandua_tetradactyla="southern tamandua" Tapirus_indicus="Asiatic tapir" Tapirus_terrestris="Brazilian tapir" Tarsius_lariang="Lariang tarsier" Tarsius_wallacei="Wallace's tarsier" Theropithecus_gelada="gelada" Thryonomys_swinderianus="greater cane rat" Tolypeutes_matacus="placentals" Tonatia_saurophila="stripe-headed round-eared bat" Trachypithecus_francoisi="Francois's langur" Trachypithecus_auratus="East Javan Langur" Trachypithecus_crepusculus="Indochinese Gray Langur" Trachypithecus_cristatus="Sunda Silvery Langur" Trachypithecus_geei="Golden langur" Trachypithecus_germaini="Germain's langur" Trachypithecus_hatinhensis="Hatinh langur" Trachypithecus_laotum="Laos langur" Trachypithecus_leucocephalus="white-headed langur" Trachypithecus_melamera="Shan langur" Trachypithecus_obscurus="Dusky langur" Trachypithecus_phayrei="Phayre's langur" Trachypithecus_pileatus="capped langur" Tragulus_javanicus="Java mouse-deer" Trichechus_manatus="Florida manatee" Tupaia_chinensis="Chinese tree shrew" Tupaia_tana="large tree shrew" Tursiops_truncatus="common bottlenose dolphin" Uropsilus_gracilis="gracile shrew mole" Ursus_maritimus="polar bear" Varecia_variegata="black-and-white ruffed lemur" Varecia_rubra="red ruffed lemur" Vicugna_pacos="alpaca" Vulpes_lagopus="Arctic fox" Xerus_inauris="South African ground squirrel" Zalophus_californianus="California sea lion" Zapus_hudsonius="meadow jumping mouse" Ziphius_cavirostris="Cuvier's beaked whale"\ subGroups view=align\ summary https://hgdownload.soe.ucsc.edu/goldenPath/hg38/cactus447way/cactus447waySummary.bb\ track cactus447way\ treeImage phylo/hg38_447way.png\ type bigMaf\ viewUi on\ caddSuper CADD 1.6 bed CADD 1.6 Score for all single-basepair mutations and selected insertions/deletions 0 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/ins.bb stdout
\ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts, \ documented in our \ makeDoc files.\
\ \\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ phenDis 1 color 100,130,160\ group phenDis\ html caddSuper.html\ longLabel CADD 1.6 Score for all single-basepair mutations and selected insertions/deletions\ shortLabel CADD 1.6\ superTrack on hide\ track caddSuper\ type bed\ visibility hide\ cadd CADD 1.6 bigWig CADD 1.6 Score for all possible single-basepair mutations (zoom in for scores) 1 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/ins.bb stdout
\ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts, \ documented in our \ makeDoc files.\
\ \\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ phenDis 0 color 100,130,160\ compositeTrack on\ group phenDis\ html caddSuper\ longLabel CADD 1.6 Score for all possible single-basepair mutations (zoom in for scores)\ maxWindowToDraw 10000000\ mouseOverFunction noAverage\ parent caddSuper\ shortLabel CADD 1.6\ track cadd\ type bigWig\ visibility dense\ caddDel CADD 1.6 Del bigBed 9 + CADD 1.6 Score: Deletions - label is length of deletion 1 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/ins.bb stdout
\ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts, \ documented in our \ makeDoc files.\
\ \\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ phenDis 1 bigDataUrl /gbdb/hg38/cadd/del.bb\ filter.score 10:100\ filterByRange.score on\ filterLabel.score Show only items with PHRED scale score of\ filterLimits.score 0:100\ html caddSuper\ longLabel CADD 1.6 Score: Deletions - label is length of deletion\ mouseOver Mutation: $change CADD Phred score: $phred\ parent caddSuper\ shortLabel CADD 1.6 Del\ track caddDel\ type bigBed 9 +\ visibility dense\ caddIns CADD 1.6 Ins bigBed 9 + CADD 1.6 Score: Insertions - label is length of insertion 1 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd/ins.bb stdout
\ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts, \ documented in our \ makeDoc files.\
\ \\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ phenDis 1 bigDataUrl /gbdb/hg38/cadd/ins.bb\ filter.score 10:100\ filterByRange.score on\ filterLabel.score Show only items with PHRED scale score of\ filterLimits.score 0:100\ html caddSuper\ longLabel CADD 1.6 Score: Insertions - label is length of insertion\ mouseOver Mutation: $change CADD Phred score: $phred\ parent caddSuper\ shortLabel CADD 1.6 Ins\ track caddIns\ type bigBed 9 +\ visibility dense\ caddSuper1_7 CADD 1.7 bed CADD 1.7 Score for all single-basepair mutations and selected insertions/deletions 0 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \Track colors
\\ This track is colored according to Table 2 in Vikas et al. The colors represent the recommended ACMG/AMP score cutoffs. \ \
| Range | \Classification | \
|---|---|
| ≥ 25.3 | \Pathogenic | \
| 25.2 - 22.6 | \Neutral | \
| ≤ 22.7 | \Benign | \
\ In CADD version 1.7, new features have been added to improve CADD scores for certain variant\ effects, boosting the overall performance of CADD and bringing new developments to the community.\ CADD v1.7 integrates annotations from recent efforts to assess variant effects, along with new\ conservation and mutation scores.
\\ CADD v1.7 supports only the major chromosomes of the hg38/GRCh38 reference genome (chromosomes 1-22,\ X, and Y) and may be the last version to support the hg19/GRCh37 human reference genome.
\\ This version includes scores derived from Evolutionary Scale Modeling (ESM) for assessing variants\ in protein-coding regions, along with scores from a convolutional neural network (CNN) trained on\ open chromatin sequences, used as a proxy for regulatory regions in the genome. The previously\ included conservation scores have been updated with data from the Zoonomia project. New annotations\ have also been added for 3' Untranslated Regions (3' UTRs), along with models of genome-wide\ mutational rates. The gene and transcript models have been updated by advancing from Ensembl version\ 95 to version 110, and the Ensembl Variant Effect Predictor (VEP) has been upgraded accordingly.
\\ The models in CADD v1.7 have been trained similarly to the version 1.6 release. The logistic\ regression uses an L2 penalty with C = 1, and training was completed after thirteen L-BFGS\ iterations using the sklearn library The new models exhibit a high degree of similarity to the\ previous release, with a Spearman correlation of 0.946 for CADD scores calculated for 100,000\ randomly selected variants between CADD GRCh38-v1.6 and CADD GRCh38-v1.7. The v1.7 models perform\ comparably to earlier versions in distinguishing known pathogenic variants (ClinVar) from common\ variants (gnomAD) across the genome. Improvements in CADD v1.7 are particularly evident when\ focusing on specific variant categories, such as missense or 3' UTR variants, where the latest\ release includes updated annotations.
\\ More information can be found at the\ CADD site\ and the Schubach et al., Nucleic Acids Res, 2024 publication.\ \ \ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts,\ documented in our \ makeDoc files.\
\ \ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/ins.bb stdout
\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ \\ Schubach M, Maass T, Nazaretyan L, Röner S, Kircher M.\ \ CADD v1.7: using protein language models, regulatory CNNs and other nucleotide-level scores to\ improve genome-wide variant predictions.\ Nucleic Acids Res. 2024 Jan 5;52(D1):D1143-D1154.\ PMID: 38183205; PMC: PMC10767851\
\ phenDis 1 color 100,130,160\ group phenDis\ html caddSuper1_7\ longLabel CADD 1.7 Score for all single-basepair mutations and selected insertions/deletions\ shortLabel CADD 1.7\ superTrack on hide\ track caddSuper1_7\ type bed\ visibility hide\ cadd1_7 CADD 1.7 bigWig CADD 1.7 Score for all possible single-basepair mutations (zoom in for scores) 1 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \Track colors
\\ This track is colored according to Table 2 in Vikas et al. The colors represent the recommended ACMG/AMP score cutoffs. \ \
| Range | \Classification | \
|---|---|
| ≥ 25.3 | \Pathogenic | \
| 25.2 - 22.6 | \Neutral | \
| ≤ 22.7 | \Benign | \
\ In CADD version 1.7, new features have been added to improve CADD scores for certain variant\ effects, boosting the overall performance of CADD and bringing new developments to the community.\ CADD v1.7 integrates annotations from recent efforts to assess variant effects, along with new\ conservation and mutation scores.
\\ CADD v1.7 supports only the major chromosomes of the hg38/GRCh38 reference genome (chromosomes 1-22,\ X, and Y) and may be the last version to support the hg19/GRCh37 human reference genome.
\\ This version includes scores derived from Evolutionary Scale Modeling (ESM) for assessing variants\ in protein-coding regions, along with scores from a convolutional neural network (CNN) trained on\ open chromatin sequences, used as a proxy for regulatory regions in the genome. The previously\ included conservation scores have been updated with data from the Zoonomia project. New annotations\ have also been added for 3' Untranslated Regions (3' UTRs), along with models of genome-wide\ mutational rates. The gene and transcript models have been updated by advancing from Ensembl version\ 95 to version 110, and the Ensembl Variant Effect Predictor (VEP) has been upgraded accordingly.
\\ The models in CADD v1.7 have been trained similarly to the version 1.6 release. The logistic\ regression uses an L2 penalty with C = 1, and training was completed after thirteen L-BFGS\ iterations using the sklearn library The new models exhibit a high degree of similarity to the\ previous release, with a Spearman correlation of 0.946 for CADD scores calculated for 100,000\ randomly selected variants between CADD GRCh38-v1.6 and CADD GRCh38-v1.7. The v1.7 models perform\ comparably to earlier versions in distinguishing known pathogenic variants (ClinVar) from common\ variants (gnomAD) across the genome. Improvements in CADD v1.7 are particularly evident when\ focusing on specific variant categories, such as missense or 3' UTR variants, where the latest\ release includes updated annotations.
\\ More information can be found at the\ CADD site\ and the Schubach et al., Nucleic Acids Res, 2024 publication.\ \ \ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts,\ documented in our \ makeDoc files.\
\ \ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/ins.bb stdout
\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ \\ Schubach M, Maass T, Nazaretyan L, Röner S, Kircher M.\ \ CADD v1.7: using protein language models, regulatory CNNs and other nucleotide-level scores to\ improve genome-wide variant predictions.\ Nucleic Acids Res. 2024 Jan 5;52(D1):D1143-D1154.\ PMID: 38183205; PMC: PMC10767851\
\ phenDis 0 color 100,130,160\ compositeTrack on\ group phenDis\ html caddSuper1_7\ longLabel CADD 1.7 Score for all possible single-basepair mutations (zoom in for scores)\ maxWindowToDraw 10000000\ mouseOverFunction noAverage\ parent caddSuper1_7\ shortLabel CADD 1.7\ track cadd1_7\ type bigWig\ visibility dense\ cadd1_7_Del CADD 1.7 Del bigBed 9 + CADD 1.7 Score: Deletions - label is length of deletion 1 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \Track colors
\\ This track is colored according to Table 2 in Vikas et al. The colors represent the recommended ACMG/AMP score cutoffs. \ \
| Range | \Classification | \
|---|---|
| ≥ 25.3 | \Pathogenic | \
| 25.2 - 22.6 | \Neutral | \
| ≤ 22.7 | \Benign | \
\ In CADD version 1.7, new features have been added to improve CADD scores for certain variant\ effects, boosting the overall performance of CADD and bringing new developments to the community.\ CADD v1.7 integrates annotations from recent efforts to assess variant effects, along with new\ conservation and mutation scores.
\\ CADD v1.7 supports only the major chromosomes of the hg38/GRCh38 reference genome (chromosomes 1-22,\ X, and Y) and may be the last version to support the hg19/GRCh37 human reference genome.
\\ This version includes scores derived from Evolutionary Scale Modeling (ESM) for assessing variants\ in protein-coding regions, along with scores from a convolutional neural network (CNN) trained on\ open chromatin sequences, used as a proxy for regulatory regions in the genome. The previously\ included conservation scores have been updated with data from the Zoonomia project. New annotations\ have also been added for 3' Untranslated Regions (3' UTRs), along with models of genome-wide\ mutational rates. The gene and transcript models have been updated by advancing from Ensembl version\ 95 to version 110, and the Ensembl Variant Effect Predictor (VEP) has been upgraded accordingly.
\\ The models in CADD v1.7 have been trained similarly to the version 1.6 release. The logistic\ regression uses an L2 penalty with C = 1, and training was completed after thirteen L-BFGS\ iterations using the sklearn library The new models exhibit a high degree of similarity to the\ previous release, with a Spearman correlation of 0.946 for CADD scores calculated for 100,000\ randomly selected variants between CADD GRCh38-v1.6 and CADD GRCh38-v1.7. The v1.7 models perform\ comparably to earlier versions in distinguishing known pathogenic variants (ClinVar) from common\ variants (gnomAD) across the genome. Improvements in CADD v1.7 are particularly evident when\ focusing on specific variant categories, such as missense or 3' UTR variants, where the latest\ release includes updated annotations.
\\ More information can be found at the\ CADD site\ and the Schubach et al., Nucleic Acids Res, 2024 publication.\ \ \ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts,\ documented in our \ makeDoc files.\
\ \ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/ins.bb stdout
\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ \\ Schubach M, Maass T, Nazaretyan L, Röner S, Kircher M.\ \ CADD v1.7: using protein language models, regulatory CNNs and other nucleotide-level scores to\ improve genome-wide variant predictions.\ Nucleic Acids Res. 2024 Jan 5;52(D1):D1143-D1154.\ PMID: 38183205; PMC: PMC10767851\
\ phenDis 1 bigDataUrl /gbdb/hg38/cadd1.7/del.bb\ filter.score 10:100\ filterByRange.score on\ filterLabel.score Show only items with PHRED scale score of\ filterLimits.score 0:100\ html caddSuper1_7\ longLabel CADD 1.7 Score: Deletions - label is length of deletion\ mouseOver Mutation: $change CADD Phred score: $phred\ parent caddSuper1_7 on\ shortLabel CADD 1.7 Del\ track cadd1_7_Del\ type bigBed 9 +\ visibility dense\ cadd1_7_Ins CADD 1.7 Ins bigBed 9 + CADD 1.7 Score: Insertions - label is length of insertion 1 100 100 130 160 177 192 207 0 0 0This track collection shows Combined Annotation Dependent Depletion scores.\ CADD is a tool for scoring the deleteriousness of single nucleotide variants as\ well as insertion/deletion variants in the human genome.
\ \\ Some mutation annotations\ tend to exploit a single information type (e.g., phastCons or phyloP for\ conservation) and/or are restricted in scope (e.g., to missense changes). Thus,\ a broadly applicable metric that objectively weights and integrates diverse\ information is needed. Combined Annotation Dependent Depletion (CADD) is a\ framework that integrates multiple annotations into one metric by contrasting\ variants that survived natural selection with simulated mutations.\
\ \\ CADD scores strongly correlate with allelic diversity, pathogenicity of both\ coding and non-coding variants, experimentally measured regulatory effects,\ and also rank causal variants within individual genome sequences with a higher\ value than non-causal variants. \ Finally, CADD scores of complex trait-associated variants from genome-wide\ association studies (GWAS) are significantly higher than matched controls and\ correlate with study sample size, likely reflecting the increased accuracy of\ larger GWAS.\
\ \\ A CADD score represents a ranking not a prediction, and no threshold is defined\ for a specific purpose. Higher scores are more likely to be deleterious: \ Scores are \ \
10 * -log of the rank\ \ so that variants with scores above 20 are \ predicted to be among the 1.0% most deleterious possible substitutions in \ the human genome. We recommend thinking carefully about what threshold is \ appropriate for your application.\ \ \
\ There are six subtracks of this track: four for single-nucleotide mutations,\ one for each base, showing all possible substitutions, \ one for insertions and one for deletions. All subtracks show the CADD Phred\ score on mouseover. Zooming in shows the exact score on mouseover, same\ basepair = score 0.0.
\\ PHRED-scaled scores are normalized to all potential ~9 billion SNVs, and\ thereby provide an externally comparable unit for analysis. For example, a\ scaled score of 10 or greater indicates a raw score in the top 10% of all\ possible reference genome SNVs, and a score of 20 or greater indicates a raw\ score in the top 1%, regardless of the details of the annotation set, model\ parameters, etc.\
\\ The four single-nucleotide mutation tracks have a default viewing range of\ score 10 to 50. As explained in the paragraph above, that results in\ slightly less than 10% of the data displayed. The \ deletion and insertion tracks have a default filter of 10-100, because they\ display discrete items and not graphical data.\
\ \\ Single nucleotide variants (SNV): For SNVs, at every\ genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing \ the reference allele, e.g., A to A, is always set to zero.\
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and instead of an actual score, the tooltip text will show\ the average score of all nucleotides under the cursor. This is indicated by\ the prefix "~" in the mouseover. Averages of scores are not useful for any\ application of CADD.\
\ \Insertions and deletions: Scores are also shown on mouseover for a\ set of insertions and deletions. On hg38, the set has been obtained from\ gnomAD3. On hg19, the set of indels has been obtained from various sources\ (gnomAD2, ExAC, 1000 Genomes, ESP). If your insertion or deleletion of interest\ is not in the track, you will need to use CADD's\ online scoring tool\ to obtain them.
\ \Track colors
\\ This track is colored according to Table 2 in Vikas et al. The colors represent the recommended ACMG/AMP score cutoffs. \ \
| Range | \Classification | \
|---|---|
| ≥ 25.3 | \Pathogenic | \
| 25.2 - 22.6 | \Neutral | \
| ≤ 22.7 | \Benign | \
\ In CADD version 1.7, new features have been added to improve CADD scores for certain variant\ effects, boosting the overall performance of CADD and bringing new developments to the community.\ CADD v1.7 integrates annotations from recent efforts to assess variant effects, along with new\ conservation and mutation scores.
\\ CADD v1.7 supports only the major chromosomes of the hg38/GRCh38 reference genome (chromosomes 1-22,\ X, and Y) and may be the last version to support the hg19/GRCh37 human reference genome.
\\ This version includes scores derived from Evolutionary Scale Modeling (ESM) for assessing variants\ in protein-coding regions, along with scores from a convolutional neural network (CNN) trained on\ open chromatin sequences, used as a proxy for regulatory regions in the genome. The previously\ included conservation scores have been updated with data from the Zoonomia project. New annotations\ have also been added for 3' Untranslated Regions (3' UTRs), along with models of genome-wide\ mutational rates. The gene and transcript models have been updated by advancing from Ensembl version\ 95 to version 110, and the Ensembl Variant Effect Predictor (VEP) has been upgraded accordingly.
\\ The models in CADD v1.7 have been trained similarly to the version 1.6 release. The logistic\ regression uses an L2 penalty with C = 1, and training was completed after thirteen L-BFGS\ iterations using the sklearn library The new models exhibit a high degree of similarity to the\ previous release, with a Spearman correlation of 0.946 for CADD scores calculated for 100,000\ randomly selected variants between CADD GRCh38-v1.6 and CADD GRCh38-v1.7. The v1.7 models perform\ comparably to earlier versions in distinguishing known pathogenic variants (ClinVar) from common\ variants (gnomAD) across the genome. Improvements in CADD v1.7 are particularly evident when\ focusing on specific variant categories, such as missense or 3' UTR variants, where the latest\ release includes updated annotations.
\\ More information can be found at the\ CADD site\ and the Schubach et al., Nucleic Acids Res, 2024 publication.\ \ \ Data were converted from the files provided on\ the CADD Downloads website,\ provided by the Kircher lab, using\ \ custom Python scripts,\ documented in our \ makeDoc files.\
\ \ \\ CADD scores are freely available for all non-commercial applications from\ the CADD website.\ For commercial applications, see\ the license instructions there.\
\ \\
The CADD data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw, ins.bb and del.bb. Individual\
regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/a.bw stdout\
\
or\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/cadd1.7/ins.bb stdout
\ Thanks to the CADD development team for providing precomputed data as simple tab-separated files.\
\ \\ Kircher M, Witten DM, Jain P, O'Roak BJ, Cooper GM, Shendure J.\ \ A general framework for estimating the relative pathogenicity of human genetic variants.\ Nat Genet. 2014 Mar;46(3):310-5.\ PMID: 24487276;\ PMC: PMC3992975\
\ \\ Rentzsch P, Witten D, Cooper GM, Shendure J, Kircher M.\ \ CADD: predicting the deleteriousness of variants throughout the human genome.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D886-D894.\ PMID: 30371827;\ PMC: PMC6323892\
\ \\ Schubach M, Maass T, Nazaretyan L, Röner S, Kircher M.\ \ CADD v1.7: using protein language models, regulatory CNNs and other nucleotide-level scores to\ improve genome-wide variant predictions.\ Nucleic Acids Res. 2024 Jan 5;52(D1):D1143-D1154.\ PMID: 38183205; PMC: PMC10767851\
\ phenDis 1 bigDataUrl /gbdb/hg38/cadd1.7/ins.bb\ filter.score 10:100\ filterByRange.score on\ filterLabel.score Show only items with PHRED scale score of\ filterLimits.score 0:100\ html caddSuper1_7\ longLabel CADD 1.7 Score: Insertions - label is length of insertion\ mouseOver Mutation: $change CADD Phred score: $phred\ parent caddSuper1_7 on\ shortLabel CADD 1.7 Ins\ track cadd1_7_Ins\ type bigBed 9 +\ visibility dense\ cancerExpr Cancer Gene Expr Gene Expression in 33 TCGA Cancer Tissues (GENCODE v23) 0 100 0 0 0 127 127 127 0 0 0\ \ The Cancer Genome Atlas (TCGA), a collaboration between the\ National Cancer Institute (NCI)\ and \ National Human Genome Research Institute (NHGRI), has generated comprehensive,\ multi-dimensional maps of the key genomic changes in 33 types of cancer. The TCGA\ dataset, 2.5 petabytes of data describing tumor tissue and matched normal tissues from\ more than 11,000 patients, is publically available and has been used widely by the\ research community.
\ \\ The Cancer Genome Atlas is a NIH-funded project to catalog genetic mutations\ responsible for cancer. The data shown here is RNA-seq expression data produced by the\ consortium.
\ \For questions or feedback on the data, please contact \ TCGA.\
\ \\ The gene track shows RNA expression level for each TCGA tissue in GENCODE canonical\ genes. The gene scores are a total of all transcripts in that gene.
\ \\ The transcript track shows RNA expression levels for each TCGA tissue using GENCODE v23\ transcripts.
\ \ \\ In Full and Pack display modes, expression for each genomic item (gene/transcript) is\ represented by a colored bar chart, where the height of each bar represents the median\ expression level across all samples for a tissue, and the bar color indicates the\ tissue.
\
\ The bar chart display has the same width and tissue order for all genomic items.\ Mouse hover over a bar will show the tissue and median expression levels.\ The Squish display mode draws a rectangle for each gene, colored to indicate the tissue\ with highest expression level if it contributes more than 10% to the overall expression\ (and colored black if no tissue predominates).\ In Dense mode, the darkness of the grayscale rectangle displayed for the gene reflects the total\ median expression level across all tissues.
\ \\ This track was designed to be used in conjunction with the GTEx expression tracks that can act as a\ control.
\ \\ The color of each cancer was derived by mapping the tissue of origin to the closest GTEx tissue,\ then taking the GTEx tissue's color. Five cancers did not have a matching GTEx tissue and were\ assigned a rainbow color scheme; these cancers are Cholangiocarcinoma, Esophageal carcinoma, Head\ and Neck squamous cell carcinoma, Sarcoma and Uveal Melanoma.
\ \\ The ordering of the cancers is based on the alphabetical ordering of their GTEx tissues. The five\ cancers that did not match were ordered alphabetically.
\ \TCGA chose cancers for study based on two broad criteria; poor prognosis/overall \ public health impact and availability of human tumor and matched normal tissue samples that meet \ TCGA\ standards.
\ \\ RNA sequencing was performed using a polyA library and the Illumina HiSeq 2000 platform. All RNA\ sequencing was performed by UNC.
\ \\ Sequence reads for this track were quantified to the hg38/GRCh38 human genome using kallisto\ assisted by the GENCODE v23 transcriptome definition. Read quantification was performed at UCSC by\ the Computational Genomics lab, using the \ Toil\ pipeline. The resulting kallisto files were combined to generate a transcript per million (tpm)\ expression matrix using the UCSC tool, kallistoToMatrix. By totaling the TPM values for all\ transcripts associated to the canonical transcript/gene, a condensed gene per million (gpm) matrix\ was made. For both matrices average expression values for each tissue were calculated and used to\ generate a bed6+5 file that is the base of each track. This was done using the UCSC tool,\ expMatrixToBarchartBed. The bed track was then converted to a bigBed file using the UCSC\ tool, bedToBigBed.
\ \\ Data shown here are in whole based upon data generated by the \ TCGA Research Network.\ John Vivian, Melissa Cline, and Benedict Paten of the UCSC Computational Genomics lab were\ responsible for the sequence read quantification used to produce this track. Chris Eisenhart \ and Kate Rosenbloom of the UCSC Genome Browser group were responsible for data file\ post-processing, track configuration and display type.
\ \\ J. Vivian et al., \ \ \ Rapid and efficient analysis of 20,000 RNA-seq samples with Toil\ bioRxiv bioRxiv, vol. 2, p. 62497, 2016.\
\ phenDis 0 group phenDis\ html tcgaExpr\ longLabel Gene Expression in 33 TCGA Cancer Tissues (GENCODE v23)\ shortLabel Cancer Gene Expr\ superTrack on\ track cancerExpr\ tcgaGeneExpr Cancer Gene Expr bigBarChart Gene Expression in 33 TCGA Cancer Tissues (GENCODE v23) 3 100 0 0 0 127 127 127 0 0 0\ \ The Cancer Genome Atlas (TCGA), a collaboration between the\ National Cancer Institute (NCI)\ and \ National Human Genome Research Institute (NHGRI), has generated comprehensive,\ multi-dimensional maps of the key genomic changes in 33 types of cancer. The TCGA\ dataset, 2.5 petabytes of data describing tumor tissue and matched normal tissues from\ more than 11,000 patients, is publically available and has been used widely by the\ research community.
\ \\ The Cancer Genome Atlas is a NIH-funded project to catalog genetic mutations\ responsible for cancer. The data shown here is RNA-seq expression data produced by the\ consortium.
\ \For questions or feedback on the data, please contact \ TCGA.\
\ \\ The gene track shows RNA expression level for each TCGA tissue in GENCODE canonical\ genes. The gene scores are a total of all transcripts in that gene.
\ \\ The transcript track shows RNA expression levels for each TCGA tissue using GENCODE v23\ transcripts.
\ \ \\ In Full and Pack display modes, expression for each genomic item (gene/transcript) is\ represented by a colored bar chart, where the height of each bar represents the median\ expression level across all samples for a tissue, and the bar color indicates the\ tissue.
\
\ The bar chart display has the same width and tissue order for all genomic items.\ Mouse hover over a bar will show the tissue and median expression levels.\ The Squish display mode draws a rectangle for each gene, colored to indicate the tissue\ with highest expression level if it contributes more than 10% to the overall expression\ (and colored black if no tissue predominates).\ In Dense mode, the darkness of the grayscale rectangle displayed for the gene reflects the total\ median expression level across all tissues.
\ \\ This track was designed to be used in conjunction with the GTEx expression tracks that can act as a\ control.
\ \\ The color of each cancer was derived by mapping the tissue of origin to the closest GTEx tissue,\ then taking the GTEx tissue's color. Five cancers did not have a matching GTEx tissue and were\ assigned a rainbow color scheme; these cancers are Cholangiocarcinoma, Esophageal carcinoma, Head\ and Neck squamous cell carcinoma, Sarcoma and Uveal Melanoma.
\ \\ The ordering of the cancers is based on the alphabetical ordering of their GTEx tissues. The five\ cancers that did not match were ordered alphabetically.
\ \TCGA chose cancers for study based on two broad criteria; poor prognosis/overall \ public health impact and availability of human tumor and matched normal tissue samples that meet \ TCGA\ standards.
\ \\ RNA sequencing was performed using a polyA library and the Illumina HiSeq 2000 platform. All RNA\ sequencing was performed by UNC.
\ \\ Sequence reads for this track were quantified to the hg38/GRCh38 human genome using kallisto\ assisted by the GENCODE v23 transcriptome definition. Read quantification was performed at UCSC by\ the Computational Genomics lab, using the \ Toil\ pipeline. The resulting kallisto files were combined to generate a transcript per million (tpm)\ expression matrix using the UCSC tool, kallistoToMatrix. By totaling the TPM values for all\ transcripts associated to the canonical transcript/gene, a condensed gene per million (gpm) matrix\ was made. For both matrices average expression values for each tissue were calculated and used to\ generate a bed6+5 file that is the base of each track. This was done using the UCSC tool,\ expMatrixToBarchartBed. The bed track was then converted to a bigBed file using the UCSC\ tool, bedToBigBed.
\ \\ Data shown here are in whole based upon data generated by the \ TCGA Research Network.\ John Vivian, Melissa Cline, and Benedict Paten of the UCSC Computational Genomics lab were\ responsible for the sequence read quantification used to produce this track. Chris Eisenhart \ and Kate Rosenbloom of the UCSC Genome Browser group were responsible for data file\ post-processing, track configuration and display type.
\ \\ J. Vivian et al., \ \ \ Rapid and efficient analysis of 20,000 RNA-seq samples with Toil\ bioRxiv bioRxiv, vol. 2, p. 62497, 2016.\
\ phenDis 1 barChartBars Adrenocortical_carcinoma Bladder_Urothelial_Carcinoma Brain_Lower_Grade_Glioma Breast_invasive_carcinoma Cervical_squamous_cell_carcinoma_and_endocervical_adenocarcinoma Colon_adenocarcinoma Glioblastoma_multiforme Kidney_Chromophobe Kidney_renal_clear_cell_carcinoma Kidney_renal_papillary_cell_carcinoma Liver_hepatocellular_carcinoma Lung_adenocarcinoma Lung_squamous_cell_carcinoma Lymphoid_Neoplasm_Diffuse_Large_B-cell_Lymphoma Mesothelioma Ovarian_serous_cystadenocarcinoma Pancreatic_adenocarcinoma Pheochromocytoma_and_Paraganglioma Prostate_adenocarcinoma Rectum_adenocarcinoma Skin_Cutaneous_Melanoma Stomach_adenocarcinoma Testicular_Germ_Cell_Tumors Thymoma Thyroid_carcinoma Uterine_Carcinosarcoma Uterine_Corpus_Endometrioid_Carcinoma Cholangiocarcinoma Esophageal_carcinoma Head_and_Neck_squamous_cell_carcinoma Sarcoma Uveal_Melanoma\ barChartColors \\#8FBC8F #8FBC8F #CDB79E #EEEE00 #EEEE00 #00CDCD #EED5D2 \\#CDB79E #CDB79E #CDB79E #CDB79E #CDB79E #CDB79E #9ACD32 #9ACD32 #9ACD32 \\#FFB6C1 #CD9B1D #D9D9D9 #1E90FF #CDB79E #FFD39B #A6A6A6 #008B45 #008B45 \\#EED5D2 #EED5D2 #ff0000 #ff8d00 #ffdb00 #00d619 #009fff\ barChartLabel Cancer types\ barChartMatrixUrl /gbdb/hgFixed/human/expMatrix/tcgaGeneMatrix.tab\ barChartMetric median\ barChartSampleUrl /gbdb/hgFixed/human/expMatrix/tcgaLargeSamples.tab\ barChartUnit GPM\ bigDataUrl /gbdb/hg38/tcga/tcgaGeneExpr.bb\ defaultLabelFields name2\ group phenDis\ html tcgaExpr\ labelFields name2, name\ longLabel Gene Expression in 33 TCGA Cancer Tissues (GENCODE v23)\ maxLimit 8000\ parent cancerExpr\ shortLabel Cancer Gene Expr\ track tcgaGeneExpr\ type bigBarChart\ visibility pack\ tcgaTranscExpr Cancer Transc Expr bigBarChart Transcript-level Expression in 33 TCGA Cancer Tissues (GENCODE v23) 3 100 0 0 0 127 127 127 0 0 0\ \ The Cancer Genome Atlas (TCGA), a collaboration between the\ National Cancer Institute (NCI)\ and \ National Human Genome Research Institute (NHGRI), has generated comprehensive,\ multi-dimensional maps of the key genomic changes in 33 types of cancer. The TCGA\ dataset, 2.5 petabytes of data describing tumor tissue and matched normal tissues from\ more than 11,000 patients, is publically available and has been used widely by the\ research community.
\ \\ The Cancer Genome Atlas is a NIH-funded project to catalog genetic mutations\ responsible for cancer. The data shown here is RNA-seq expression data produced by the\ consortium.
\ \For questions or feedback on the data, please contact \ TCGA.\
\ \\ The gene track shows RNA expression level for each TCGA tissue in GENCODE canonical\ genes. The gene scores are a total of all transcripts in that gene.
\ \\ The transcript track shows RNA expression levels for each TCGA tissue using GENCODE v23\ transcripts.
\ \ \\ In Full and Pack display modes, expression for each genomic item (gene/transcript) is\ represented by a colored bar chart, where the height of each bar represents the median\ expression level across all samples for a tissue, and the bar color indicates the\ tissue.
\
\ The bar chart display has the same width and tissue order for all genomic items.\ Mouse hover over a bar will show the tissue and median expression levels.\ The Squish display mode draws a rectangle for each gene, colored to indicate the tissue\ with highest expression level if it contributes more than 10% to the overall expression\ (and colored black if no tissue predominates).\ In Dense mode, the darkness of the grayscale rectangle displayed for the gene reflects the total\ median expression level across all tissues.
\ \\ This track was designed to be used in conjunction with the GTEx expression tracks that can act as a\ control.
\ \\ The color of each cancer was derived by mapping the tissue of origin to the closest GTEx tissue,\ then taking the GTEx tissue's color. Five cancers did not have a matching GTEx tissue and were\ assigned a rainbow color scheme; these cancers are Cholangiocarcinoma, Esophageal carcinoma, Head\ and Neck squamous cell carcinoma, Sarcoma and Uveal Melanoma.
\ \\ The ordering of the cancers is based on the alphabetical ordering of their GTEx tissues. The five\ cancers that did not match were ordered alphabetically.
\ \TCGA chose cancers for study based on two broad criteria; poor prognosis/overall \ public health impact and availability of human tumor and matched normal tissue samples that meet \ TCGA\ standards.
\ \\ RNA sequencing was performed using a polyA library and the Illumina HiSeq 2000 platform. All RNA\ sequencing was performed by UNC.
\ \\ Sequence reads for this track were quantified to the hg38/GRCh38 human genome using kallisto\ assisted by the GENCODE v23 transcriptome definition. Read quantification was performed at UCSC by\ the Computational Genomics lab, using the \ Toil\ pipeline. The resulting kallisto files were combined to generate a transcript per million (tpm)\ expression matrix using the UCSC tool, kallistoToMatrix. By totaling the TPM values for all\ transcripts associated to the canonical transcript/gene, a condensed gene per million (gpm) matrix\ was made. For both matrices average expression values for each tissue were calculated and used to\ generate a bed6+5 file that is the base of each track. This was done using the UCSC tool,\ expMatrixToBarchartBed. The bed track was then converted to a bigBed file using the UCSC\ tool, bedToBigBed.
\ \\ Data shown here are in whole based upon data generated by the \ TCGA Research Network.\ John Vivian, Melissa Cline, and Benedict Paten of the UCSC Computational Genomics lab were\ responsible for the sequence read quantification used to produce this track. Chris Eisenhart \ and Kate Rosenbloom of the UCSC Genome Browser group were responsible for data file\ post-processing, track configuration and display type.
\ \\ J. Vivian et al., \ \ \ Rapid and efficient analysis of 20,000 RNA-seq samples with Toil\ bioRxiv bioRxiv, vol. 2, p. 62497, 2016.\
\ phenDis 1 barChartBars Adrenocortical_carcinoma Bladder_Urothelial_Carcinoma Brain_Lower_Grade_Glioma Breast_invasive_carcinoma Cervical_squamous_cell_carcinoma_and_endocervical_adenocarcinoma Colon_adenocarcinoma Glioblastoma_multiforme Kidney_Chromophobe Kidney_renal_clear_cell_carcinoma Kidney_renal_papillary_cell_carcinoma Liver_hepatocellular_carcinoma Lung_adenocarcinoma Lung_squamous_cell_carcinoma Lymphoid_Neoplasm_Diffuse_Large_B-cell_Lymphoma Mesothelioma Ovarian_serous_cystadenocarcinoma Pancreatic_adenocarcinoma Pheochromocytoma_and_Paraganglioma Prostate_adenocarcinoma Rectum_adenocarcinoma Skin_Cutaneous_Melanoma Stomach_adenocarcinoma Testicular_Germ_Cell_Tumors Thymoma Thyroid_carcinoma Uterine_Carcinosarcoma Uterine_Corpus_Endometrioid_Carcinoma Cholangiocarcinoma Esophageal_carcinoma Head_and_Neck_squamous_cell_carcinoma Sarcoma Uveal_Melanoma\ barChartColors \\#8FBC8F #8FBC8F #CDB79E #EEEE00 #EEEE00 #00CDCD #EED5D2 \\#CDB79E #CDB79E #CDB79E #CDB79E #CDB79E #CDB79E #9ACD32 #9ACD32 #9ACD32 \\#FFB6C1 #CD9B1D #D9D9D9 #1E90FF #CDB79E #FFD39B #A6A6A6 #008B45 #008B45 \\#EED5D2 #EED5D2 #ff0000 #ff8d00 #ffdb00 #00d619 #009fff\ barChartLabel Cancer types\ barChartMatrixUrl /gbdb/hgFixed/human/expMatrix/tcgaMatrix.tab\ barChartMetric median\ barChartSampleUrl /gbdb/hgFixed/human/expMatrix/tcgaLargeSamples.tab\ barChartUnit TPM\ bigDataUrl /gbdb/hg38/tcga/tcgaTranscExpr.bb\ defaultLabelFields name2\ group phenDis\ html tcgaExpr\ labelFields name2, name\ longLabel Transcript-level Expression in 33 TCGA Cancer Tissues (GENCODE v23)\ maxLimit 8000\ parent cancerExpr\ shortLabel Cancer Transc Expr\ track tcgaTranscExpr\ type bigBarChart\ visibility pack\ ccdsGene CCDS genePred Consensus CDS 0 100 12 120 12 133 187 133 0 0 0\ This track shows human genome high-confidence gene annotations from the\ Consensus \ Coding Sequence (CCDS) project. This project is a collaborative effort \ to identify a core set of \ human protein-coding regions that are consistently annotated and of high \ quality. The long-term goal is to support convergence towards a standard set \ of gene annotations on the human genome.\
\Collaborators include:\
\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ CDS annotations of the human genome were obtained from two sources:\ NCBI \ RefSeq and a union of the gene annotations from \ Ensembl and \ Vega, collectively known \ as Hinxton.
\\ Genes with identical CDS genomic coordinates in both sets become CCDS \ candidates. The genes undergo a quality evaluation, which must be approved by \ all collaborators. The following criteria are currently used to assess each\ gene: \
\ A unique CCDS ID is assigned to the CCDS, which links together all gene \ annotations with the same CDS. CCDS gene annotations are under continuous\ review, with periodic updates to this track.\
\ \\ This track was produced at UCSC from data downloaded from the\ CCDS project \ web site.\
\ \\ Hubbard T, Barker D, Birney E, Cameron G, Chen Y, Clark L, Cox T, Cuff J, Curwen V, Down T et\ al.\ The Ensembl genome database project.\ Nucleic Acids Res. 2002 Jan 1;30(1):38-41.\ PMID: 11752248; PMC: PMC99161\
\\ Pruitt KD, Harrow J, Harte RA, Wallin C, Diekhans M, Maglott DR, Searle S, Farrell CM, Loveland JE,\ Ruef BJ et al.\ \ The consensus coding sequence (CCDS) project: Identifying a common protein-coding gene set for the\ human and mouse genomes.\ Genome Res. 2009 Jul;19(7):1316-23.\ PMID: 19498102; PMC: PMC2704439\
\\ Pruitt KD, Tatusova T, Maglott DR.\ \ NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts\ and proteins.\ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4.\ PMID: 15608248; PMC: PMC539979\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ color 12,120,12\ group genes\ longLabel Consensus CDS\ shortLabel CCDS\ track ccdsGene\ type genePred\ visibility hide\ adult_cpoola_models Cell Line Pool models bigBed 12 + Cell Line Pool transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-CpoolA.bb\ longLabel Cell Line Pool transcript models\ parent sample_models_view on\ shortLabel Cell Line Pool models\ subGroups view=sample_models_view sample=adult_cpoola type=models\ track adult_cpoola_models\ type bigBed 12 +\ visibility squish\ adult_cpoola_ont_post_models Cell Line Pool ONT post models bigBed 12 + Cell Line Pool ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_CpoolA01Rep1.bb\ itemRgb on\ longLabel Cell Line Pool ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Cell Line Pool ONT post models\ subGroups view=per_expr_models_view sample=adult_cpoola type=post_capture_ont_models\ track adult_cpoola_ont_post_models\ type bigBed 12 +\ visibility hide\ adult_cpoola_ont_post_reads Cell Line Pool ONT post reads bam Cell Line Pool ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_CpoolA01Rep1.bam\ longLabel Cell Line Pool ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Cell Line Pool ONT post reads\ subGroups view=per_expr_reads_view sample=adult_cpoola type=post_capture_ont_reads\ track adult_cpoola_ont_post_reads\ type bam\ visibility hide\ adult_cpoola_ont_pre_models Cell Line Pool ONT pre models bigBed 12 + Cell Line Pool ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_CpoolA01Rep1.bb\ itemRgb on\ longLabel Cell Line Pool ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Cell Line Pool ONT pre models\ subGroups view=per_expr_models_view sample=adult_cpoola type=pre_capture_ont_models\ track adult_cpoola_ont_pre_models\ type bigBed 12 +\ visibility hide\ adult_cpoola_ont_pre_reads Cell Line Pool ONT pre reads bam Cell Line Pool ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_CpoolA01Rep1.bam\ longLabel Cell Line Pool ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Cell Line Pool ONT pre reads\ subGroups view=per_expr_reads_view sample=adult_cpoola type=pre_capture_ont_reads\ track adult_cpoola_ont_pre_reads\ type bam\ visibility hide\ adult_cpoola_pacbio_post_models Cell Line Pool PB post models bigBed 12 + Cell Line Pool PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_CpoolA01Rep1.bb\ itemRgb on\ longLabel Cell Line Pool PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Cell Line Pool PB post models\ subGroups view=per_expr_models_view sample=adult_cpoola type=post_capture_pacbio_models\ track adult_cpoola_pacbio_post_models\ type bigBed 12 +\ visibility hide\ adult_cpoola_pacbio_post_reads Cell Line Pool PB post reads bam Cell Line Pool PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_CpoolA01Rep1.bam\ longLabel Cell Line Pool PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Cell Line Pool PB post reads\ subGroups view=per_expr_reads_view sample=adult_cpoola type=post_capture_pacbio_reads\ track adult_cpoola_pacbio_post_reads\ type bam\ visibility hide\ adult_cpoola_pacbio_pre_models Cell Line Pool PB pre models bigBed 12 + Cell Line Pool PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_CpoolA01Rep1.bb\ itemRgb on\ longLabel Cell Line Pool PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Cell Line Pool PB pre models\ subGroups view=per_expr_models_view sample=adult_cpoola type=pre_capture_pacbio_models\ track adult_cpoola_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ adult_cpoola_pacbio_pre_reads Cell Line Pool PB pre reads bam Cell Line Pool PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_CpoolA01Rep1.bam\ longLabel Cell Line Pool PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Cell Line Pool PB pre reads\ subGroups view=per_expr_reads_view sample=adult_cpoola type=pre_capture_pacbio_reads\ track adult_cpoola_pacbio_pre_reads\ type bam\ visibility hide\ gnomADPextCells_Culturedfibroblasts Cells-Cultured Fibroblasts bigWig 0 1 gnomAD pext Cells-Cultured Fibroblasts 0 100 170 238 255 212 246 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Cells_Culturedfibroblasts.bw\ color 170,238,255\ longLabel gnomAD pext Cells-Cultured Fibroblasts\ parent gnomadPext off\ shortLabel Cells-Cultured Fibroblasts\ track gnomADPextCells_Culturedfibroblasts\ visibility hide\ gnomADPextCells_EBV_transformedlymphocytes Cells-EBV-transformed Lymphocytes bigWig 0 1 gnomAD pext Cells-EBV-transformed Lymphocytes 0 100 204 102 255 229 178 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Cells_EBV_transformedlymphocytes.bw\ color 204,102,255\ longLabel gnomAD pext Cells-EBV-transformed Lymphocytes\ parent gnomadPext off\ shortLabel Cells-EBV-transformed Lymphocytes\ track gnomADPextCells_EBV_transformedlymphocytes\ visibility hide\ centromeres Centromeres bed 4 . Centromere Locations 0 100 255 0 0 255 127 127 0 0 24 chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX,chrY, https://www.ncbi.nlm.nih.gov/nuccore/$$\ Track indicating the location of the centromere sequences.\ Centromeres are specialized chromatin structures that are required for cell division. These\ genomic regions are normally defined by long tracts of tandem repeats, or satellite DNA, that\ contain a limited number of sequence differences to distinguish the linear order of repeat copies.\ The size and repetitive nature of these regions mean they are typically not represented in\ reference assemblies. Unlike all previous versions of the human reference assembly, where the\ centromere regions have been represented by a multi-megabase gap, GRCh38 incorporates centromere\ reference models that provide an initial genomic description derived from chromosome-assigned whole\ genome shotgun (WGS) read libraries of alpha satellite.\
\ \\ Each reference model provides an approximation of the true array sequence organization.\ Although the long-range repeat ordering is not expected to represent the true organization,\ the submissions are expected to provide a biologically rich description of array variants and\ local-monomer organization as observed in the initial WGS read dataset. As a result, these\ sequences serve as a useful mapping target to extend sequence-based studies to sites previously\ omitted from the human reference genome.\
\ \\ The sequences are generated based on second-order Markov models of monomer\ variants, and graphical models of larger scale higher order repeats.\ The graphical models are based on an analysis of Sanger reads from the\ HuRef sequencing project (Assembly\ GCA_000002125.1; BioProject\ PRJNA19621),\ and their local-ordering is supported by observed same-read monomer\ adjacencies. The Markov models are generated by the program linearSat, which\ was written for this project and that also generates a linear representation\ of monomer order. The software linearSat generates a second-order Markov\ chain to the size of a given array provided by sequence coverage normalization\ estimates. The sequence definitions of transposable element insertions are\ limited to the sequences directly adjacent to alpha satellite within the read\ database, and incomplete representations are noted with an adjacent\ 100 bp gap. In total, these sequences provide a more complete reference\ of sequence composition and higher order repeat variation inherent to a\ given alpha satellite array, used to assemble centromeric regions of the\ human chromosomes.\
\ \\ The data for this track was supplied by\ Karen Miga.\
\ \\ Miga KH, Newton Y, Jain M, Altemose N, Willard HF, Kent WJ.\ \ Centromere reference models for human chromosomes X and Y satellite arrays.\ Genome Res. 2014 Apr;24(4):697-707.\ PMID: 24501022; PMC: PMC3975068\
\ map 1 chromosomes chr1,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chrX,chrY\ color 255,0,0\ group map\ longLabel Centromere Locations\ shortLabel Centromeres\ track centromeres\ type bed 4 .\ url https://www.ncbi.nlm.nih.gov/nuccore/$$\ urlLabel NCBI accession record:\ visibility hide\ hprcChainNetViewchain Chains bed 3 Human Genomes, Chain/Net pairwise alignments, as mapped by the HPRC project 3 100 0 0 0 255 255 0 1 0 0 hprc 1 longLabel Human Genomes, Chain/Net pairwise alignments, as mapped by the HPRC project\ parent hprcChainNet\ shortLabel Chains\ spectrum on\ track hprcChainNetViewchain\ view chain\ visibility pack\ chm13LiftOver CHM13 alignments bigChain GCA_009914755.4 CHM13 (GCA_009914755.4) v1_nfLO liftOver alignments 0 100 120 20 0 187 137 127 0 0 0\ These tracks show the one-to-one v1_nfLO alignments of the GRCh38/hg38 to the\ T2T-CHM13 v2.0 assembly.\
\ \\ The track displays boxes joined together by either single or double lines,\ with the boxes represent aligning regions, single lines indicating gaps that\ are largely due to a deletion in the CHM13 v2.0 assembly or an insertion in\ the GRCh38/hg38, and double lines representing more complex gaps that involve\ substantial sequence in both assembly.\
\ \ \\
\ To prevent ambiguous alignments, all false duplications, as determined by the Genome in a Bottle Consortium\ (GCA_000001405.15_GRCh38_GRC_exclusions_T2Tv2.bed), \ as well as the GRCh38 modeled centromeres,\ were masked from the GRCh38/hg38 primary assembly. In addition, unlocalized and unplaced (random) contigs were removed.\
\ \\ For the minimap2-based pipeline, the initial chain file was generated using\ nf-LO v1.5.1 with\ minimap2 v2.24 alignments. These \ chains were then split at all locations that contained unaligned segments greater than 1kbp or \ gaps greater than 10kbp. Split chain files were then converted to PAF format\ with extended CIGAR strings using chaintools (v0.1),\ and alignments between nonhomologous chromosomes were removed. The trim-paf operation of\ rustybam (v0.1.29) \ was next used to remove overlapping alignments \ in the query sequence, and then the target sequence, to create 1:1 alignments. PAF alignments \ were converted back to the chain format with paf2chain commit f68eeca, and finally, \ chaintools was used to generate the inverted chain file.\
\ \\ Full commands with parameters used were:\
\\
nextflow run main.nf --source GRCh38.fa --target chm13v2.0.fasta --outdir dir -profile local --aligner minimap2\
python chaintools/src/split.py -c input.chain -o input-split.chain\
python chaintools/src/to_paf.py -c input-split.chain -t target.fa -q query.fa -o input-split.paf\
awk '$1==$6' input-split.paf | rb break-paf --max-size 10000 | rb trim-paf -r | rb invert | rb trim-paf -r | rb invert > out.paf\
paf2chain -i out.paf > out.chain\
python chaintools/src/invert.py -c out.chain -o out_inverted.chain\
\
\
\
The above process does not add chain ids or scores. The UCSC utilities\
chainMergeSort and chainScore are used to update the\
chains:\
\
\
chainMergeSort out.chain | chainScore stdin chm13v2.0.2bit hg38.2bit chm13v2.0-hg38.chain\
chainMergeSort out_inverted.chain | chainScore stdin hg38.2bit chm13v2.0.2bit hg38-chm13v2.0.chain\
\
\
\
\ Rustybam trim-paf\ uses dynamic programming and the CIGAR string to find an optimal\ splitting point between overlapping alignments in the query sequence. It\ starts its trimming with the largest overlap and then recursively trims\ smaller overlaps.\
\ \\ Results were validated by using chaintools to confirm that there were no\ overlapping sequences with respect to both CHM13v2.0 and GRCh38 in the\ released chain file. In addition, trimmed alignments were visually inspected\ with SafFire to confirm their quality.\
\ \\ Chains were swapped to make GRCh38/hg38 the target.\
\ \ \\ The v1_nflo chains were generated by Nae-Chyun Chen<naechyun.chen@gmail.com>\ and Mitchell Vollger<mvollger@uw.edu>\
\ \\
Nurk S, Koren S, Rhie A, Rautiainen M, et al. The complete sequence of a human genome. bioRxiv, 2021.
\ \ compGeno 1 bigDataUrl /gbdb/hg38/bbi/chm13LiftOver/hg38-chm13v2.ncbi-qnames.over.chain.bb\ color 120,20,0\ group compGeno\ linkDataUrl /gbdb/hg38/bbi/chm13LiftOver/hg38-chm13v2.ncbi-qnames.over.link.bb\ longLabel CHM13 (GCA_009914755.4) v1_nfLO liftOver alignments\ shortLabel CHM13 alignments\ track chm13LiftOver\ type bigChain GCA_009914755.4\ visibility hide\ cytoBand Chromosome Band bed 4 + Chromosome Bands Localized by FISH Mapping Clones 0 100 0 0 0 127 127 127 0 0 0\ The chromosome band track represents the approximate \ location of bands seen on Giemsa-stained chromosomes.\ Chromosomes are displayed in the browser with the short arm first. \ Cytologically identified bands on the chromosome are numbered outward \ from the centromere on the short (p) and long (q) arms. At low resolution, \ bands are classified using the nomenclature \ [chromosome][arm][band], where band is a \ single digit. Examples of bands on chromosome 3 include 3p2, 3p1, cen, 3q1, \ and 3q2. At a finer resolution, some of the bands are subdivided into \ sub-bands, adding a second digit to the band number, e.g. 3p26. This \ resolution produces about 500 bands. A final subdivision into a \ total of 862 sub-bands is made by adding a period and another digit to the \ band, resulting in 3p26.3, 3p26.2, etc.
\ \\ Chromosome band information was downloaded from NCBI\ using the ideogram.gz file for the respective assembly. These data were then \ transformed into our visualization format. See our \ assembly creation documentation for the organism of interest\ to see the specific steps taken to transform these data.\ Band lengths are typically estimated based on FISH or other\ molecular markers interpreted via microscopy.
\\ For some of our older assemblies, greater than 10 years old, the tracks were\ created as detailed below and in Furey and Haussler, 2003.
\\ Barbara Trask, Vivian Cheung, Norma Nowak and others in the BAC Resource\ Consortium used fluorescent in-situ hybridization (FISH) to determine a \ cytogenetic location for large genomic clones on the chromosomes.\ The results from these experiments are the primary source of information used\ in estimating the chromosome band locations.\ For more information about the process, see the paper, Cheung,\ et al., 2001. and the accompanying web site,\ Human BAC Resource.
\\ BAC clone placements in the human sequence are determined at UCSC using a \ combination of full BAC clone sequence, BAC end sequence, and STS marker \ information.
\ \\ We would like to thank all the labs that have contributed to this resource:\
\ Cheung VG, Nowak N, Jang W, Kirsch IR, Zhao S, Chen XN, Furey TS, Kim UJ, Kuo WL, Olivier M et\ al.\ \ Integration of cytogenetic landmarks into the draft sequence of the human genome.\ Nature. 2001 Feb 15;409(6822):953-8.\ PMID: 11237021\
\ \\ Furey TS, Haussler D.\ \ Integration of the cytogenetic map with the draft human genome sequence.\ Hum Mol Genet. 2003 May 1;12(9):1037-44.\ PMID: 12700172\
\ \ map 1 group map\ longLabel Chromosome Bands Localized by FISH Mapping Clones\ shortLabel Chromosome Band\ track cytoBand\ type bed 4 +\ visibility hide\ cytoBandIdeo Chromosome Band (Ideogram) bed 4 + Chromosome Bands Localized by FISH Mapping Clones (for Ideogram) 1 100 0 0 0 127 127 127 0 0 0 map 1 group map\ longLabel Chromosome Bands Localized by FISH Mapping Clones (for Ideogram)\ shortLabel Chromosome Band (Ideogram)\ track cytoBandIdeo\ type bed 4 +\ visibility dense\ civic CIViC bigBed 12 + CIViC - Expert & crowd-sourced cancer variant interpretation 0 100 0 0 0 127 127 127 0 0 0\ This track shows genomic locations for variants in the\ CIViC (Clinical\ Interpretation of Variants in Cancer) database. These clinically\ relevant variant interpretations are expert and crowd-sourced from\ peer-reviewed literature, clinical trials, and some conference\ abstracts.\
\ \\ Each variant's interpretation is in the context of a broader molecular\ profile: one or more variants grouped together. For example, clinical\ evidence may be relevant to a KRAS G12 mutation on its own, but other\ clinical evidence may relevant for cases with either a mutation in\ KRAS G12 or G13.\
\ \\ The primary points of data from the scientific literature are curated\ as Clinical Evidence, which connects to a molecular profile, which in\ turn connects to the variants shown in this track. Groups of evidence\ can become curator Assertions about the relevence of a molecular\ profile.\
\ \\ The detail for a feature will list diseases and therapies that have\ been associated with a genomic variant. Visiting the CIViC page for a\ variant will allow browsing the Molecular Profiles associated with\ that variant, and in turn each Molecular Profile shows the Clinical\ Evidence and Assertions for various diseases and therapies.\
\ \ \\ There are three types of variant feature types in CIViC: gene, fusion, and\ factor, of which only the gene and fusion fetaures have a genomic location.\
\ \\ Gene variants are shown as a single item, with a name indicating the\ variant's mode: sequence change, gene expression, gene deletion,\ etc.\
\ \\ Fusion variants connect two genes via a structural DNA rearrangement,\ typically in the introns or promotors of genes. For CIViC fusions that\ have an annotated transcript and exon, the exon will be shown as a\ thick bar. If there is an intron associated with the fusion, it will\ be annotated as a thin bar on the feature.\
\ \\ This track reflects the monthly data summaries published by CIViC. The\ latest information is always available directly on the CIViC website\ or by its API.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be\ accessed from scripts through our API, via the track name\ "civic".\
\ \\ The monthly CIViC Variant Summaries were reformatted at UCSC\ to bigBed format. The\ data is updated every month, the week after CIViC data summary\ release. The diseases and therapies associated with a variant are\ collected from the corresponding TSV files from CIViC, using the\ molecular profile summaries as a mapping.\
\ \\ Thanks to the CIViC contributors and organizers for curating the\ database and making the data available for download.\
\ \\ Griffith M, Spies NC, Krysiak K, McMichael JF, Coffman AC, Danos AM, Ainscough BJ, Ramirez CA,\ Rieke DT, Kujan L et al. CIViC is a community knowledgebase for expert crowdsourcing the clinical\ interpretation of variants in cancer. Nat Genet. 2017Jan31;49(2):170-174. PMID:\ 28138153; PMC:\ PMC5367263\
\ \ phenDis 1 bigDataUrl /gbdb/hg38/civic/civic.bb\ group phenDis\ longLabel CIViC - Expert & crowd-sourced cancer variant interpretation\ mouseOverField mouseOverHTML\ shortLabel CIViC\ track civic\ type bigBed 12 +\ urls origVariant="https://civicdb.org/variants/$$/summary" alleleRegistryId="https://reg.clinicalgenome.org/redmine/projects/registry/genboree_registry/by_canonicalid?canonicalid=$$" clinvarId="https://www.ncbi.nlm.nih.gov/clinvar/variation/$$/" diseaseLink="https://www.disease-ontology.org/?id=DOID:$$"\ clinGenComp ClinGen bigBed 9 + ClinGen curation activities (Dosage Sensitivity and Gene-Disease Validity) 0 100 0 0 0 127 127 127 0 0 0\ \
NOTE:
\
These data are for research purposes only. While the ClinGen data are \
open to the public, users seeking information about a personal medical or \
genetic condition are urged to consult with a qualified physician for \
diagnosis and for answers to personal medical questions.\
\ UCSC presents these data for use by qualified professionals, and even \ such professionals should use caution in interpreting the significance of \ information found here. No single data point should be taken at face \ value and such data should always be used in conjunction with as much \ corroborating data as possible. No treatment protocols should be \ developed or patient advice given on the basis of these data without \ careful consideration of all possible sources of information.\
\\ No attempt to identify individual patients should \ be undertaken. No one is authorized to attempt to identify patients \ by any means.\
\ \\ The Clinical Genome Resource (ClinGen)\ tracks display data generated from several key curation activities related to gene-disease validity,\ dosage sensitivity, and variant pathogenicity.\ ClinGen is a National Institute of Health (NIH)-funded initiative dedicated to \ identifying clinically relevant genes and variants for use in precision medicine and research. \ This is accomplished by harnessing the data from both research efforts and clinical genetic \ testing and using it to propel expert and machine-driven curation activities. \ ClinGen works closely with the National Center for Biotechnology Information (NCBI) of the \ National Library of Medicine (NLM)\ which distributes part of this information through its ClinVar database.\
\ \\ The available data tracks are:\
\ A rating system is used to classify the evidence supporting or refuting dosage\ sensitivity for individual genes and regions, which takes in consideration the following criteria:\ number of causative variants reported, patterns of inheritance, consistency of phenotype, evidence\ from large-scale case-control studies, mutational mechanisms, data from public genome variation \ databases, and expert consensus opinion.\
\\ The system is intended to be of a "dynamic nature", with regions being reevaluated periodically to \ incorporate emerging evidence. The evidence collected is displayed within a publicly available \ database. \ Evidence that haploinsufficiency or triplosensitivity of a gene is associated with a specific \ phenotype will aid in the interpretive assessment of CNVs including that gene or genomic region.\
\\ Similarly, a qualitative classification system is used to correlate the evidence of \ a gene-disease relationship: "Definitive", "Strong", "Moderate", \ "Limited", "Animal Model Only", \ "No Known Disease Relationship", "Disputed", or "Refuted".\
\ \\ Items are shaded according to dosage sensitivity type, red \ for haploinsufficiency score 3, blue for triplosensitivity score 3, \ and grey for other evidence scores or \ not yet evaluated).\ Mouseover on items shows the supporting evidence of dosage sensitivity.\ Tracks can be filtered according to the supporting evidence of dosage sensitivity.\ \
\ Dosage Scores are used to classify the evidence of the supporting dosage sensitivity map:\
\ For more information on the use of the scores see the ClinGen\ FAQs.\
\ \\ The gene-disease validity classifications are labeled with the disease entity and hovering \ over the features shows the associated gene. Items are color coded based on the strength of their \ classification as provided below:\
\| Color | \Classifications | \
|---|---|
| \ | Definitive: The role of this gene in this particular disease has been \ repeatedly demonstrated and has been upheld over time | \
| \ | Strong: The role of this gene in disease has been independently\ demonstrated typically in at least two separate studies, including both strong variant-level\ evidence in unrelated probands and compelling gene-level evidence from experimental data | \
| \ | Moderate: There is moderate evidence to support a causal role for this\ gene in this disease, typically including both several probands with variants and moderate \ experimental data supporting the gene-disease assertion | \
| \ | Limited: There is limited evidence to support a causal role for this \ gene in this disease, such as few probands with variants and limited experimental data supporting \ the gene-disease assertion | \
| \ | Animal Model Only: There are no published human probands with variants \ but there is animal model data supporting the gene-disease assertion | \
| \ | No Known Disease Relationship: Evidence for a causal role in disease \ has not been reported | \
| \ | Disputed: Conflicting evidence disputing a role for this gene in this \ disease has arisen since the initial report identifying an association between the gene and disease | \
| \ | Refuted: Evidence refuting the role of the gene in the specified \ disease has been reported and significantly outweighs any evidence supporting the role | \
\ The version of the ClinGen Standard Operating Procedures (SOPs) that each gene-disease \ classification was performed with is provided as well. An older or newer SOP version does not \ necessarily mean the classification is any more or less valid but is provided for clarity. \ Each details page also contains a direct link to an evidence summary detailing the rationale behind\ the specific classification and information such as a breakdown of the semi-qualitative framework, \ relevant PubMed IDs, the type of data (Genetic vs Experimental Evidence), and a detailed summary.\
\ \\ These tracks are multi-view composite tracks that contain multiple data types (views). Each view \ within a track has separate display controls, as described \ here.\
\ \\ Item names correspond to the VCEP loci, usually the gene symbol. Mouseovers display the disease with a\ link to the CSpec, the VCEP panel with a link to the ClinGen VCEP page, and the current expert panel status.
\ \\ The raw data can be explored interactively with the Table Browser,\ or the Data Integrator. For automated analysis, the data may \ be queried from our REST API. Please refer to our \ mailing list archives\ for questions, or our Data Access FAQ for more\ information.\
\ \\ Data is also freely available on the ClinGen website \ (gene-disease curation methods) \ and FTP (dosage curations). \
\ \ \\ Thank you to ClinGen and NCBI, especially Erin Rooney Riggs, Christa Lese Martin, Tristan Nelson,\ May Flowers, Scott Goehringer, and Phillip Weller for technical coordination and \ consultation, and to Christopher Lee, Luis Nassar, and Anna Benet-Pages of the Genome \ Browser team.\
\ \\ Rehm HL, Berg JS, Brooks LD, Bustamante CD, Evans JP, Landrum MJ, Ledbetter DH, Maglott DR, Martin\ CL, Nussbaum RL et al.\ \ ClinGen--the Clinical Genome Resource.\ N Engl J Med. 2015 Jun 4;372(23):2235-42.\ PMID: 26014595; PMC: PMC4474187\
\ \\ Richards S, Aziz N, Bale S, Bick D, Das S, Gastier-Foster J, Grody WW, Hegde M, Lyon E, Spector E\ et al.\ \ Standards and guidelines for the interpretation of sequence variants: a joint consensus\ recommendation of the American College of Medical Genetics and Genomics and the Association for\ Molecular Pathology.\ Genet Med. 2015 May;17(5):405-24.\ PMID: 25741868; PMC: PMC4544753\
\ \\ Riggs ER, Church DM, Hanson K, Horner VL, Kaminsky EB, Kuhn RM, Wain KE, Williams ES, Aradhya S,\ Kearney HM et al.\ \ Towards an evidence-based process for the clinical interpretation of copy number variation.\ Clin Genet. 2012 May;81(5):403-12.\ PMID: 22097934; PMC: PMC5008023\
\ \\ Strande NT, Riggs ER, Buchanan AH, Ceyhan-Birsoy O, DiStefano M, Dwight SS, Goldstein J, Ghosh R,\ Seifert BA, Sneddon TP et al.\ \ Evaluating the Clinical Validity of Gene-Disease Associations: An Evidence-Based Framework Developed\ by the Clinical Genome Resource.\ Am J Hum Genet. 2017 Jun 1;100(6):895-906.\ PMID: 28552198; PMC: PMC5473734\
\ \ phenDis 1 compositeTrack on\ dataVersion /gbdb/$D/bbi/clinGen/clinGenVersion.txt\ group phenDis\ html clinGen\ itemRgb on\ longLabel ClinGen curation activities (Dosage Sensitivity and Gene-Disease Validity)\ noParentConfig on\ shortLabel ClinGen\ track clinGenComp\ type bigBed 9 +\ visibility hide\ iscaComposite ClinGen CNVs bed 3 Clinical Genome Resource (ClinGen) CNVs 0 100 0 0 0 127 127 127 0 0 0The ClinGen CNVs track is no longer being updated. These data, along with updates,\ can be found in the \ ClinVar Copy Number Variants (ClinVar CNVs) track.
\\ See our \ news archive for more information.
\\
NOTE:
\
These data are for research purposes only. While the ClinGen data are\
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal medical questions.\
UCSC presents these data for use by qualified professionals, and even\ such professionals should use caution in interpreting the significance of \ information found here. No single data point should be taken at face \ value and such data should always be used in conjunction with as much \ corroborating data as possible. No treatment protocols should be \ developed or patient advice given on the basis of these data without \ careful consideration of all possible sources of information.\
\ \No attempt to identify individual patients should\ be undertaken. No one is authorized to attempt to identify patients \ by any means.\
\ \\ The Clinical Genome Resource (ClinGen)\ is a National Institutes of Health (NIH)-funded program dedicated to building a genomic\ knowledge base to improve patient care. \ This will be accomplished by harnessing the data from both research efforts and clinical genetic\ testing, and using it to propel expert and machine-driven curation activities. \ By facilitating collaboration within the genomics community,\ we will all better understand the relationship between genomic variation and human health. \ ClinGen will work closely with the National\ Center for Biotechnology Information (NCBI) of the National Library of Medicine (NLM), \ which will distribute this information through its\ ClinVar database.\
\ \\ The ClinGen dataset displays clinical microarray data submitted to dbGaP/dbVar at NCBI\ by ClinGen member laboratories (dbVar study\ nstd37),\ as well as clinical data reported in Kaminsky et al., 2011 (dbVar study\ ntsd101)\ (see reference below). This track shows copy number variants (CNVs) found in patients referred\ for genetic testing for indications such as intellectual disability, developmental delay,\ autism and congenital anomalies. Additionally, the ClinGen "Curated Pathogenic" and\ "Curated Benign" tracks represent genes/genomic regions reviewed for dosage sensitivity\ in an evidence-based manner by the ClinGen Structural Variation Working Group (dbVar study\ nstd45).\
\ \The CNVs in this study have been reviewed for their clinical significance by\ the submitting ClinGen laboratory. Some of the deletions and duplications in the track\ have been reported as causative for a phenotype by the submitting clinical \ laboratory; this information was based on current knowledge at the time of submission.\ However, it should be noted that phenotype information is often vague and imprecise and\ should be used with caution. While all samples were submitted because of a phenotype in \ a patient, only 15% of patients had variants determined to be causal, \ and most patients will have additional variants that are not causal.\
\ \CNVs are separated into subtracks and are labeled as:\
Two subtracks, "Path Gain" and "Path Loss", are aggregate tracks\ showing graphically the accumulated level of gains and losses in the \ Pathogenic subtrack across the genome. Similarly, "Benign Gain" and\ "Benign Loss" show the accumulated level of gains and losses in the\ Benign subtrack. These tracks are collectively called "Coverage"\ tracks.\
\ \Many samples have multiple variants, not all of which are causative \ of the phenotype. The CNVs in these samples have been decoupled, so it is not\ possible to connect multiple imbalances as coming from a single patient.\ It is therefore not possible to identify individuals via their genotype. \
\ \ \\ The samples were analyzed by arrays from patients referred for \ cytogenetic testing due to clinical phenotypes. Samples were analyzed with a \ probe spacing of 20-75 kb. The minimum CNV breakpoints are shown; if available,\ the maximum CNV breakpoints are provided in the details page, but are not shown \ graphically on the Browser image.\
\ \Data were submitted to \ dbGaP at NCBI and thence decoupled as described into\ dbVar for unrestricted release.\
\ \\ The entries are colored red for loss and \ blue for gain. The names of items use the \ ClinVar convention of appending "_inheritance" indicating the mechanism of \ inheritance, if known: "_pat, _mat, _dnovo, _unk" as paternal, maternal, \ de novo and unknown, respectively. \
\ \\ Most data were validated by the submitting laboratory using various methods, \ including FISH, G-banded karyotype, MLPA and qPCR.\
\ \\ Thank you to ClinGen and NCBI for technical coordination and consultation, and to\ the UCSC Genome Browser staff for engineering the track display.\
\ \\ Miller DT, Adam MP, Aradhya S, Biesecker LG, Brothman AR, Carter NP, Church DM, Crolla JA, Eichler\ EE, Epstein CJ et al.\ \ Consensus statement: chromosomal microarray is a first-tier clinical diagnostic test for individuals\ with developmental disabilities or congenital anomalies.\ Am J Hum Genet. 2010 May 14;86(5):749-64.\ PMID: 20466091; PMC: PMC2869000\
\ \\ Kaminsky EB, Kaul V, Paschall J, Church DM, Bunke B, Kunig D, Moreno-De-Luca D, Moreno-De-Luca A,\ Mulle JG, Warren ST et al.\ \ An evidence-based approach to establish the functional and clinical significance of copy number\ variants in intellectual and developmental disabilities.\ Genet Med. 2011 Sep;13(9):777-84.\ PMID: 21844811; PMC: PMC3661946\
\ phenDis 1 compositeTrack on\ dimensions dimensionY=class dimensionX=level\ group phenDis\ longLabel Clinical Genome Resource (ClinGen) CNVs\ pennantIcon snowflake.png /goldenPath/newsarch.html#093020b "ClinGen CNV data are now updated on ClinVar Variants track. See news archive for details."\ shortLabel ClinGen CNVs\ sortOrder class=+ level=+ view=+\ subGroup1 view Views cov=Coverage cnv=CNVs dose=Dose\ subGroup2 class Class path=Pathogenic likP=Likely_Pathogenic unc=Uncertain likB=Likely_Benign ben=Benign\ subGroup3 level Evidence cur=Curated sub=Submitted\ track iscaComposite\ type bed 3\ visibility hide\ clinGenCspec ClinGen VCEP Specifications bigBed 9 + Clingen CSpec Variant Interpretation VCEP Specifications 3 100 0 0 0 127 127 127 0 0 0 phenDis 1 bigDataUrl /gbdb/hg38/bbi/clinGen/clinGenCspec.bb\ longLabel Clingen CSpec Variant Interpretation VCEP Specifications\ mouseOver Disease: $diseaseNOTE:
\
ClinVar is intended for use primarily by physicians and other\
professionals concerned with genetic disorders, by genetics researchers, and\
by advanced students in science and medicine. Research data is not easy to interpret, and not\
everything shown is necessarily useful. While the ClinVar\
database is open to all academic users, users seeking information about a\
personal medical or genetic condition are urged to consult with a qualified\
physician for diagnosis and for answers to personal questions.
\ These tracks show the genomic positions of variants in the\ ClinVar database. \ ClinVar is a free, public archive of reports\ of the relationships among human variations and phenotypes, with supporting\ evidence.
\ \\ The ClinVar SNVs track displays substitutions and indels shorter than 50 bp, and \ the ClinVar CNVs track displays copy number variants (CNVs) equal to or larger than 50 bp.\
\ \\ The ClinVar Interpretations track displays the genomic positions of individual variant \ submissions and interpretations of the clinical significance and their relationship to disease in \ the ClinVar database.\
\ \\ Note on the start position of variants: The data in the track are obtained directly from ClinVar's FTP site.\ We display the data obtained from ClinVar as-is to avoid discrepancies between UCSC and NCBI. \ However, be aware that the ClinVar conventions are different from the VCF standard. \ Variants may be right-aligned or may contain additional context, e.g. for\ inserts. The VCF position is also available in this track,\ as an additional field, at the end of the list of fields, when you click any variant.\ It can be extracted using our table browser, the API,\ or the bigBedToBed tool (see the Data access section below). \ And GnomAD has a converter.\
\ \\ Items can be filtered according to the size of the variant, variant type, clinical significance,\ allele origin, phenotype, and molecular consequence, using the track Configure options.\ Each subtrack has separate display controls, as described\ here.\
\ \\ Entries in the ClinVar SNVs and ClinVar Interpretations tracks are colored by clinical \ significance:\
\ Entries in the ClinVar CNVs track are colored by type of variant, among others:\
\ In the ClinVar SNV track, an option to show triangles for protein-truncating mutations is available\ under the Decoration settings, using the Glyph decoration placement option. Triangles can be placed\ using either the Overlay or Adjacent display. Variants with the following molecular consequences\ are considered protein-truncating: nonsense, frameshift variant, splice acceptor variant, and\ splice donor variant.
\ \\ Mouseover on the genomic locations of ClinVar variants shows variant details, clinical \ interpretation, and associated conditions. Further information on each variant is displayed on \ the details page by clicking onto any variant. ClinVar is an archive for assertions of clinical \ significance made by the submitters. The level of review supporting the assertion of clinical \ significance for the variation is reported as the \ review status. \ Stars (0 to 4) provide a graphical representation of the aggregate review status. \
\ \\ The variants in the ClinVar Interpretations track are arranged from top to bottom by the variant \ classification of each submission:\
\ More information about using and understanding the ClinVar data can be found \ here.\
\ \\ For the human genome version hg19, the hg19 genome released by UCSC in 2009 had a \ mitochondrial genome "chrM" that was not the same as the one later used for most\ databases like ClinVar. As a result, we added the official mitochondrial genome\ in 2020 as "chrMT", and all mitochondrial annotations of ClinVar and most other\ databases are shown on the mitochondrial genome called "chrMT". For a full description\ of the issue of the mitochondrial genome in hg19, please see the \ hg19 README file \ on our download site. \
\ \ \ClinVar tries to publish a new release on the \ first Thursday of every month. \ In practice, the exact day can move by a few days.\ Our track is updated on the day after any ClinVar release, and copied to our public site one day later.\ The exact date of our last update is shown on the track configuration page. \ You can find the previous versions of the track organized by month on our\ downloads server in the \ archive\ directory. To display a previous version of the track, paste the URL to one of\ the older files into the custom track text input field under "My Data > Custom Tracks".
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be\ accessed from scripts through our API, the track names are\ "clinVarMain" and "clinVarCnv".\ \
\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ The files for this track are called clinvarMain.bb and clinvarCnv.bb. Individual\ regions or the whole genome annotation can be obtained using our tool bigBedToBed,\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool\ can also be used to obtain only features within a given range, e.g. \ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg19/bbi/clinvar/clinvarMain.bb -chrom=chr21 -start=0 -end=100000000 stdout\
\ \\ ClinVar files were reformatted at UCSC to the bigBed format.\ The data is updated every month, one week after the ClinVar release date.\ The program that performs the update is available on\ GitHub.\
\ \\ Thanks to NCBI for making the ClinVar data available on their FTP site as a tab-separated file.\ If you email them (clinvar@ncbi.nlm.nih.gov), feel free to CC us, it is always good to learn more about ClinVar.\
\ \\ Landrum MJ, Lee JM, Benson M, Brown G, Chao C, Chitipiralla S, Gu B, Hart J, Hoffman D, Hoover J\ et al.\ \ ClinVar: public archive of interpretations of clinically relevant variants.\ Nucleic Acids Res. 2016 Jan 4;44(D1):D862-8.\ PMID: 26582918; PMC: PMC4702865\
\ \\ Azzariti DR, Riggs ER, Niehaus A, Rodriguez LL, Ramos EM, Kattman B, Landrum MJ, Martin CL, Rehm HL.\ \ Points to consider for sharing variant-level information from clinical genetic testing with\ ClinVar.\ Cold Spring Harb Mol Case Stud. 2018 Feb;4(1).\ PMID: 29437798; PMC: PMC5793773\
\ \ phenDis 1 compositeTrack on\ dataVersion /gbdb/$D/bbi/clinvar/version.txt\ group phenDis\ itemRgb on\ longLabel ClinVar Variants\ noParentConfig on\ scoreLabel ClinVar Star-Rating (0-4)\ shortLabel ClinVar Variants\ track clinvar\ type bed 12 +\ urls rcvAcc="https://www.ncbi.nlm.nih.gov/clinvar/$$/" geneId="https://www.ncbi.nlm.nih.gov/gene/$$" snpId="https://www.ncbi.nlm.nih.gov/snp/$$" nsvId="https://www.ncbi.nlm.nih.gov/dbvar/variants/$$/" origName="https://www.ncbi.nlm.nih.gov/clinvar/variation/$$/"\ visibility hide\ cloneEndSuper Clone Ends bed 3 Mapping of clone libraries end placements 0 100 0 0 0 127 127 127 0 0 0\ This track shows the NCBI clone end mappings from the\ NCBI Clone DB database. Libraries with more than\ 30,000 clones are included in this track display.\ While the NCBI Clone DB database interface has been retired and is no longer\ available, they were archived and are still accessible for download at NCBI and through the\ UCSC Genome Browser.
\\ Clone availability: most of the clone libraries shown here can\ no longer be ordered. Two librarires that we show are exceptions and are still available\ for ordering from\ BACPAC\ Genomics who still sells the libraries made by and formerly distributed by\ Children's Hospital Oakland Research Institute (CHORI): the \ BCGSC Human 32k BAC Re-Array\ (minimal tiling set, mostly RP11 and CTD clones) and the CHORI-17 (CH17)\ BAC library from a hydatidiform mole.
\\ Bacterial artificial chromosomes (BACs) are a key part of many\ large-scale sequencing projects. A BAC typically consists of 50 - 300 kb of\ DNA. During the early phase of a sequencing project, it is common\ to sequence a single read (approximately 500 bases) off each end of\ a large number of BACs. Later on in the project, these BAC end reads\ can be mapped to the genome sequence.
\\ These BAC end pairs can be useful for validating the assembly over\ relatively long ranges. In some cases, the BACs are useful biological\ reagents. This track can also be used for determining which BAC\ contains a given gene, useful information for certain wet lab experiments.
\\ The scoring scheme used for this annotation assigns 1000 to an alignment\ when the BAC end pair aligns to only one location in the genome (after\ filtering). When a BAC end pair or clone aligns to multiple locations, the\ score is calculated as 1500/(number of alignments).
\ \\ Items in this track are colored according to their strand orientation. Blue indicates alignment to the forward strand, \ and green indicates alignment to the negative strand.\
\ \ \\ The mappings of these BAC end sequences are taken directly from the\ NCBI Clone DB FTP site\ ftp.ncbi.nih.gov/repository/clone/reports/Homo_sapiens/\ *.GCF_000001405.26.106.*.gff files.
\\ UCSC filtered the NCBI Clone DB mapped ends to drop ends that mapped to a\ region that was three times longer than the median size of the clones in\ the library. Only libraries with more than\ 30,000 clones are included in this track display.
\\ Click through on displayed items to the Clone DB database information,\ including\ Clone DB distributor references.
\| clone information from NCBI Clone DB and UCSC mapping statistics | ||||||
|---|---|---|---|---|---|---|
| library name | \
total clones | \
total end sequences | \
NCBI mapped ends | \
UCSC filtered ends | \
UCSC dropped | \
per-cent dropped | \
| ABC8 | 2,007,047 | 3,888,476 | 1,205,466 | 1,192,784 | 12,682 | % 1.05 |
| WI2 | 1,122,564 | 2,298,885 | 589,547 | 582,843 | 6,704 | % 1.14 |
| ABC12 | 1,120,939 | 2,169,280 | 778,216 | 771,827 | 6,389 | 0.82 |
| ABC7 | 1,116,966 | 2,152,975 | 650,329 | 644,071 | 6,258 | 0.96 |
| ABC9 | 1,065,503 | 2,084,892 | 757,644 | 750,648 | 6,996 | 0.92 |
| ABC10 | 1,062,082 | 2,121,489 | 788,344 | 781,331 | 7,013 | 0.89 |
| ABC14 | 1,042,929 | 2,089,193 | 846,055 | 839,126 | 6,929 | 0.82 |
| ABC13 | 1,009,643 | 2,057,345 | 811,829 | 803,589 | 8,240 | 1.01 |
| ABC11 | 998,880 | 1,966,644 | 730,565 | 724,864 | 5,701 | 0.78 |
| ABC23 | 942,133 | 1,535,766 | 437,098 | 433,896 | 3,202 | 0.73 |
| ABC16 | 907,948 | 1,534,288 | 452,316 | 449,101 | 3,215 | 0.71 |
| ABC24 | 835,600 | 1,383,475 | 399,056 | 395,776 | 3,280 | 0.82 |
| ABC27 | 768,336 | 1,229,804 | 334,232 | 331,822 | 2,410 | 0.72 |
| ABC18 | 743,640 | 1,204,811 | 325,150 | 322,904 | 2,246 | 0.69 |
| COR2A | 723,569 | 1,441,881 | 583,327 | 578,578 | 4,749 | 0.81 |
| ABC22 | 519,274 | 780,151 | 189,988 | 188,743 | 1,245 | 0.66 |
| ABC21 | 436,930 | 680,160 | 182,214 | 180,973 | 1,241 | 0.68 |
| RP11 | 292,975 | 394,813 | 86,875 | 85,903 | 972 | 1.12 |
| COR02 | 272,396 | 546,984 | 208,377 | 206,782 | 1,595 | 0.77 |
| CTD | 226,848 | 403,688 | 96,594 | 94,941 | 1,653 | 1.71 |
| CH17 | 176,209 | 325,659 | 105,805 | 105,060 | 745 | 0.70 |
| ABC20 | 49,132 | 80,350 | 24,720 | 24,474 | 246 | 1.00 |
| UCSC dropped | 152,979 | n/a | n/a | n/a | n/a | n/a |
| multiple mappings | 775,629 | n/a | n/a | n/a | n/a | n/a |
\ Many of the libraries shown here were constructed by\ Pieter J. de Jong\ and colleagues, including the RPCI-11 (RP11) library at the Roswell Park\ Cancer Institute, and the CHORI-17 (CH17) and BCGSC 32k Re-Array libraries\ at BACPAC Genomics (formerly at the Children's Hospital Oakland Research\ Institute, CHORI). For background on de Jong's role in building these\ clone libraries, see this\ Undark profile.
\\ Additional information about the clone, including how it\ can be obtained, may be found at the\ NCBI Clone Registry. To view the registry entry for a\ specific clone, open the details page for the clone and click on its name at\ the top of the page.
\ map 1 compositeTrack on\ dimensions dimensionX=source\ dragAndDrop on\ group map\ longLabel Mapping of clone libraries end placements\ noInherit on\ shortLabel Clone Ends\ sortOrder source=+\ subGroup1 source Source agencourt=Agencourt chori=Chori corielle=Coriell caltech=CalTech rpci=RPCI wibr=WIBR placements=Placements\ track cloneEndSuper\ type bed 3\ visibility hide\ clsLongReadRnaTrack CLS long-read RNAs bigBed 12 Capture long-seq long-read lncRNAs 3 100 0 0 0 127 127 127 0 0 0\ These tracks represent the results of targeted long-read RNA sequencing\ aimed at identifying lowly expressed lncRNAs in adult and embryonic\ tissues. The track consists of capture target regions, mappings of pre- and\ post-capture reads, and transcript models built from the data.\
\ \\ Portions of this dataset were used to develop the lncRNA annotations\ introduced in GENCODE v47. The data are a superset of the data incorporated\ into GENCODE. The transcript models for a given RNA do not necessarily match\ those in GENCODE and are provided as a guide to exploring the sequencing data.\
\ \\ Detailed descriptions of the data are available at the\ GENCODE CLS Project site.
\ \\
This is a multi-view composite track containing multiple data types (views). Each view includes subtracks that are displayed individually in the browser. Instructions for configuring multi-view tracks are \
here.
\
\
\
Views:
\
Model Color Coding
\
\ Model annotations are color-coded based on their incorporation into GENCODE V47\ and the assigned GENCODE V47 BioType:\
\\ This project, led by the \ GENCODE consortium,\ employed the Capture Long-read Sequencing (CLS) protocol to enrich transcripts from targeted genomic regions. It used a large capture array with orthologous probes in human and mouse genomes, targeting non-GENCODE lncRNA annotations and regions suspected of unannotated transcription. CapTrap-Seq, a cDNA library preparation protocol, was used to enrich for full-length RNA molecules (5′ to 3′).\
\ \\ Matched adult and embryonic tissues from human and mouse were selected to maximize transcriptome complexity. Libraries were sequenced pre- and post-capture using PacBio and Oxford Nanopore Technologies (ONT) long-read platforms, as well as short-read technologies.\
\ \\ Transcript isoform models were built from reads using the LyRic analysis software. These were merged using intron chains, with transcription start and end sites anchored using CAGE and poly(A) data.\
\ \\ Data and metadata is discoverable via Array Express entry E-MTAB-14562\
\ \\
This dataset was developed by the \
Guigó Lab, Centre for Genomic Regulation (CRG)\
and the GENCODE consortium.
\
The track set was constructed by Sílvia Carbonell-Sala, Andrea Tanzer, and Mark Diekhans.
\ Kaur G, Perteghella T, Carbonell-Sala S, Gonzalez-Martinez J, Hunt T, Mądry T, Jungreis I, Arnan C,\ Lagarde J, Borsari B et al.\ \ GENCODE: massively expanding the lncRNA catalog through capture long-read RNA sequencing.\ bioRxiv. 2024 Oct 31;.\ PMID: 39554180;\ PMC: PMC11565817\
\ \\ Mudge JM, Carbonell-Sala S, Diekhans M, Martinez JG, Hunt T, Jungreis I, Loveland JE, Arnan C,\ Barnes I, Bennett R et al.\ \ GENCODE 2025: reference gene annotation for human and mouse.\ Nucleic Acids Res. 2025 Jan 6;53(D1):D966-D975.\ PMID: 39565199;\ PMC: PMC11701607\
\ \\ Pardo-Palacios FJ, Wang D, Reese F, Diekhans M, Carbonell-Sala S, Williams B, Loveland JE, De María\ M, Adams MS, Balderrama-Gutierrez G et al.\ \ Systematic assessment of long-read RNA-seq methods for transcript identification and\ quantification.\ Nat Methods. 2024 Jul;21(7):1349-1363.\ PMID: 38849569;\ PMC: PMC11543605\
\ \\ Carbonell-Sala S, Perteghella T, Lagarde J, Nishiyori H, Palumbo E, Arnan C, Takahashi H, Carninci\ P, Uszczynska-Ratajczak B, Guigó R.\ \ CapTrap-seq: a platform-agnostic and quantitative approach for high-fidelity full-length RNA\ sequencing.\ Nat Commun. 2024 Jun 27;15(1):5278.\ PMID: 38937428;\ PMC: PMC11211341\
\ \\ LyRic: Long RNA-seq analysis workflow \ https://github.com/guigolab/LyRic\
\ rna 1 compositeTrack on\ dimensions dimX=type dimY=sample\ html clsLongReadRna.html\ longLabel Capture long-seq long-read lncRNAs\ parent long_read_transcripts\ shortLabel CLS long-read RNAs\ subGroup1 view Views targets_view=Targets models_view=Models sample_models_view=Sample_models per_expr_models_view=Per-experiment_models per_expr_reads_view=Per-experiment_reads\ subGroup2 sample Sample combined=Combined adult_brain=Adult_Brain embryo_brain=Embryonic_Brain adult_cpoola=Cell_Line_Pool adult_heart=Adult_Heart embryo_heart=Embryonic_Heart adult_liver=Adult_Liver embryo_liver=Embryonic_Liver placenta_placenta=Placenta adult_testis=Adult_Testis adult_tpoola=Tissue_Pool adult_wblood=Adult_Blood embryo_ipsc=Embryonic_iPSC\ subGroup3 type Type targets=Targets models=Models pre_capture_ont_models=Pre-capture_ONT_models pre_capture_pacbio_models=Pre-capture_PacBio_models post_capture_ont_models=Post-capture_ONT_models post_capture_pacbio_models=Post-capture_PacBio_models pre_capture_ont_reads=Pre-capture_ONT_reads pre_capture_pacbio_reads=Pre-capture_PacBio_reads post_capture_ont_reads=Post-capture_ONT_reads post_capture_pacbio_reads=Post-capture_PacBio_reads\ track clsLongReadRnaTrack\ type bigBed 12\ visibility pack\ ghClusteredInteraction Clustered Interactions bigInteract GeneHancer Regulatory Elements and Gene Interactions 3 100 0 0 0 127 127 127 0 0 0 https://www.genecards.org/cgi-bin/carddisp.pl?gene=$\ This track shows data from Single-cell transcriptome analysis reveals differential\ nutrient absorption functions in human intestine. Droplet-based\ single-cell RNA sequencing (scRNA-seq) was used to survey gene expression\ profiles of the epithelium in the human ileum, colon, and rectum. A total of 7\ cell clusters were identified: enterocytes (EC), goblet cells (G), paneth-like\ cells (PLC), enteroendocrine cells (EEC), progenitor cells (PRO),\ transient-amplifying cells (TA) and stem cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in colon\ cells where cells are grouped by cell type \ (Colon Cells) or donor \ (Colon Donor). The default track \ displayed is Colon Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated \ with those classes. Note that the Colon Donor track \ is colored by donor for improved clarity.
\ \\ Using scRNA-seq, RNA profiles of intestinal epithelial cells were obtained for \ 4,472 cells from two human colon samples. Tissue samples belonged to a male \ donor age 54 (Colon-1) and a female donor age 67 (Colon-2) both diagnosed with \ Adenocarcinoma. The healthy intestinal mucous membranes used for each sample \ were cut away from the tumor border in surgically removed ascending colon tissue. \ Additionally, the intestinal tissues were washed in Hank's balanced salt solution \ (HBSS) to remove mucus, blood cells, and muscle tissue. The sample was enriched \ for epithelial cells through centrifugation before being dissociated with Tryple \ to obtain single-cell suspensions. RNA-seq libraries were prepared using 10x \ Genomics 3' v2 kit and sequenced on an Illumina Hiseq X Ten PE150.
\ \The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used \ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The\ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720
\ \ \ singleCell 1 barChartBars enteroendocrine_cell enterocyte goblet_cell paneth-like_cell progenitor_cell stem_cell transit-amplifying_cell\ barChartColors #c7d2e5 #0198c0 #0251fc #7197d7 #4d689b #9e9fa2 #949dae\ barChartLimit 1.6\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/colonWang/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/colonWang/cell_type.bb\ defaultLabelFields name\ html colonWang\ labelFields name,name2\ longLabel Colon cells binned by cell type from Wang et al 2020\ parent colonWang\ shortLabel Colon Cells\ track colonWangCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-intestine+colon&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ colonWangDonor Colon Donor bigBarChart Colon cells binned by organ donor from Wang et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-intestine+colon&gene=$$\ This track shows data from Single-cell transcriptome analysis reveals differential\ nutrient absorption functions in human intestine. Droplet-based\ single-cell RNA sequencing (scRNA-seq) was used to survey gene expression\ profiles of the epithelium in the human ileum, colon, and rectum. A total of 7\ cell clusters were identified: enterocytes (EC), goblet cells (G), paneth-like\ cells (PLC), enteroendocrine cells (EEC), progenitor cells (PRO),\ transient-amplifying cells (TA) and stem cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in colon\ cells where cells are grouped by cell type \ (Colon Cells) or donor \ (Colon Donor). The default track \ displayed is Colon Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated \ with those classes. Note that the Colon Donor track \ is colored by donor for improved clarity.
\ \\ Using scRNA-seq, RNA profiles of intestinal epithelial cells were obtained for \ 4,472 cells from two human colon samples. Tissue samples belonged to a male \ donor age 54 (Colon-1) and a female donor age 67 (Colon-2) both diagnosed with \ Adenocarcinoma. The healthy intestinal mucous membranes used for each sample \ were cut away from the tumor border in surgically removed ascending colon tissue. \ Additionally, the intestinal tissues were washed in Hank's balanced salt solution \ (HBSS) to remove mucus, blood cells, and muscle tissue. The sample was enriched \ for epithelial cells through centrifugation before being dissociated with Tryple \ to obtain single-cell suspensions. RNA-seq libraries were prepared using 10x \ Genomics 3' v2 kit and sequenced on an Illumina Hiseq X Ten PE150.
\ \The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used \ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The\ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720
\ \ \ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/colonWang/donor.colors\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/colonWang/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/colonWang/donor.bb\ defaultLabelFields name\ html colonWang\ labelFields name,name2\ longLabel Colon cells binned by organ donor from Wang et al 2020\ parent colonWang\ shortLabel Colon Donor\ track colonWangDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-intestine+colon&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ colonWang Colon Wang Colon single cell sequencing from Wang et al 2020 0 100 0 0 0 127 127 127 0 0 0\ This track shows data from Single-cell transcriptome analysis reveals differential\ nutrient absorption functions in human intestine. Droplet-based\ single-cell RNA sequencing (scRNA-seq) was used to survey gene expression\ profiles of the epithelium in the human ileum, colon, and rectum. A total of 7\ cell clusters were identified: enterocytes (EC), goblet cells (G), paneth-like\ cells (PLC), enteroendocrine cells (EEC), progenitor cells (PRO),\ transient-amplifying cells (TA) and stem cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in colon\ cells where cells are grouped by cell type \ (Colon Cells) or donor \ (Colon Donor). The default track \ displayed is Colon Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated \ with those classes. Note that the Colon Donor track \ is colored by donor for improved clarity.
\ \\ Using scRNA-seq, RNA profiles of intestinal epithelial cells were obtained for \ 4,472 cells from two human colon samples. Tissue samples belonged to a male \ donor age 54 (Colon-1) and a female donor age 67 (Colon-2) both diagnosed with \ Adenocarcinoma. The healthy intestinal mucous membranes used for each sample \ were cut away from the tumor border in surgically removed ascending colon tissue. \ Additionally, the intestinal tissues were washed in Hank's balanced salt solution \ (HBSS) to remove mucus, blood cells, and muscle tissue. The sample was enriched \ for epithelial cells through centrifugation before being dissociated with Tryple \ to obtain single-cell suspensions. RNA-seq libraries were prepared using 10x \ Genomics 3' v2 kit and sequenced on an Illumina Hiseq X Ten PE150.
\ \The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used \ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The\ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720
\ \ \ singleCell 0 group singleCell\ longLabel Colon single cell sequencing from Wang et al 2020\ shortLabel Colon Wang\ superTrack on\ track colonWang\ visibility hide\ gnomADPextColon_Sigmoid Colon-Sigmoid bigWig 0 1 gnomAD pext Colon-Sigmoid 0 100 238 187 119 246 221 187 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Colon_Sigmoid.bw\ color 238,187,119\ longLabel gnomAD pext Colon-Sigmoid\ parent gnomadPext off\ shortLabel Colon-Sigmoid\ track gnomADPextColon_Sigmoid\ visibility hide\ gnomADPextColon_Transverse Colon-Transverse bigWig 0 1 gnomAD pext Colon-Transverse 0 100 204 153 85 229 204 170 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Colon_Transverse.bw\ color 204,153,85\ longLabel gnomAD pext Colon-Transverse\ parent gnomadPext off\ shortLabel Colon-Transverse\ track gnomADPextColon_Transverse\ visibility hide\ cons470wayViewelements Conserved Elements bed 4 Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group, 0 100 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group,\ parent cons470way\ shortLabel Conserved Elements\ track cons470wayViewelements\ view elements\ visibility hide\ constraintSuper Constraint scores bed Human constraint scores 0 100 0 0 0 127 127 127 0 0 0\ The "Constraint scores" container track includes several subtracks showing the results of\ constraint prediction algorithms. These try to find regions of negative\ selection, where variations likely have functional impact. The algorithms do\ not use multi-species alignments to derive evolutionary constraint, but use\ primarily human variation, usually from variants collected by gnomAD (see the\ gnomAD V2 or V3 tracks on hg19 and hg38) or TOPMED (contained in our dbSNP\ tracks and available as a filter). One of the subtracks is based on UK Biobank\ variants, which are not available publicly, so we have no track with the raw data.\ The number of human genomes that are used as the input for these scores are\ 76k, 53k and 110k for gnomAD, TOPMED and UK Biobank, respectively.\
\ \Note that another important constraint score, gnomAD\ constraint, is not part of this container track but can be found in the hg38 gnomAD\ track.\
\ \ The algorithms included in this track are:\\ JARVIS scores are shown as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The scores were downloaded and converted to a single bigWig file.\ Move the mouse over the bars to display the exact values. A horizontal line is shown at the 0.733\ value which signifies the 90th percentile.
\ See hg19 makeDoc and\ hg38 makeDoc.\\ Interpretation: The authors offer a suggested guideline of > 0.9998 for identifying\ higher confidence calls and minimizing false positives. In addition to that strict threshold, the \ following two more relaxed cutoffs can be used to explore additional hits. Note that these\ thresholds are offered as guidelines and are not necessarily representative of pathogenicity.
\ \\
| Percentile | JARVIS score threshold |
|---|---|
| 99th | 0.9998 |
| 95th | 0.9826 |
| 90th | 0.7338 |
\ HMC scores are displayed as a signal ("wiggle") track, with one score per genome position.\ Mousing over the bars displays the exact values. The highly-constrained cutoff\ of 0.8 is indicated with a line.
\\ Interpretation: \ A protein residue with HMC score <1 indicates that missense variants affecting\ the homologous residues are significantly under negative selection (P-value <\ 0.05) and likely to be deleterious. A more stringent score threshold of HMC<0.8\ is recommended to prioritize predicted disease-associated variants.\
\ \\ Interpretation: The authors suggest the following guidelines for evaluating\ intolerance. By default, the MetaDome track displays a horizontal line at 0.7 which \ signifies the first intolerant bin. For more information see the MetaDome publication.
\ \\
| Classification | MetaDome Tolerance Score |
|---|---|
| Highly intolerant | ≤ 0.175 |
| Intolerant | ≤ 0.525 |
| Slightly intolerant | ≤ 0.7 |
\ MTR data can be found on two tracks, MTR All data and MTR Scores. In the\ MTR Scores track the data has been converted into 4 separate signal tracks\ representing each base pair mutation, with the lowest possible score shown when\ multiple transcripts overlap at a position. Overlaps can happen since this score\ is derived from transcripts and multiple transcripts can overlap. \ A horizontal line is drawn on the 0.8 score line\ to roughly represent the 25th percentile, meaning the items below may be of particular\ interest. It is recommended that the data be explored using\ this version of the track, as it condenses the information substantially while\ retaining the magnitude of the data.
\ \Any specific point mutations of interest can then be researched in the \ MTR All data track. This track contains all of the information from\ \ MTRV2 including more than 3 possible scores per base when transcripts overlap.\ A mouse-over on this track shows the ref and alt allele, as well as the MTR score\ and the MTR score percentile. Filters are available for MTR score, False Discovery Rate\ (FDR), MTR percentile, and variant consequence. By default, only items in the bottom\ 25 percentile are shown. Items in the track are colored according\ to their MTR percentile:
\\ Interpretation: Regions with low MTR scores were seen to be enriched with\ pathogenic variants. For example, ClinVar pathogenic variants were seen to\ have an average score of 0.77 whereas ClinVar benign variants had an average score\ of 0.92. Further validation using the FATHMM cancer-associated training dataset saw\ that scores less than 0.5 contained 8.6% of the pathogenic variants while only containing\ 0.9% of neutral variants. In summary, lower scores are more likely to represent\ pathogenic variants whereas higher scores could be pathogenic, but have a higher chance\ to be a false positive. For more information see the MTR-Viewer publication.
\ \\ Scores were downloaded and converted to a single bigWig file. See the\ hg19 makeDoc and the\ hg38 makeDoc for more info.\
\ \\ Scores were downloaded and converted to .bedGraph files with a custom Python \ script. The bedGraph files were then converted to bigWig files, as documented in our \ makeDoc hg19 build log.
\ \\
The authors provided a bed file containing codon coordinates along with the scores. \
This file was parsed with a python script to create the two tracks. For the first track\
the scores were aggregated for each coordinate, then the lowest score chosen for any\
overlaps and the result written out to bedGraph format. The file was then converted\
to bigWig with the bedGraphToBigWig utility. For the second track the file\
was reorganized into a bed 4+3 and conveted to bigBed with the bedToBigBed\
utility.
\ See the hg19 makeDoc for details including the build script.
\\ The raw MetaDome data can also be accessed via their Zenodo handle.
\ \\ V2\ file was downloaded and columns were reshuffled as well as itemRgb added for the\ MTR All data track. For the MTR Scores track the file was parsed with a python\ script to pull out the highest possible MTR score for each of the 3 possible mutations\ at each base pair and 4 tracks built out of these values representing each mutation.
\\ See the hg19 makeDoc entry on MTR for more info.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/hmc/hmc.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \ \\ Thanks to Jean-Madeleine Desainteagathe (APHP Paris, France) for suggesting the JARVIS, MTR, HMC tracks. Thanks to Xialei Zhang for providing the HMC data file and to Dimitrios Vitsios and Slave Petrovski for helping clean up the hg38 JARVIS files for providing guidance on interpretation. Additional\ thanks to Laurens van de Wiel for providing the MetaDome data as well as guidance on the track development and interpretation. \
\ \ \\ Vitsios D, Dhindsa RS, Middleton L, Gussow AB, Petrovski S.\ \ Prioritizing non-coding regions based on human genomic constraint and sequence context with deep\ learning.\ Nat Commun. 2021 Mar 8;12(1):1504.\ PMID: 33686085; PMC: PMC7940646\
\ \\ Xiaolei Zhang, Pantazis I. Theotokis, Nicholas Li, the SHaRe Investigators, Caroline F. Wright, Kaitlin E. Samocha, Nicola Whiffin, James S. Ware\ \ Genetic constraint at single amino acid resolution improves missense variant prioritisation and gene discovery.\ Medrxiv 2022.02.16.22271023\
\ \\ Wiel L, Baakman C, Gilissen D, Veltman JA, Vriend G, Gilissen C.\ \ MetaDome: Pathogenicity analysis of genetic variants through aggregation of homologous human protein\ domains.\ Hum Mutat. 2019 Aug;40(8):1030-1038.\ PMID: 31116477; PMC: PMC6772141\
\ \\ Silk M, Petrovski S, Ascher DB.\ \ MTR-Viewer: identifying regions within genes under purifying selection.\ Nucleic Acids Res. 2019 Jul 2;47(W1):W121-W126.\ PMID: 31170280; PMC: PMC6602522\
\ \\ Halldorsson BV, Eggertsson HP, Moore KHS, Hauswedell H, Eiriksson O, Ulfarsson MO, Palsson G,\ Hardarson MT, Oddsson A, Jensson BO et al.\ \ The sequences of 150,119 genomes in the UK Biobank.\ Nature. 2022 Jul;607(7920):732-740.\ PMID: 35859178; PMC: PMC9329122\
\ \ \\ Huang YF, Gulko B, Siepel A.\ \ Fast, scalable prediction of deleterious noncoding variants from functional and population genomic\ data.\ Nat Genet. 2017 Apr;49(4):618-624.\ PMID: 28288115; PMC: PMC5395419\
\ \ phenDis 1 group phenDis\ longLabel Human constraint scores\ shortLabel Constraint scores\ superTrack on hide\ track constraintSuper\ type bed\ visibility hide\ constraintV2 Constraint V2 bigBed 12 + gnomAD Constraint Metrics V2 3 100 0 0 0 127 127 127 0 0 0 varRep 1 longLabel gnomAD Constraint Metrics V2\ parent gnomadPLI off\ shortLabel Constraint V2\ track constraintV2\ type bigBed 12 +\ view v2\ visibility pack\ constraintV4 Constraint V4 bigBed 12 + gnomAD Constraint Metrics V4 0 100 0 0 0 127 127 127 0 0 0 varRep 1 longLabel gnomAD Constraint Metrics V4\ parent gnomadPLI off\ shortLabel Constraint V4\ track constraintV4\ type bigBed 12 +\ view v4\ visibility hide\ constraintV4_1 Constraint V4.1 bigBed 12 + gnomAD Constraint Metrics V4.1 3 100 0 0 0 127 127 127 0 0 0 varRep 1 longLabel gnomAD Constraint Metrics V4.1\ parent gnomadPLI off\ shortLabel Constraint V4.1\ track constraintV4_1\ type bigBed 12 +\ view v4_1\ visibility pack\ coriellDelDup Coriell CNVs bed 9 + Coriell Cell Line Copy Number Variants 0 100 0 0 0 127 127 127 0 0 0 http://ccr.coriell.org/Sections/Search/Search.aspx?q=$$\ The Coriell Cell Line Copy Number Variants track displays\ copy-number variants (CNVs) in chromosomal aberration and inherited disorder\ cell lines in the NIGMS Human Genetic Cell Repository. The Repository,\ sponsored by the National Institute of General Medical Sciences, provides\ scientists around the world with resources for cell and genetic research.\ The samples include highly characterized cell lines and high quality DNA.\ NIGMS Repository samples represent a variety of disease states, chromosomal\ abnormalities, apparently healthy individuals and many distinct human\ populations.\
\ \\ Approximately 1000 samples from the Chromosomal Aberrations and Heritable\ Diseases collections of the NIGMS Repository were genotyped on the Affymetrix\ Genome-Wide Human SNP 6.0 Array and analyzed for CNVs at the Coriell Institute\ for Medical Research. Genotyping data for many of these samples is available\ through dbGaP.\
\ \\ The genotyped samples represent a diverse set of copy-number variants. The\ selection was weighted to over-sample commonly manifested types of aberrations.\ Karyotyping was performed on all NIGMS Repository cell lines that were\ submitted with reported chromosome abnormalities. When available, the ISCN\ description of the sample, based on G-banding and FISH analysis, is included\ in the phenotypic data. Karyotypes for these cells can be viewed in the\ online Repository catalog.\
\ \\ Field definitions for an item description:\
\ CN State item coloring:\
\ We thank Dorit Berlin and Zhenya Tang of the NIGMS Human Genetic Cell\ Repository at the\ Coriell Institute for Medical\ Research for these data.\
\ \\
NCBI dbGaP:\
\
Genotyping NIGMS Chromosomal Aberration and Inherited Disorder Samples.\
\
\
NIGMS Human Genetic Cell Repository\
online catalog at the Coriell Institute for Medical Research.\
\
\ This track displays data from Single-cell genomics identifies cell type-specific\ molecular changes in autism. Single-nucleus RNA sequencing (snRNA-seq)\ was performed on post-mortem cortical tissue samples from patients with autism\ spectrum disorder (ASD) as well as control donors. A total of 17 cell clusters\ were identified using known cell type markers found in Velmeshev et\ al., 2019.
\ \\ This track collection contains five bar chart tracks of RNA expression in the human\ cerebral cortex where cells are grouped by cell type \ (Cortex Cells), diagnosis\ (Cortex Diagnosis), donor \ (Cortex Donor), sample \ (Cortex Sample), and sex\ (Cortex Sex). \ The default track displayed is Cortex Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| immune | |
| endothelial | |
| glia |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Cortex Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy cortical samples were taken from 16 controls (ages 4-22) without \ neurological disorders and 15 ASD patients (ages 7-21). A total of 41 post-mortem\ tissue samples were obtained from both the prefrontal cortex (PFC) and anterior\ cingulate cortex (ACC). When present, subcortical white matter was removed\ prior to collection from cortical samples containing all layers of cortical\ grey matter. ASD and control samples were matched for sex and age and processed\ together to minimize batch effects. Nuclei were isolated from brain tissue\ using a glass dounce homogenizer in lysis buffer and then filtered twice\ through a 30 µm cell strainer. Next, samples were processed\ using 10x Genomics 3' library kit and the resulting single-nucleus libraries\ were pooled together and sequenced on an Illumina NovaSeq 6000. This process\ generated 104,559 single-nuclei gene expression profiles in total.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The\ UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Dmitry Velmeshev and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by by Daniel Schmelter. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Velmeshev D, Schirmer L, Jung D, Haeussler M, Perez Y, Mayer S, Bhaduri A, Goyal N, Rowitch DH,\ Kriegstein AR.\ \ Single-cell genomics identifies cell type-specific molecular changes in autism.\ Science. 2019 May 17;364(6441):685-689.\ PMID: 31097668; PMC: PMC7678724\
\ singleCell 1 barChartBars astrocyte_(fibrous) astrocyte_(protoplasmic) endothelial_cell interneuron_PVALB+ interneuron_SST+ interneuron_SV2C+ interneuron_VIP+ neuron_L2/3_cortex neuron_L4_cortex neuron_L5/6_corticofugal neuron_L5/6_cortico-cortical microglial_cell neuron_NRGN+_I neuron_NRGN+_II neuron_maturing oligodendrocyte_precursor oligodendrocyte\ barChartColors #81ce00 #81cd00 #01c000 #ebbf00 #ebbf00 #eabe00 #ebbf00 #ecbf00 #ecbf00 #ecbf00 #edbf00 #ef1211 #c8b701 #c5b701 #ebbf00 #c5be01 #86c601\ barChartLimit 4\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/cortexVelmeshev/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/cortexVelmeshev/cell_type.bb\ defaultLabelFields name2\ html cortexVelmeshev\ labelFields name,name2\ longLabel Cerebral cortex RNA binned by cell type from Velmeshev et al 2019\ parent cortexVelmeshev\ shortLabel Cortex Cells\ track cortexVelmeshevCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=autism&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ cortexVelmeshevDiagnosis Cortex Diagnosis bigBarChart Cerebral cortex RNA binned by ASD/control diagnosis from Velmeshev et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=autism&gene=$$\ This track displays data from Single-cell genomics identifies cell type-specific\ molecular changes in autism. Single-nucleus RNA sequencing (snRNA-seq)\ was performed on post-mortem cortical tissue samples from patients with autism\ spectrum disorder (ASD) as well as control donors. A total of 17 cell clusters\ were identified using known cell type markers found in Velmeshev et\ al., 2019.
\ \\ This track collection contains five bar chart tracks of RNA expression in the human\ cerebral cortex where cells are grouped by cell type \ (Cortex Cells), diagnosis\ (Cortex Diagnosis), donor \ (Cortex Donor), sample \ (Cortex Sample), and sex\ (Cortex Sex). \ The default track displayed is Cortex Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| immune | |
| endothelial | |
| glia |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Cortex Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy cortical samples were taken from 16 controls (ages 4-22) without \ neurological disorders and 15 ASD patients (ages 7-21). A total of 41 post-mortem\ tissue samples were obtained from both the prefrontal cortex (PFC) and anterior\ cingulate cortex (ACC). When present, subcortical white matter was removed\ prior to collection from cortical samples containing all layers of cortical\ grey matter. ASD and control samples were matched for sex and age and processed\ together to minimize batch effects. Nuclei were isolated from brain tissue\ using a glass dounce homogenizer in lysis buffer and then filtered twice\ through a 30 µm cell strainer. Next, samples were processed\ using 10x Genomics 3' library kit and the resulting single-nucleus libraries\ were pooled together and sequenced on an Illumina NovaSeq 6000. This process\ generated 104,559 single-nuclei gene expression profiles in total.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The\ UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Dmitry Velmeshev and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by by Daniel Schmelter. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Velmeshev D, Schirmer L, Jung D, Haeussler M, Perez Y, Mayer S, Bhaduri A, Goyal N, Rowitch DH,\ Kriegstein AR.\ \ Single-cell genomics identifies cell type-specific molecular changes in autism.\ Science. 2019 May 17;364(6441):685-689.\ PMID: 31097668; PMC: PMC7678724\
\ singleCell 1 barChartBars ASD Control\ barChartColors #ebbf00 #e9bf00\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/cortexVelmeshev/diagnosis.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/cortexVelmeshev/diagnosis.bb\ defaultLabelFields name2\ html cortexVelmeshev\ labelFields name,name2\ longLabel Cerebral cortex RNA binned by ASD/control diagnosis from Velmeshev et al 2019\ parent cortexVelmeshev\ shortLabel Cortex Diagnosis\ track cortexVelmeshevDiagnosis\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=autism&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ cortexVelmeshevDonor Cortex Donor bigBarChart Cerebral cortex RNA binned by organ donor from Velmeshev et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=autism&gene=$$\ This track displays data from Single-cell genomics identifies cell type-specific\ molecular changes in autism. Single-nucleus RNA sequencing (snRNA-seq)\ was performed on post-mortem cortical tissue samples from patients with autism\ spectrum disorder (ASD) as well as control donors. A total of 17 cell clusters\ were identified using known cell type markers found in Velmeshev et\ al., 2019.
\ \\ This track collection contains five bar chart tracks of RNA expression in the human\ cerebral cortex where cells are grouped by cell type \ (Cortex Cells), diagnosis\ (Cortex Diagnosis), donor \ (Cortex Donor), sample \ (Cortex Sample), and sex\ (Cortex Sex). \ The default track displayed is Cortex Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| immune | |
| endothelial | |
| glia |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Cortex Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy cortical samples were taken from 16 controls (ages 4-22) without \ neurological disorders and 15 ASD patients (ages 7-21). A total of 41 post-mortem\ tissue samples were obtained from both the prefrontal cortex (PFC) and anterior\ cingulate cortex (ACC). When present, subcortical white matter was removed\ prior to collection from cortical samples containing all layers of cortical\ grey matter. ASD and control samples were matched for sex and age and processed\ together to minimize batch effects. Nuclei were isolated from brain tissue\ using a glass dounce homogenizer in lysis buffer and then filtered twice\ through a 30 µm cell strainer. Next, samples were processed\ using 10x Genomics 3' library kit and the resulting single-nucleus libraries\ were pooled together and sequenced on an Illumina NovaSeq 6000. This process\ generated 104,559 single-nuclei gene expression profiles in total.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The\ UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Dmitry Velmeshev and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by by Daniel Schmelter. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Velmeshev D, Schirmer L, Jung D, Haeussler M, Perez Y, Mayer S, Bhaduri A, Goyal N, Rowitch DH,\ Kriegstein AR.\ \ Single-cell genomics identifies cell type-specific molecular changes in autism.\ Science. 2019 May 17;364(6441):685-689.\ PMID: 31097668; PMC: PMC7678724\
\ singleCell 1 barChartBars 1823 4341 4849 4899 5144 5163 5242 5278 5294 5387 5391 5403 5408 5419 5531 5538 5554 5565 5577 5841 5864 5879 5893 5936 5939 5945 5958 5976 5978 6032 6033\ barChartColors #e5be00 #e7bf00 #e8bf00 #e9bf00 #c6c200 #ecbf00 #bec100 #e9bf00 #e8bf00 #ebbf00 #e9bf00 #adc600 #e8be00 #e2be00 #e8bf00 #c1c200 #dfbd00 #ebbf00 #e4bf00 #e9bf00 #ecbf00 #e3be00 #e5be00 #d9bf00 #ebbf00 #e3bf00 #eabf00 #ebbf00 #eabf00 #e9bf00 #e8bf00\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/cortexVelmeshev/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/cortexVelmeshev/donor.bb\ defaultLabelFields name2\ html cortexVelmeshev\ labelFields name,name2\ longLabel Cerebral cortex RNA binned by organ donor from Velmeshev et al 2019\ parent cortexVelmeshev\ shortLabel Cortex Donor\ track cortexVelmeshevDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=autism&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ cortexVelmeshevSample Cortex Sample bigBarChart Cerebral cortex RNA binned by biosample from Velmeshev et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=autism&gene=$$\ This track displays data from Single-cell genomics identifies cell type-specific\ molecular changes in autism. Single-nucleus RNA sequencing (snRNA-seq)\ was performed on post-mortem cortical tissue samples from patients with autism\ spectrum disorder (ASD) as well as control donors. A total of 17 cell clusters\ were identified using known cell type markers found in Velmeshev et\ al., 2019.
\ \\ This track collection contains five bar chart tracks of RNA expression in the human\ cerebral cortex where cells are grouped by cell type \ (Cortex Cells), diagnosis\ (Cortex Diagnosis), donor \ (Cortex Donor), sample \ (Cortex Sample), and sex\ (Cortex Sex). \ The default track displayed is Cortex Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| immune | |
| endothelial | |
| glia |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Cortex Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy cortical samples were taken from 16 controls (ages 4-22) without \ neurological disorders and 15 ASD patients (ages 7-21). A total of 41 post-mortem\ tissue samples were obtained from both the prefrontal cortex (PFC) and anterior\ cingulate cortex (ACC). When present, subcortical white matter was removed\ prior to collection from cortical samples containing all layers of cortical\ grey matter. ASD and control samples were matched for sex and age and processed\ together to minimize batch effects. Nuclei were isolated from brain tissue\ using a glass dounce homogenizer in lysis buffer and then filtered twice\ through a 30 µm cell strainer. Next, samples were processed\ using 10x Genomics 3' library kit and the resulting single-nucleus libraries\ were pooled together and sequenced on an Illumina NovaSeq 6000. This process\ generated 104,559 single-nuclei gene expression profiles in total.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The\ UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Dmitry Velmeshev and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by by Daniel Schmelter. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Velmeshev D, Schirmer L, Jung D, Haeussler M, Perez Y, Mayer S, Bhaduri A, Goyal N, Rowitch DH,\ Kriegstein AR.\ \ Single-cell genomics identifies cell type-specific molecular changes in autism.\ Science. 2019 May 17;364(6441):685-689.\ PMID: 31097668; PMC: PMC7678724\
\ singleCell 1 barChartBars 1823_BA24 4341_BA24 4341_BA46 4849_BA24 4899_BA24 5144_PFC 5163_BA24 5242_BA24 5278_BA24 5278_PFC 5294_BA24 5294_BA9 5387_BA9 5391_BA24 5403_PFC 5408_PFC_Nova 5419_PFC 5531_BA24 5531_BA9 5538_PFC_Nova 5554_BA24 5565_BA24 5565_BA9 5577_BA9 5841_BA9 5864_BA9 5879_PFC_Nova 5893_BA24 5893_PFC 5936_PFC_Nova 5939_BA24 5939_BA9 5945_PFC 5958_BA24 5958_BA9 5976_BA9 5978_BA24 5978_BA9 6032_BA24 6033_BA24 6033_BA9\ barChartColors #e5be00 #e7bf00 #e6bf00 #e8bf00 #e9bf00 #c6c200 #ecbf00 #bec100 #e7bf00 #e6be00 #e4be00 #e9bf00 #ebbf00 #e9bf00 #adc600 #e8be00 #e2be00 #e3be00 #e9bf00 #c1c200 #dfbd00 #ebbf00 #ebbf00 #e4bf00 #e9bf00 #ecbf00 #e3be00 #ebbf00 #cbc000 #d9bf00 #e8bf00 #ecbf00 #e3bf00 #e6bf00 #ecbf00 #ebbf00 #eabf00 #e9bf00 #e9bf00 #e6bf00 #e9be00\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/cortexVelmeshev/sample.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/cortexVelmeshev/sample.bb\ defaultLabelFields name2\ html cortexVelmeshev\ labelFields name,name2\ longLabel Cerebral cortex RNA binned by biosample from Velmeshev et al 2019\ parent cortexVelmeshev\ shortLabel Cortex Sample\ track cortexVelmeshevSample\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=autism&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ cortexVelmeshevSex Cortex Sex bigBarChart Cerebral cortex RNA binned by sex of donor from Velmeshev et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=autism&gene=$$\ This track displays data from Single-cell genomics identifies cell type-specific\ molecular changes in autism. Single-nucleus RNA sequencing (snRNA-seq)\ was performed on post-mortem cortical tissue samples from patients with autism\ spectrum disorder (ASD) as well as control donors. A total of 17 cell clusters\ were identified using known cell type markers found in Velmeshev et\ al., 2019.
\ \\ This track collection contains five bar chart tracks of RNA expression in the human\ cerebral cortex where cells are grouped by cell type \ (Cortex Cells), diagnosis\ (Cortex Diagnosis), donor \ (Cortex Donor), sample \ (Cortex Sample), and sex\ (Cortex Sex). \ The default track displayed is Cortex Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| immune | |
| endothelial | |
| glia |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Cortex Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy cortical samples were taken from 16 controls (ages 4-22) without \ neurological disorders and 15 ASD patients (ages 7-21). A total of 41 post-mortem\ tissue samples were obtained from both the prefrontal cortex (PFC) and anterior\ cingulate cortex (ACC). When present, subcortical white matter was removed\ prior to collection from cortical samples containing all layers of cortical\ grey matter. ASD and control samples were matched for sex and age and processed\ together to minimize batch effects. Nuclei were isolated from brain tissue\ using a glass dounce homogenizer in lysis buffer and then filtered twice\ through a 30 µm cell strainer. Next, samples were processed\ using 10x Genomics 3' library kit and the resulting single-nucleus libraries\ were pooled together and sequenced on an Illumina NovaSeq 6000. This process\ generated 104,559 single-nuclei gene expression profiles in total.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The\ UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Dmitry Velmeshev and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by by Daniel Schmelter. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Velmeshev D, Schirmer L, Jung D, Haeussler M, Perez Y, Mayer S, Bhaduri A, Goyal N, Rowitch DH,\ Kriegstein AR.\ \ Single-cell genomics identifies cell type-specific molecular changes in autism.\ Science. 2019 May 17;364(6441):685-689.\ PMID: 31097668; PMC: PMC7678724\
\ singleCell 1 barChartBars F M\ barChartColors #e8bf00 #ebbf00\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/cortexVelmeshev/sex.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/cortexVelmeshev/sex.bb\ defaultLabelFields name2\ html cortexVelmeshev\ labelFields name,name2\ longLabel Cerebral cortex RNA binned by sex of donor from Velmeshev et al 2019\ parent cortexVelmeshev\ shortLabel Cortex Sex\ track cortexVelmeshevSex\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=autism&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ cortexVelmeshev Cortex Velmeshev Cerebral cortex single cell data from Velmeshev et al 2019 0 100 0 0 0 127 127 127 0 0 0\ This track displays data from Single-cell genomics identifies cell type-specific\ molecular changes in autism. Single-nucleus RNA sequencing (snRNA-seq)\ was performed on post-mortem cortical tissue samples from patients with autism\ spectrum disorder (ASD) as well as control donors. A total of 17 cell clusters\ were identified using known cell type markers found in Velmeshev et\ al., 2019.
\ \\ This track collection contains five bar chart tracks of RNA expression in the human\ cerebral cortex where cells are grouped by cell type \ (Cortex Cells), diagnosis\ (Cortex Diagnosis), donor \ (Cortex Donor), sample \ (Cortex Sample), and sex\ (Cortex Sex). \ The default track displayed is Cortex Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| immune | |
| endothelial | |
| glia |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Cortex Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy cortical samples were taken from 16 controls (ages 4-22) without \ neurological disorders and 15 ASD patients (ages 7-21). A total of 41 post-mortem\ tissue samples were obtained from both the prefrontal cortex (PFC) and anterior\ cingulate cortex (ACC). When present, subcortical white matter was removed\ prior to collection from cortical samples containing all layers of cortical\ grey matter. ASD and control samples were matched for sex and age and processed\ together to minimize batch effects. Nuclei were isolated from brain tissue\ using a glass dounce homogenizer in lysis buffer and then filtered twice\ through a 30 µm cell strainer. Next, samples were processed\ using 10x Genomics 3' library kit and the resulting single-nucleus libraries\ were pooled together and sequenced on an Illumina NovaSeq 6000. This process\ generated 104,559 single-nuclei gene expression profiles in total.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The\ UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Dmitry Velmeshev and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by by Daniel Schmelter. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Velmeshev D, Schirmer L, Jung D, Haeussler M, Perez Y, Mayer S, Bhaduri A, Goyal N, Rowitch DH,\ Kriegstein AR.\ \ Single-cell genomics identifies cell type-specific molecular changes in autism.\ Science. 2019 May 17;364(6441):685-689.\ PMID: 31097668; PMC: PMC7678724\
\ singleCell 0 group singleCell\ longLabel Cerebral cortex single cell data from Velmeshev et al 2019\ shortLabel Cortex Velmeshev\ superTrack on\ track cortexVelmeshev\ visibility hide\ cosmicMuts COSMIC bigBed 6 + 3 Catalogue of Somatic Mutations in Cancer V101 0 100 0 0 0 127 127 127 0 0 0 https://cancer.sanger.ac.uk/cosmic/search?q=$$COSMIC, \ the "Catalogue Of Somatic Mutations In Cancer," is an online database of somatic mutations found in \ human cancer. Focused exclusively on non-inherited acquired mutations, COSMIC combines information \ from a range of sources, curating the described relationships between cancer phenotypes and gene \ (and genomic) mutations. These data are then made available in a number of ways including here in the \ UCSC genome browser, on the COSMIC website with custom analytical tools, or via the\ COSMIC sftp server.\ Publications using COSMIC as a data source may cite our reference below.
\ \The data in COSMIC are curated from a number of high-quality sources and combined into a single\ resource. The sources include:
\ \Information on known cancer genes, selected from the \ Cancer Gene Census is curated manually to maximize its descriptive content. \ \
\ UCSC was provided with the COSMIC annotations directly, and the file was converted to a bigBed\ for display using the bedToBigBed utility.\
\ \\ Clicking into any item also displays the reference allele, alternate allele, and the\ Cosmic legacy mutation identifier (COSNnnnnn). Outlinks can also be found directly to COSMIC\ for additional information.\
\ \\ The limited data available to UCSC can be explored interactively \ with the Table Browser,\ or the Data Integrator. For automated analysis, the data may be\ queried from our REST API. Please refer to our\ mailing list archives\ for questions, or our Data Access FAQ for more\ information.
\\ The complete data can be explored and downloaded via the COSMIC \ website.\
\ \For further information on COSMIC, or for help with the information provided, please contact\ \ cosmic@sanger.\ ac.\ uk.\
\ \\ Forbes SA, Beare D, Boutselakis H, Bamford S, Bindal N, Tate J, Cole CG, Ward S, Dawson E, Ponting L\ et al.\ \ COSMIC: somatic cancer genetics at high-resolution.\ Nucleic Acids Res. 2017 Jan 4;45(D1):D777-D783.\ PMID: 27899578; PMC: PMC5210583\
\ phenDis 1 bigDataUrl /gbdb/hg38/cosmic/cosmic.bb\ dataVersion COSMIC v101\ group phenDis\ longLabel Catalogue of Somatic Mutations in Cancer V101\ noScoreFilter on\ shortLabel COSMIC\ track cosmicMuts\ type bigBed 6 + 3\ url https://cancer.sanger.ac.uk/cosmic/search?q=$$\ urlLabel Genomic Mutation ID:\ cosmicRegions COSMIC Regions bigBed 8 + Catalogue of Somatic Mutations in Cancer V82 0 100 200 0 0 227 127 127 0 0 0 http://cancer.sanger.ac.uk/cosmic/mutation/overview?id=$$COSMIC, \ the "Catalogue Of Somatic Mutations In Cancer," is an online database of somatic mutations found in \ human cancer. Focused exclusively on non-inherited acquired mutations, COSMIC combines information \ from a range of sources, curating the described relationships between cancer phenotypes and gene \ (and genomic) mutations. These data are then made available in a number of ways including here in the \ UCSC genome browser, on the COSMIC website with custom analytical tools, or via the\ COSMIC sftp server.\ Publications using COSMIC as a data source may cite our reference below.
\ \The data in COSMIC are curated from a number of high-quality sources and combined into a single\ resource. The sources include:
\ \Information on known cancer genes, selected from the \ Cancer Gene Census is curated manually to maximize its descriptive content. \ \
\ The data was downloaded from the COSMIC sftp server. It was first converted to a bed file using\ the UCSC utility cosmicToBed, then converted into a bigBed file using the UCSC utility bedToBigBed.\ The bigBed file is used to generate the track. \
\ \\
Due to licensed material, we do not allow downloads or Table Browser access for the bigBed data. The\
raw data underlying this track can be explored and downloaded via the COSMIC \
website. The\
CosmicMutantExport.tsv.gz file was converted to a BED file using the cosmicToBed\
utility, and then converted into a bigBed file using the bedToBigBed utility. You can\
download these tools from the\
utilities directory.\
For further information on COSMIC, or for help with the information provided, please contact\ \ cosmic@sanger.\ ac.\ uk.\
\ \\ Forbes SA, Beare D, Boutselakis H, Bamford S, Bindal N, Tate J, Cole CG, Ward S, Dawson E, Ponting L\ et al.\ \ COSMIC: somatic cancer genetics at high-resolution.\ Nucleic Acids Res. 2017 Jan 4;45(D1):D777-D783.\ PMID: 27899578; PMC: PMC5210583\
\ phenDis 1 bigDataUrl /gbdb/hg38/cosmic/cosMutHg38V82.bb\ color 200, 0, 0\ group phenDis\ html cosmicRegions\ labelFields cosmLabel\ longLabel Catalogue of Somatic Mutations in Cancer V82\ mouseOverField _mouseOver\ noScoreFilter on\ pennantIcon snowflake.png ../goldenPath/newsarch.html#091523 "COSMIC data is now updated on the COSMIC track (not COSMIC Regions). See news archive for details."\ searchIndex name,cosmLabel\ shortLabel COSMIC Regions\ tableBrowser off\ track cosmicRegions\ type bigBed 8 +\ url http://cancer.sanger.ac.uk/cosmic/mutation/overview?id=$$\ urlLabel COSMIC ID:\ iscaViewTotal Coverage (Graphical) bedGraph 4 Clinical Genome Resource (ClinGen) CNVs 2 100 0 0 0 127 127 127 0 0 0 phenDis 0 alwaysZero on\ longLabel Clinical Genome Resource (ClinGen) CNVs\ maxHeightPixels 128:57:16\ parent iscaComposite\ shortLabel Coverage (Graphical)\ track iscaViewTotal\ type bedGraph 4\ view cov\ viewLimits 0:100\ viewUi on\ visibility full\ covid COVID Data Container of SARS-CoV-2 data 0 100 0 0 0 127 127 127 0 0 0\ This is a container track for all data related to SARS-CoV-2 for hg38 \ in the UCSC Genome Browser. Click into any of the sub-tracks to see information\ details on the specific annotations.
\ phenDis 0 cartVersion 4\ group phenDis\ longLabel Container of SARS-CoV-2 data\ shortLabel COVID Data\ superTrack on\ track covid\ cpc1Sv CPC 58 SVs bigBed 9 + Structural Variants from the Chinese Pangenome Consortium (58 samples, CPC-only) 0 100 0 0 0 127 127 127 0 0 0\ This track displays structural variants (SVs) at least 50 bp long\ (deletions, insertions, and complex substitutions) identified by the\ Chinese Pangenome Consortium (CPC) in 58 samples representing 36 Chinese\ minority ethnic groups.
\ \\ The upstream release combined the 58 CPC samples with 47 samples from\ Phase 1 of the Human Pangenome Reference Consortium (HPRC) into a single\ pangenome graph built on the T2T-CHM13v2 assembly with Minigraph-Cactus.\ For this track we recomputed allele counts (AC), allele numbers (AN) and\ sample counts (NS) using only the 58 CPC sample columns (those with\ HIFI032* or RY* prefixes in the source VCF) and dropped\ all snarls that no CPC sample carries (HPRC-specific SVs). To see the\ HPRC data on its own, use the HPRC SV tracks elsewhere in this collection.
\ \\ A pangenome is a graph that represents many genomes simultaneously, letting\ variants that are missing from a single linear reference be captured and\ typed directly. Variants are shown natively on the hs1 browser and lifted\ to hg38 using the UCSC hs1ToHg38.over.chain.gz chain. The track\ contains 46,092 snarl sites on hs1 and 36,030 lifted to hg38 (10,062 did\ not lift, typically in T2T-added repetitive regions).
\ \Items are colored by SV type:
\\ Each bed item spans from the start of the REF allele to its end on the\ reference. Pure insertions (where REF is a single base) therefore appear\ as narrow single-base marks; DELs and CPX items span the affected reference\ interval.
\ \\ The name field is the graph snarl ID (two node identifiers separated\ by strand arrows, e.g. >2541>2547). It is stable across the\ graph but has no meaning outside the CPC pangenome graph file.
\ \\ The source VCF was decomposed with bcftools norm -m -any, so each\ graph snarl appears as one VCF row per alternative allele (a single\ bubble in the graph may have 2-20+ alt paths). For this track we first\ compute the CPC-only allele count per alt, drop any alt that no CPC sample\ carries, then collapse all remaining alts sharing the same snarl ID into\ one track item:
\Available filters:
\\ Gao et al. 2023 generated PacBio HiFi long reads (mean ~30.65x,\ Sequel II/IIe platforms) for 58 QC-passed samples representing 36\ minority Chinese ethnic groups, complemented with Illumina short reads\ and Oxford Nanopore ultralong reads. Haplotype-phased de novo assemblies\ were produced with\ hifiasm\ v0.16.1 (116 high-quality haplotype assemblies retained after QC) and\ combined with 47 HPRC Phase 1 assemblies into a single variation graph\ built on T2T-CHM13v2 with the Minigraph-Cactus pipeline (Minigraph v0.19\ for the SV skeleton, Cactus v2.1.1 base alignment, hal2vg).\ Graph bubbles were decomposed into variant records with vcfwave\ and normalized with bcftools norm -m -any, yielding the source\ VCF (CPC.HPRC.Phase1.processed.SVs.normed.vcf.gz). The upstream\ Gao et al. release identified 78,072 SVs across the combined 105-sample\ graph. For this track we restrict to the 58 CPC samples (columns matching\ HIFI032* or RY*), recompute AC/AN/NS from those columns\ only, drop snarls with no CPC carrier (HPRC-specific sites), filter to\ alts with ≥50 bp REF/ALT length difference, and collapse by graph snarl\ ID. The final track contains 46,092 snarl sites on hs1; the hg38 version\ is lifted with the UCSC hs1ToHg38.over.chain.gz chain (36,030\ sites, 10,062 did not lift).
\ \\ The source VCF is distributed by the\ \ Chinese-Pangenome-Consortium-Phase-I GitHub repository.
\ \\ The step-by-step build commands (CPC-only recount, liftOver, snarl\ collapse, bigBed build) are recorded in the UCSC makeDoc for this track\ container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \The data can be explored interactively with the\ Table Browser or\ Data Integrator, and accessed from\ scripts via our API\ (track=cpc1Sv).
\ \For automated download, the bigBed files are at\ \ http://hgdownload.soe.ucsc.edu/gbdb/hs1/lrSv/cpc1.bb (native) and\ \ http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/cpc1.bb (lifted).\ Use bigBedToBed to extract features: e.g.\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hs1/lrSv/cpc1.bb -chrom=chr21 -start=0 -end=100000000 stdout
\ \The original pangenome VCF is distributed by the Chinese Pangenome\ Consortium; see the\ \ CPC Phase I repository.
\ \Thanks to the Chinese Pangenome Consortium and the HPRC Phase 1 team\ for producing and releasing the combined pangenome and its decomposed\ variant calls.
\ \\ Gao Y, Yang X, Chen H, Tan X, Yang Z, Deng L, Wang B, Kong S, Li S, Cui Y et al.\ \ A pangenome reference of 36 Chinese populations.\ Nature. 2023 Jul;619(7968):112-121.\ PMID: 37316654; PMC: PMC10322713\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/cpc1.bb\ filter.AC 0:116\ filter.insLen 0:376583\ filter.svLen 0:8998096\ filterByRange.AC on\ filterByRange.alleleFreq on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Allele Count\ filterLabel.alleleFreq Allele Frequency\ filterLabel.insLen Insertion Length\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.alleleFreq 0:1\ filterType.svType multipleListOr\ filterValues.svType INS,DEL,CPX,MIXED\ itemRgb on\ longLabel Structural Variants from the Chinese Pangenome Consortium (58 samples, CPC-only)\ mouseOver Var: $name ($svType)CpG islands are associated with genes, particularly housekeeping\ genes, in vertebrates. CpG islands are typically common near\ transcription start sites and may be associated with promoter\ regions. Normally a C (cytosine) base followed immediately by a \ G (guanine) base (a CpG) is rare in\ vertebrate DNA because the Cs in such an arrangement tend to be\ methylated. This methylation helps distinguish the newly synthesized\ DNA strand from the parent strand, which aids in the final stages of\ DNA proofreading after duplication. However, over evolutionary time,\ methylated Cs tend to turn into Ts because of spontaneous\ deamination. The result is that CpGs are relatively rare unless\ there is selective pressure to keep them or a region is not methylated\ for some other reason, perhaps having to do with the regulation of gene\ expression. CpG islands are regions where CpGs are present at\ significantly higher levels than is typical for the genome as a whole.
\ \\ The unmasked version of the track displays potential CpG islands\ that exist in repeat regions and would otherwise not be visible\ in the repeat masked version.\
\ \\ By default, only the masked version of the track is displayed. To view the\ unmasked version, change the visibility settings in the track controls at\ the top of this page.\
\ \CpG islands were predicted by searching the sequence one base at a\ time, scoring each dinucleotide (+17 for CG and -1 for others) and\ identifying maximally scoring segments. Each segment was then\ evaluated for the following criteria:\ \
\ The entire genome sequence, masking areas included, was\ used for the construction of the track Unmasked CpG.\ The track CpG Islands is constructed on the sequence after\ all masked sequence is removed.\
\ \The CpG count is the number of CG dinucleotides in the island. \ The Percentage CpG is the ratio of CpG nucleotide bases\ (twice the CpG count) to the length. The ratio of observed to expected \ CpG is calculated according to the formula (cited in \ Gardiner-Garden et al. (1987)):\ \
Obs/Exp CpG = Number of CpG * N / (Number of C * Number of G)\ \ where N = length of sequence.\
\ The calculation of the track data is performed by the following command sequence:\
\
twoBitToFa assembly.2bit stdout | maskOutFa stdin hard stdout \\\
| cpg_lh /dev/stdin 2> cpg_lh.err \\\
| awk '{$2 = $2 - 1; width = $3 - $2; printf("%s\\t%d\\t%s\\t%s %s\\t%s\\t%s\\t%0.0f\\t%0.1f\\t%s\\t%s\\n", $1, $2, $3, $5, $6, width, $6, width*$7*0.01, 100.0*2*$6/width, $7, $9);}' \\\
| sort -k1,1 -k2,2n > cpgIsland.bed\
\
The unmasked track data is constructed from\
twoBitToFa -noMask output for the twoBitToFa command.\
\
\
\ CpG islands and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator.\ All the tables can also be queried directly from our public MySQL\ servers, with more information available on our\ help page as well as on\ our blog.
\\ The source for the cpg_lh program can be obtained from\ src/utils/cpgIslandExt/.\ The cpg_lh program binary can be obtained from: http://hgdownload.soe.ucsc.edu/admin/exe/linux.x86_64/cpg_lh (choose "save file")\
\ \This track was generated using a modification of a program developed by G. Micklem and L. Hillier \ (unpublished).
\ \\ Gardiner-Garden M, Frommer M.\ \ CpG islands in vertebrate genomes.\ J Mol Biol. 1987 Jul 20;196(2):261-82.\ PMID: 3656447\
\ regulation 1 altColor 128,228,128\ color 0,100,0\ group regulation\ html cpgIslandSuper\ longLabel CpG Islands (Islands < 300 Bases are Light Green)\ shortLabel CpG Islands\ superTrack on\ track cpgIslandSuper\ type bed 4 +\ crisprAllTargets CRISPR Targets bigBed 9 + CRISPR/Cas9 -NGG Targets, whole genome 0 100 0 0 0 127 127 127 0 0 0 http://crispor.gi.ucsc.edu/crispor.py?org=$D&pos=$S:${&pam=NGG\ This track shows the DNA sequences targetable by CRISPR RNA guides using\ the Cas9 enzyme from S. pyogenes (PAM: NGG) over the entire\ human (hg38) genome. CRISPR target sites were annotated with\ predicted specificity (off-target effects) and predicted efficiency\ (on-target cleavage) by various\ algorithms through the tool CRISPOR. Sp-Cas9 usually cuts double-stranded DNA three or \ four base pairs 5' of the PAM site.\
\ \\ The track "CRISPR Targets" shows all potential -NGG target sites across the genome.\ The target sequence of the guide is shown with a thick (exon) bar. The PAM\ motif match (NGG) is shown with a thinner bar. Guides\ are colored to reflect both predicted specificity and efficiency. Specificity\ reflects the "uniqueness" of a 20mer sequence in the genome; the less unique a\ sequence is, the more likely it is to cleave other locations of the genome\ (off-target effects). Efficiency is the frequency of cleavage at the target\ site (on-target efficiency).
\ \Shades of gray stand for sites that are hard to target specifically, as the\ 20mer is not very unique in the genome:
\| impossible to target: target site has at least one identical copy in the genome and was not scored | |
| hard to target: many similar sequences in the genome that alignment stopped, repeat? | |
| hard to target: target site was aligned but results in a low specificity score <= 50 (see below) |
Colors highlight targets that are specific in the genome (MIT specificity > 50) but have different predicted efficiencies:
\| unable to calculate Doench/Fusi 2016 efficiency score | |
| low predicted cleavage: Doench/Fusi 2016 Efficiency percentile <= 30 | |
| medium predicted cleavage: Doench/Fusi 2016 Efficiency percentile > 30 and < 55 | |
| high predicted cleavage: Doench/Fusi 2016 Efficiency > 55 |
\
Mouse-over a target site to show predicted specificity and efficiency scores:
\
Click onto features to show all scores and predicted off-targets with up to\ four mismatches. The Out-of-Frame score by Bae et al. 2014\ is correlated with\ the probability that mutations induced by the guide RNA will disrupt the open\ reading frame. The authors recommend out-of-frame scores > 66 to create\ knock-outs with a single guide efficiently.
\ \
Off-target sites are sorted by the CFD (Cutting Frequency Determination)\ score (Doench et al. 2016).\ The higher the CFD score, the more likely there is off-target cleavage at that site.\ Off-targets with a CFD score < 0.023 are not shown on this page, but are available when\ following the link to the external CRISPOR tool.\ When compared against experimentally validated off-targets by\ Haeussler et al. 2016, the large majority of predicted\ off-targets with CFD scores < 0.023 were false-positives. For storage and performance\ reasons, on the level of individual off-targets, only CFD scores are available.
\ \\ Like most algorithms, the MIT specificity score is not always a perfect\ predictor of off-target effects. Despite low scores, many tested guides\ caused few and/or weak off-target cleavage when tested with whole-genome assays\ (Figure 2 from Haeussler\ et al. 2016), as shown below, and the published data contains few data points\ with high specificity scores. Overall though, the assays showed that the higher\ the specificity score, the lower the off-target effects.
\ \
\
\
Similarly, efficiency scoring is not very accurate: guides with low\ scores can be efficient and vice versa. As a general rule, however, the higher\ the score, the less likely that a guide is very inefficient. The\ following histograms illustrate, for each type of score, how the share of\ inefficient guides drops with increasing efficiency scores:\
\ \
\
\
When reading this plot, keep in mind that both scores were evaluated on\ their own training data. Especially for the Moreno-Mateos score, the\ results are too optimistic, due to overfitting. When evaluated on independent\ datasets, the correlation of the prediction with other assays was around 25%\ lower, see Haeussler et al. 2016. At the time of\ writing, there is no independent dataset available yet to determine the\ Moreno-Mateos accuracy for each score percentile range.
\ \\ The entire human (hg38) genome was scanned for the -NGG motif. Flanking 20mer\ guide sequences were\ aligned to the genome with BWA and scored with MIT Specificity scores using the\ command-line version of crispor.org. Non-unique guide sequences were skipped.\ Flanking sequences were extracted from the genome and input for Crispor\ efficiency scoring, available from the Crispor downloads page, which\ includes the Doench 2016, Moreno-Mateos 2015 and Bae\ 2014 algorithms, among others.
\\ Note that the Doench 2016 scores were updated by\ the Broad institute in 2017 ("Azimuth" update). As a result, earlier versions of\ the track show the old Doench 2016 scores and this version of the track shows new\ Doench 2016 scores. Old and new scores are almost identical, they are\ correlated to 0.99 and for more than 80% of the guides the difference is below 0.02.\ However, for very few guides, the difference can be bigger. In case of doubt, we recommend\ the new scores. Crispor.org can display both\ scores and many more with the "Show all scores" link.
\ \\ Positional data can be explored interactively with the \ Table\ Browser or the Data Integrator.\ For small programmatic positional queries, the track can be accessed using our \ REST API. For genome-wide data or \ automated analysis, CRISPR genome annotations can be downloaded from\ our download server\ as a bigBedFile.
\\ The files for this track are called crispr.bb, which lists positions and\ scores, and crisprDetails.tab, which has information about off-target matches. Individual\ regions or whole genome annotations can be obtained using our tool bigBedToBed,\ which can be compiled from the source code or downloaded as a pre-compiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here. The tool\ can also be used to obtain only features within a given range, e.g.
\\ bigBedToBed\ http://hgdownload.soe.ucsc.edu/gbdb/hg38/crisprAllTargets/crispr.bb -chrom=chr21\ -start=0 -end=1000000 stdout
\ \\ Track created by Maximilian Haeussler, with helpful input\ from Jean-Paul Concordet (MNHN Paris) and Alberto Stolfi (NYU).\
\ \\ Haeussler M, Schönig K, Eckert H, Eschstruth A, Mianné J, Renaud JB, Schneider-Maunoury S,\ Shkumatava A, Teboul L, Kent J et al.\ Evaluation of off-target and on-target scoring algorithms and integration into the\ guide RNA selection tool CRISPOR.\ Genome Biol. 2016 Jul 5;17(1):148.\ PMID: 27380939; PMC: PMC4934014\
\ \\ Bae S, Kweon J, Kim HS, Kim JS.\ \ Microhomology-based choice of Cas9 nuclease target sites.\ Nat Methods. 2014 Jul;11(7):705-6.\ PMID: 24972169\
\ \\ Doench JG, Fusi N, Sullender M, Hegde M, Vaimberg EW, Donovan KF, Smith I, Tothova Z, Wilen C,\ Orchard R et al.\ \ Optimized sgRNA design to maximize activity and minimize off-target effects of CRISPR-Cas9.\ Nat Biotechnol. 2016 Feb;34(2):184-91.\ PMID: 26780180; PMC: PMC4744125\
\ \\ Hsu PD, Scott DA, Weinstein JA, Ran FA, Konermann S, Agarwala V, Li Y, Fine EJ, Wu X, Shalem O\ et al.\ \ DNA targeting specificity of RNA-guided Cas9 nucleases.\ Nat Biotechnol. 2013 Sep;31(9):827-32.\ PMID: 23873081; PMC: PMC3969858\
\ \\ Moreno-Mateos MA, Vejnar CE, Beaudoin JD, Fernandez JP, Mis EK, Khokha MK, Giraldez AJ.\ \ CRISPRscan: designing highly efficient sgRNAs for CRISPR-Cas9 targeting in vivo.\ Nat Methods. 2015 Oct;12(10):982-8.\ PMID: 26322839; PMC: PMC4589495\
\ genes 1 bigDataUrl /gbdb/hg38/crisprAll/crispr.bb\ denseCoverage 0\ detailsTabUrls _offset=/gbdb/$db/crisprAll/crisprDetails.tab\ group genes\ html crisprAll\ itemRgb on\ longLabel CRISPR/Cas9 -NGG Targets, whole genome\ mouseOverField _mouseOver\ noGenomeReason This track is too big for whole-genome Table Browser access, it would lead to a timeout in your internet browser. Small regional queries can work, but large regions, such as entire chromosomes, will fail. Please see the CRISPR Track documentation, the section "Data Access", for bulk-download options and remote access via the bedToBigBed tool. API access should always work. Contact us if you encounter difficulties with accessing the data.\ scoreFilterMax 100\ scoreLabel MIT Guide Specificity Score\ shortLabel CRISPR Targets\ tableBrowser tbNoGenome\ track crisprAllTargets\ type bigBed 9 +\ url http://crispor.gi.ucsc.edu/crispor.py?org=$D&pos=$S:${&pam=NGG\ urlLabel Click here to show this guide on Crispor.org, with expression oligos, validation primers and more\ visibility hide\ crossTissueMaps Cross Tissue Nuclei Single Nuclei sequenced across many tissues 0 100 0 0 0 127 127 127 0 0 0\
This track collection shows data from \
Single-nucleus cross-tissue molecular reference maps toward\
understanding disease gene function. The dataset covers ~200,000 single nuclei\
from a total of 16 human donors across 25 samples, using 4 different sample preparation\
protocols followed by droplet based single-cell RNA-seq. The samples were obtained from\
frozen tissue as part of the Genotype-Tissue Expression (GTEx) project.\
Samples were taken from the esophagus, skeletal muscle, heart, lung, prostate, breast,\
and skin. The dataset includes 43 broad cell classes, some specific to certain tissues\
and some shared across all tissue types.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ This track collection contains three bar chart tracks of RNA expression. The first track,\ Cross Tissue Nuclei, allows\ cells to be grouped together and faceted on up to 4 categories: tissue, cell class, cell subclass,\ and cell type. The second track,\ Cross Tissue Details, allows\ cells to be grouped together and faceted on up to 7 categories: tissue, cell class, cell subclass,\ cell type, granular cell type, sex, and donor. The third track,\ GTEx Immune Atlas,\ allows cells to be grouped together and faceted on up to 5 categories: tissue, cell type, cell\ class, sex, and donor.\
\ \\ Please see the\ GTEx portal\ for further interactive displays and additional data.
\ \\ Tissue-cell type combinations in the Full and Combined tracks are\ colored by which cell type they belong to in the below table:\
\
| Color | \Cell Type | \
|---|---|
| Endothelial | |
| Epithelial | |
| Glia | |
| Immune | |
| Neuron | |
| Stromal | |
| Other |
\ Tissue-cell type combinations in the Immune Atlas track are shaded according\ to the below table:\
| Color | \Cell Type | \
|---|---|
| Inflammatory Macrophage | |
| Lung Macrophage | |
| Monocyte/Macrophage FCGR3A High | |
| Monocyte/Macrophage FCGR3A Low | |
| Macrophage HLAII High | |
| Macrophage LYVE1 High | |
| Proliferating Macrophage | |
| Dendritic Cell 1 | |
| Dendritic Cell 2 | |
| Mature Dendritic Cell | |
| Langerhans | |
| CD14+ Monocyte | |
| CD16+ Monocyte | |
| LAM-like | |
| Other |
\ Using the previously collected tissue samples from the Genotype-Tissue Expression\ project, nuclei were isolated using four different protocols and sequenced\ using droplet based single cell RNA-seq. CellBender v2.1 and other standard quality\ control techniques were applied, resulting in 209,126 nuclei profiles across eight\ tissues, with a mean of 918 genes and 1519 transcripts per profile.\
\ \\ Data from all samples was integrated with a conditional variation autoencoder\ in order to correct for multiple sources of variation like sex, and protocol\ while preserving tissue and cell type specific effects.\
\ \\ For detailed methods, please refer to Eraslan et al, or the\ \ GTEx portal website.\
\ \\
The gene expression files were downloaded from the\
\
GTEx portal. The UCSC command line utilities matrixClusterColumns,\
matrixToBarChartBed, and bedToBigBed were used to transform\
these into a bar chart format bigBed file that can be visualized.\
The UCSC utilities can be found on\
our download server.\
\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions or our Data Access FAQ for more\ information.
\ \\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the GTEx Consortium for creating and analyzing these data.
\ \\ Eraslan G, Drokhlyansky E, Anand S, Fiskin E, Subramanian A, Slyper M, Wang J, Van Wittenberghe N,\ Rouhana JM, Waldman J et al.\ \ Single-nucleus cross-tissue molecular reference maps toward understanding disease gene function.\ Science. 2022 May 13;376(6594):eabl4290.\ PMID: 35549429; PMC: PMC9383269\
\ singleCell 0 configureByPopup off\ group singleCell\ longLabel Single Nuclei sequenced across many tissues\ shortLabel Cross Tissue Nuclei\ superTrack on\ track crossTissueMaps\ visibility hide\ dbSnpArchive dbSNP Archive bed 6 + dbSNP Track Archive 0 100 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This composite track contains information about single nucleotide polymorphisms (SNPs)\ and small insertions and deletions (indels) — collectively Simple\ Nucleotide Polymorphisms — from\ dbSNP, available from\ ftp.ncbi.nih.gov/snp.\ You can click into each track for a version/subset-specific description.
\\ This collection includes numbered versions of the entire dbSNP datasets\ (All SNP) as well as three tracks with subsets of the items in that version. \ Here is information on each of the subsets:\
\ The default maximum weight for this track is 1, so unless\ the setting is changed in the track controls, SNPs that map to multiple genomic \ locations will be omitted from display. When a SNP's flanking sequences \ map to multiple locations in the reference genome, it calls into question \ whether there is true variation at those sites, or whether the sequences\ at those sites are merely highly similar but not identical.\
\ \\ Variants are shown as single tick marks at most zoom levels.\ When viewing the track at or near base-level resolution, the displayed\ width of the SNP corresponds to the width of the variant in the reference\ sequence. Insertions are indicated by a single tick mark displayed between\ two nucleotides, single nucleotide polymorphisms are displayed as the width \ of a single base, and multiple nucleotide variants are represented by a \ block that spans two or more bases.\
\ \\ On the track controls page, SNPs can be colored and/or filtered from the \ display according to several attributes:\
\\ You can configure this track such that the details page displays\ the function and coding differences relative to \ particular gene sets. Choose the gene sets from the list on the SNP \ configuration page displayed beneath this heading: On details page,\ show function and coding differences relative to. \ When one or more gene tracks are selected, the SNP details page \ lists all genes that the SNP hits (or is close to), with the same keywords \ used in the function category. The function usually \ agrees with NCBI's function, except when NCBI's functional annotation is \ relative to an XM_* predicted RefSeq (not included in the UCSC Genome \ Browser's RefSeq Genes track) and/or UCSC's functional annotation is \ relative to a transcript that is not in RefSeq.\
\ \\ dbSNP uses a class called 'in-del'. We compare the length of the\ reference allele to the length(s) of observed alleles; if the\ reference allele is shorter than all other observed alleles, we change\ 'in-del' to 'insertion'. Likewise, if the reference allele is longer\ than all other observed alleles, we change 'in-del' to 'deletion'.\
\ \\ dbSNP determines the genomic locations of SNPs by aligning their flanking \ sequences to the genome.\ UCSC displays SNPs in the locations determined by dbSNP, but does not\ have access to the alignments on which dbSNP based its mappings.\ Instead, UCSC re-aligns the flanking sequences \ to the neighboring genomic sequence for display on SNP details pages. \ While the recomputed alignments may differ from dbSNP's alignments,\ they often are informative when UCSC has annotated an unusual condition.\
\\ Non-repetitive genomic sequence is shown in upper case like the flanking \ sequence, and a "|" indicates each match between genomic and flanking bases.\ Repetitive genomic sequence (annotated by RepeatMasker and/or the\ Tandem Repeats Finder with period <= 12) is shown in lower case, and matching\ bases are indicated by a "+".\
\ \\ The data that comprise this track were extracted from database dump files \ and headers of fasta files downloaded from NCBI. \ The database dump files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/database/\ (for human, organism_tax_id = human_9606;\ for mouse, organism_tax_id = mouse_10090).\ The fasta files were downloaded from \ ftp://ftp.ncbi.nih.gov/snp/organisms/\ organism_tax_id/rs_fasta/\
\\ Note: It is not recommeneded to use LiftOver to convert SNPs between assemblies,\ and more information about how to convert SNPs between assemblies can be found on the following\ FAQ entry.
\\ The raw data can be explored interactively with the \ Table Browser,\ Data Integrator, or \ Variant Annotation Integrator.\ For automated analysis, the genome annotation files can be downloaded in their entirety for \ hg38,\ hg19, \ and mm10 as\ (snp*.txt.gz). \ You can also make queries using the UCSC Genome Browser \ JSON API or \ public MySQL server. Please refer to our \ mailing list archives\ for questions and example queries, or our \ Data Access FAQ for more information.\
\ \\ For the human assembly, we provide a related table that contains\ orthologous alleles in the chimpanzee, orangutan and rhesus macaque\ reference genome assemblies. \ We use our liftOver utility to identify the orthologous alleles. \ The candidate human SNPs are a filtered list that meet the criteria:\
\ Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. \ dbSNP: the NCBI database of genetic variation.\ Nucleic Acids Res. 2001 Jan 1;29(1):308-11.\ PMID: 11125122; PMC: PMC29783\
\ varRep 1 cartVersion 3\ group varRep\ html ../../dbSnpArchive\ longLabel dbSNP Track Archive\ maxWindowToDraw 10000000\ shortLabel dbSNP Archive\ superTrack on\ track dbSnpArchive\ type bed 6 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ dbVar_common dbVar Common SV bigBed 9 + . NCBI dbVar Curated Common Structural Variants 3 100 0 0 0 127 127 127 0 0 0\ This track displays common structural variants (SVs) from\ nstd186\ (NCBI Curated Common Structural Variants), divided into subtracks by source study and by\ population.\
\ \\ nstd186 is a curated collection of structural variants in\ dbVar from studies with at least\ 100 samples, that include allele frequency data, and that have an allele frequency of >=0.01\ in at least one population. It includes copy number gains and losses, copy number variations,\ duplications, deletions, insertions, and mobile element variants (ALU, LINE1, SVA, HERV).\
\ \\ The dataset aggregates variants from six source studies:\
\\ For the latest nstd186 variant call counts and version history, see the\ nstd186\ summary page at NCBI.\
\ \\ Per-source-study subtracks (variants from nstd186 attributed to one of the six component\ studies):\
\\ Per-population subtracks (variants with AF >= 0.01 aggregated across nstd186 source\ studies for each super-population):\
\\ The NCBI dbVar\ Track Hub additionally provides population-only variants (variants common in one\ population but not in any other): African only, American only, East Asian only, European only,\ and South Asian only. These are not loaded as native Genome Browser tracks; connect to the hub to\ view them.\
\ \\ Items in all subtracks follow the same conventions. Variants are colored by type, using the dbVar\ color scheme described in the\ dbVar Overview\ page:\
\| Color | \Variant Type(s) | \
|---|---|
| copy number loss, deletion (including mobile element deletions) | |
| copy number gain, duplication, insertion (including mobile element insertions) | |
| copy number variation |
\ Mouseover on items shows genes affected, size, variant type, allele count (AC), allele\ number (AN), allele frequency (AF), and population (in per-population subtracks).\
\ \\ Subtracks can be filtered by:\
\\ The Hide empty subtracks option on the track configuration page hides subtracks that have\ no data in the current viewing window. This is enabled by default and can be toggled off.\
\ \\ The raw data can be explored interactively with the\ Table Browser, or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API.\
\\ The data can also be found directly at the\ dbVar\ nstd186 data access page, or in the\ dbVar\ Track Hub. For questions about dbVar track data, please contact\ dbvar@ncbi.nlm.nih.gov.\ \
\ \\ Thanks to the dbVar team at NCBI, especially John Lopez and Timothy Hefferon for technical\ coordination and consultation, and to Christopher Lee, Anna Benet-Pages, and Daniel Schmelter, of\ the Genome Browser team for engineering the track display.\
\ \\ Lappalainen I, Lopez J, Skipper L, Hefferon T, Spalding JD, Garner J, Chen C, Maguire M, Corbett M,\ Zhou G et al.\ \ DbVar and DGVa: public archives for genomic structural variation.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D936-41.\ PMID: 23193291;\ PMC: PMC3531204\
\ varRep 1 compositeTrack on\ filterLabel.freq_range Frequency Range\ filterLabel.length Variant Size\ filterLabel.type Variant Type\ filterValues.freq_range Under 0.02,0.02 to 0.05,0.05 to 0.1,0.1 to 0.2,0.2 to 0.5,Over 0.5\ filterValues.length Under 10KB,10KB to 100KB,100KB to 1MB,Over 1MB\ filterValues.type alu deletion,alu insertion,copy number gain,copy number loss,copy number variation,deletion,duplication,herv deletion,insertion,line1 deletion,line1 insertion,mobile element deletion,mobile element insertion,sva deletion,sva insertion\ hideEmptySubtracks on\ html dbVarCommon\ itemRgb on\ longLabel NCBI dbVar Curated Common Structural Variants\ mouseOverField label\ searchIndex name\ shortLabel dbVar Common SV\ superTrack dbVarSv pack\ track dbVar_common\ type bigBed 9 + .\ visibility pack\ dbVar_conflict dbVar Conflict SV bigBed 9 + . NCBI dbVar Curated Conflict Variants 3 100 0 0 0 127 127 127 0 0 0\ The track NCBI dbVar Curated Common SVs: Conflicts with Pathogenic highlights loci where\ common copy number variants from\ nstd186 (NCBI Curated\ Common Structural Variants) overlap with structural variants with clinical assertions,\ submitted to ClinVar by external labs (Clinical Structural\ Variants - nstd102).\
\ \\ Overlap in the track refers to reciprocal overlap between variants in the common\ (NCBI Curated Common Structural Variants) versus clinical (ClinVar CNVs)\ tracks. Reciprocal overlap values can be anywhere from 10% to 100%.\
\ \\ For more information on the number of variant calls and latest statistics for nstd186 see\ Summary of nstd186\ (NCBI Curated Common Structural Variants).\
\ \\ Items in this track follow the same conventions as the parent Common SV track: items are colored\ by variant type, based on the dbVar colors described in the\ dbVar Overview page.\ The variant types present in this track are copy number gain, copy number loss, copy number\ variation, deletion, and duplication.\
\| Color | \Variant Type(s) | \
|---|---|
| copy number loss, deletion | |
| copy number gain, duplication | |
| copy number variation |
\ Mouseover on items indicates genes affected, size, variant type, and allele frequencies (AF). \ All tracks can be filtered according to the variant length, variant type and \ variant overlap. The overlap filter defines five bins within that range (10-25,\ 25-50, 50-75, 75-90, 90-100 percent reciprocal overlap; intervals are inclusive of the upper bound).\
\ \ \\ The raw data can be explored interactively with the\ Table Browser, or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API.\
\ \\ The data can also be found directly from the dbVar \ nstd186 data access, as well as in the\ \ dbVar Track Hub, where additional subtracks are included. For questions about\ dbVar track data, please contact\ dbvar@ncbi.nlm.nih.gov.\ \
\ \\ Thanks to the dbVar team at NCBI, especially John Lopez and Timothy Hefferon for technical\ coordination and consultation, and to Christopher Lee, Anna Benet-Pages, and Daniel Schmelter of\ the Genome Browser team for engineering the track display.\
\ \\ Lappalainen I, Lopez J, Skipper L, Hefferon T, Spalding JD, Garner J, Chen C, Maguire M, Corbett M,\ Zhou G et al.\ \ DbVar and DGVa: public archives for genomic structural variation.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D936-41.\ PMID: 23193291; PMC: PMC3531204\
\ \ varRep 1 compositeTrack on\ filterLabel.length Variant Size\ filterLabel.overlap Variant Overlap\ filterLabel.type Variant Type\ filterValues.length Under 10KB,10KB to 100KB,100KB to 1MB,Over 1MB\ filterValues.overlap 10 to 25,25 to 50,50 to 75,75 to 90,90 to 100\ filterValues.type copy number gain,copy number loss,copy number variation,deletion,duplication\ html dbVarConflict\ itemRgb on\ longLabel NCBI dbVar Curated Conflict Variants\ mouseOverField label\ searchIndex name\ shortLabel dbVar Conflict SV\ superTrack dbVarSv pack\ track dbVar_conflict\ type bigBed 9 + .\ visibility pack\ dbVar_other dbVar Other SV bigBed 9 + . NCBI dbVar Other Structural Variants 0 100 0 0 0 127 127 127 0 0 0\ This track displays structural variants (SVs) in\ dbVar that are not classified as\ common, somatic, or clinical. The track is defined by exclusion: it contains dbVar SVs minus\
\\ NCBI sometimes refers to this category as presumed normal SVs in their hub documentation\ and source files. We use the term Other here to avoid implying that the variants are\ clinically normal — the track is purely a residual bucket of dbVar SVs that don't fit the\ other three composites.\
\ \\ This track is updated with every monthly dbVar release.\
\ \\ The Other SVs are split into two subtracks:\
\\ The Healthy subtrack is considerably larger than the Phenotype subtrack. Turning on\ Hide empty subtracks (default) limits the display to subtracks with data in the current\ viewing window.\
\ \\ Variants are colored by type, using the dbVar color scheme described in the\ dbVar Overview\ page:\
\| Color | \Variant Type(s) | \
|---|---|
| deletion, delins, copy number loss | |
| duplication, copy number gain, insertion | |
| copy number variation | |
| inversion | |
| complex substitution | |
| tandem duplication | |
| sequence alteration |
\ Mouseover on items shows gene(s) affected, size, variant type, dbVar study of origin,\ discovery method, phenotype (in the Phenotype subtrack), and population code (if available).\
\ \\ Subtracks can be filtered by:\
\\
Per NCBI's dbVar processing pipeline, variant calls are extracted from the\
variant_calls.gvf files on the dbVar FTP site, reciprocally overlapped with the\
pathogenic clinical SV file using bedtools, filtered by the exclusion criteria described above,\
and converted to bigBed format. See the\
dbVar\
Overview for full methods.\
\ The raw data can be explored interactively with the\ Table Browser, or the\ Data Integrator. Due to the size of the Healthy subtrack (over\ 5 million items), Table Browser queries on large regions may be slow — narrow by\ chromosome or region where possible.\
\\ The data can also be downloaded from the\ dbVar Track Hub.\ For questions about dbVar track data, please contact\ dbvar@ncbi.nlm.nih.gov.\ \
\ \\ Thanks to the dbVar team at NCBI, especially John Lopez and Timothy Hefferon for technical\ coordination and consultation.\
\ \\ Lappalainen I, Lopez J, Skipper L, Hefferon T, Spalding JD, Garner J, Chen C, Maguire M, Corbett M,\ Zhou G et al.\ \ DbVar and DGVa: public archives for genomic structural variation.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D936-41.\ PMID: 23193291;\ PMC: PMC3531204\
\ varRep 1 compositeTrack on\ filterLabel.length Variant Size\ filterLabel.method Discovery Method\ filterLabel.overlap Pathogenic Reciprocal Overlap\ filterLabel.population Population Code\ filterLabel.type Variant Type\ filterValues.length Under 10KB,10KB to 100KB,100KB to 1MB,Over 1MB\ filterValues.method Curated,Merging,Multiple,Oligo aCGH,Optical mapping,SNP array,Sequencing,other\ filterValues.overlap none,10 to 25,25 to 50,50 to 75,75 to 90,90 to 100\ filterValues.population AFR,AMR,EAS,EUR,OTH,SAS,mixed,multiple,none,unknown\ filterValues.type alu deletion,alu insertion,complex substitution,copy-neutral loss of heterozygosity,copy number gain,copy number loss,copy number variation,deletion,delins,duplication,herv deletion,herv insertion,insertion,inversion,line1 deletion,line1 insertion,mobile element deletion,mobile element insertion,novel sequence insertion,sequence alteration,sva deletion,sva insertion,tandem duplication\ hideEmptySubtracks on\ html dbVarOther\ itemRgb on\ longLabel NCBI dbVar Other Structural Variants\ mouseOverField label\ searchIndex name\ shortLabel dbVar Other SV\ superTrack dbVarSv\ track dbVar_other\ type bigBed 9 + .\ visibility hide\ dbVar_somatic dbVar Somatic SV bigBed 9 + . NCBI dbVar Somatic Structural Variants 0 100 0 0 0 127 127 127 0 0 0\ This track displays structural variants (SVs) in\ dbVar with somatic origin,\ aggregated from six dbVar studies.\
\ \\ Source studies:\
\\ This track is updated with every monthly dbVar release.\
\ \\ Variants are colored by type, using the dbVar color scheme described in the\ dbVar Overview\ page:\
\| Color | \Variant Type(s) | \
|---|---|
| deletion, copy number loss | |
| duplication, copy number gain, insertion, mobile element insertion | |
| inversion | |
| complex substitution | |
| tandem duplication |
\ Mouseover on items shows gene(s) affected, size, variant type, source dbVar study, and\ discovery method.\
\ \\ The track can be filtered by:\
\\
Per NCBI's dbVar processing pipeline, somatic variant calls are extracted from the\
variant_calls.somatic.gvf files on the dbVar FTP site, reciprocally overlapped\
with the pathogenic clinical SV file using bedtools, and converted to bigBed format. See the\
dbVar\
Overview for full methods.\
\ The raw data can be explored interactively with the\ Table Browser, or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API.\
\\ The data can also be downloaded from the\ dbVar Track Hub,\ or via the dbVar FTP in VCF, GVF, or tab-delimited formats. For questions about dbVar track data,\ please contact\ dbvar@ncbi.nlm.nih.gov.\ \
\ \\ Thanks to the dbVar team at NCBI, especially John Lopez and Timothy Hefferon for technical\ coordination and consultation.\
\ \\ Lappalainen I, Lopez J, Skipper L, Hefferon T, Spalding JD, Garner J, Chen C, Maguire M, Corbett M,\ Zhou G et al.\ \ DbVar and DGVa: public archives for genomic structural variation.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D936-41.\ PMID: 23193291;\ PMC: PMC3531204\
\\ Tate JG, Bamford S, Jubb HC, Sondka Z, Beare DM, Bindal N, Boutselakis H, Cole CG, Creatore C,\ Dawson E et al.\ \ COSMIC: the Catalogue Of Somatic Mutations In Cancer.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D941-D947.\ PMID: 30371878;\ PMC: PMC6323903\
\ varRep 1 compositeTrack on\ filterLabel.length Variant Size\ filterLabel.method Discovery Method\ filterLabel.overlap Pathogenic Reciprocal Overlap\ filterLabel.type Variant Type\ filterValues.length Under 10KB,10KB to 100KB,100KB to 1MB,Over 1MB\ filterValues.method Curated,Multiple,SNP array,Sequencing\ filterValues.overlap none,10 to 25,25 to 50,50 to 75,75 to 90,90 to 100\ filterValues.type complex substitution,copy number gain,copy number loss,copy-neutral loss of heterozygosity,deletion,duplication,insertion,inversion,mobile element insertion,tandem duplication\ html dbVarSomatic\ itemRgb on\ longLabel NCBI dbVar Somatic Structural Variants\ mouseOverField label\ searchIndex name\ shortLabel dbVar Somatic SV\ superTrack dbVarSv\ track dbVar_somatic\ type bigBed 9 + .\ visibility hide\ dbVarSv dbVar Struct Var NCBI dbVar Structural Variants 0 100 0 0 0 127 127 127 0 0 0\ This super-track groups structural variant (SV) tracks from\ dbVar, NCBI's archive of human\ genomic structural variation. The data are mirrored from the\ NCBI dbVar track\ hub.\
\ \\ There are four track collections in this super-track:\
\\ Clinical structural variants from dbVar study nstd102 are not duplicated here; they are available\ in our dedicated ClinVar track (subtrack\ ClinVar CNVs), which pulls from the same underlying ClinVar XML release.\
\ \\ nstd186 is a\ curated collection of SVs from studies with at least 100 samples and allele frequency >= 0.01\ in at least one population. It aggregates data from six source studies:\
\\ Variants must be of a qualifying structural variant type (deletions, duplications, insertions,\ copy number variants, and mobile element variants). For the latest statistics and version\ history, see the\ nstd186 summary\ page at NCBI.\
\ \\ These tracks are composite tracks that contain multiple subtracks. Each subtrack has its own\ display controls, as described here. Items are\ colored by variant type using the dbVar color scheme\ (dbVar Overview):\
\| Color | \Variant Type(s) | \
|---|---|
| deletion, copy number loss | |
| duplication, copy number gain, insertion | |
| copy number variation |
\ Some composites display additional colors for less common variant types. Refer to each composite\ track's description page for the full legend.\
\ \\ The raw data can be explored interactively with the\ Table Browser, or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API.\
\\ The data can also be found directly at the\ dbVar\ nstd186 data access page, or in the\ dbVar\ Track Hub, where additional subtracks (e.g., population-exclusive variants, ClinVar SVs) are\ available. For questions about dbVar track data, please contact\ dbvar@ncbi.nlm.nih.gov.\ \
\ \\ Thanks to the dbVar team at NCBI, especially John Lopez and Timothy Hefferon for technical\ coordination and consultation, and to Christopher Lee, Anna Benet-Pages, and Daniel Schmelter of\ the Genome Browser team for engineering the track display.\
\ \\ Lappalainen I, Lopez J, Skipper L, Hefferon T, Spalding JD, Garner J, Chen C, Maguire M, Corbett M,\ Zhou G et al.\ \ DbVar and DGVa: public archives for genomic structural variation.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D936-41.\ PMID: 23193291;\ PMC: PMC3531204\
\ varRep 0 dataVersion /gbdb/$D/bbi/dbVar/version.txt\ group varRep\ html dbVarCurated\ longLabel NCBI dbVar Structural Variants\ shortLabel dbVar Struct Var\ superTrack on\ track dbVarSv\ decipherContainer DECIPHER DECIPHER 0 100 0 0 0 127 127 127 0 0 0NOTE:
\
While the DECIPHER database is \
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions.\
Because the UCSC Genes mappings for CNVs are based on associations from\ RefSeq and UniProt, they are dependent on any interpretations from those\ sources. Furthermore, because many DECIPHER records refer to multiple gene\ names, or syndromes not tightly mapped to individual genes, the associations\ in this track should be treated with skepticism and any conclusions\ based on them should be carefully scrutinized using independent\ resources.\
\Data Display Agreement Notice
\
The CNV/SNV data are only available for display in the Browser, and not for bulk\
download. Access to bulk data may be obtained directly from DECIPHER\
(https://www.deciphergenomics.org/about/data-sharing) and is subject to a\
Data Access Agreement, in which the user certifies that no attempt to\
identify individual patients will be undertaken. The same restrictions\
apply to the public data displayed at UCSC in the UCSC Genome Browser;\
no one is authorized to attempt to identify patients by any means.\
These data are made available as soon as possible and may be a\ pre-publication release. For information on the proper use of DECIPHER\ data, please see https://www.deciphergenomics.org/about/data-sharing.\
\The DECIPHER consortium provides these data in good faith as a research\ tool, but without verifying the accuracy, clinical validity, or utility of\ the data. The DECIPHER consortium makes no warranty, express or implied,\ nor assumes any legal liability or responsibility for any purpose for\ which the data are used.\
\\ The \ DECIPHER\ database of submicroscopic chromosomal imbalance \ collects clinical information about chromosomal \ microdeletions/duplications/insertions, translocations and inversions, \ and displays this information on the human genome map.\
\ The CNVs and SNVs tracks show genomic regions of reported cases and their \ associated phenotype information. All data have passed the strict\ consent requirements of the DECIPHER project and are approved for\ unrestricted public release. Clicking the Patient View ID link\ brings up a more detailed informational page on the patient at the \ DECIPHER web site.
\ \\ The Population CNVs track shows common copy-number variants (CNVs) and their\ population frequencies, lifted over from the hg19 assembly.
\ \\ The genomic locations of DECIPHER variants are labeled with the DECIPHER variant descriptions. \ Mouseover on items shows variant details, clinical interpretation, and associated conditions. \ Further information on each variant is displayed on the details page by a click onto any variant. \
\ \\ For the CNVs track, the entries are colored by the type of variant:\
\ A light-to-dark color gradient indicates the clinical significance of each variant, with \ the lightest shade being benign, to the darkest shade being pathogenic. Detailed information on the \ CNV color code is described here.\ Items can be filtered according to the size of the variant, variant type, and clinical significance \ using the track Configure options.\
\ \\ For the SNVs track, the entries are colored according to the estimated clinical significance \ of the variant:\
\ For the Population CNVs track, genomic variants are visually differentiated to facilitate quick and\ clear identification. Variants are colored according to their clinical significance and type:\
\\ The Population CNVs track's mouseover tooltip provides the following information\ about the data:\
\\ Data provided by the DECIPHER project group are imported and processed\ to create a simple BED track to annotate the genomic regions associated\ with individual patients.\
\ \ \\ For more information on DECIPHER, please contact\ \ contact@deciphergenomics.\ org\
\ \\ The DECIPHER data access and documentation can be found at\ DECIPHER Downloads.\
\ \\ Firth HV, Richards SM, Bevan AP, Clayton S, Corpas M, Rajan D, Van Vooren S, Moreau Y, Pettett RM,\ Carter NP.\ \ DECIPHER: Database of Chromosomal Imbalance and Phenotype in Humans Using Ensembl Resources.\ Am J Hum Genet. 2009 Apr;84(4):524-33.\ PMID: 19344873; PMC: PMC2667985\
\ phenDis 0 cartVersion 7\ dataVersion /gbdb/$D/decipher/version.txt\ group phenDis\ longLabel DECIPHER\ shortLabel DECIPHER\ superTrack on\ track decipherContainer\ decodeSv deCODE 3622 SVs bigBed 9 + High-confidence Structural Variants from 3,622 Icelanders (deCODE, Oxford Nanopore) 0 100 0 0 0 127 127 127 0 0 0\ This track shows high-confidence structural variants (SVs) identified by\ Oxford Nanopore long-read sequencing of 3,622 Icelanders recruited through\ the deCODE genetics population cohort. The track contains 119,453 high-confidence\ SVs (41,216 deletions, 75,050 insertions and 3,187 combined insertion/deletion\ events), deduplicated from a 133,886-record upstream release. Variants are\ site-level (no per-sample genotypes) and have been\ filtered to a high-confidence subset validated in the accompanying\ population-scale analysis.\
\\ Note that this release does not include allele counts or allele frequencies:\ each row represents a site that was called with high confidence in the\ cohort, but the number of carrier samples is not provided, so the track\ cannot be filtered by AF/AC.\
\ \\ Items are colored by SV type:\
\ Insertions are placed at the insertion site with a width of 1 bp; deletions\ span the deleted interval; INSDEL events span the affected reference region\ and have SVLEN=0 because the reference and alternate alleles differ in both\ sequence and length. Filters are available for SV type and SV length.\
\\ Where a variant falls inside an annotated tandem-repeat region, the detail\ page also shows the coordinates of that region (TRRBEGIN / TRREND from the\ source VCF), which can be useful context for repeat-mediated insertions and\ deletions.\
\ \\ Beyter et al. 2021 performed Oxford Nanopore long-read sequencing of 3,622\ Icelanders recruited through deCODE genetics and detected a median of\ 22,636 SVs per individual (13,353 insertions and 9,474 deletions). Across\ the cohort they derived a set of 133,886 reliably genotyped SV alleles,\ imputed those alleles into 166,281 chip-typed Icelanders, and tested them\ for association with disease and quantitative traits (notably including a\ rare PCSK9 deletion associated with lower LDL-cholesterol and a\ multi-allelic 57-bp VNTR in ACAN associated with adult height). The\ track shown here displays 119,453 unique high-confidence SV sites (exact-duplicate\ records present in the release have been collapsed): 41,216 deletions, 75,050\ insertions and 3,187 combined insertion/deletion events.\ The release is site-only (no per-sample genotypes or allele frequencies),\ so the track cannot be filtered by AF/AC.\
\\ The VCF ont_sv_high_confidence_SVs.sorted.vcf.gz was downloaded\ from the deCODE genetics\ \ LRS_SV_sets GitHub repository.\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator and exported from there\ to spreadsheet or tab-sep tables. From scripts, the data can be accessed\ through our API, track=decodeSv.\
\\ The annotation is stored as a bigBed file that can be downloaded from\ our\ download server as decodeSv.bb. Individual regions or the whole\ annotation can be obtained with the bigBedToBed utility, available\ from our\ utilities\ page. Example:\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/decodeSv.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\\ The original VCF is available from the deCODE genetics\ LRS_SV_sets\ GitHub repository.\
\ \\ Thanks to the deCODE genetics team and the Icelandic study participants for\ making this dataset publicly available.\
\ \\ Beyter D, Ingimundardottir H, Oddsson A, Eggertsson HP, Bjornsson E, Jonsson H, Atlason BA,\ Kristmundsdottir S, Mehringer S, Hardarson MT et al.\ \ Long-read sequencing of 3,622 Icelanders provides insight into the role of structural variants in\ human diseases and other traits.\ Nat Genet. 2021 Jun;53(6):779-786.\ PMID: 33972781\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/decodeSv.bb\ filter.insLen 0:22130\ filter.svLen 0:861080\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.insLen Insertion Length\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterType.svType multipleListOr\ filterValues.svType DEL,INS,INSDEL\ itemRgb on\ longLabel High-confidence Structural Variants from 3,622 Icelanders (deCODE, Oxford Nanopore)\ mouseOver Var: $name ($svType)\ The "Prediction Scores" container track contains subtracks showing the results of variant impact prediction\ scores. Usually these are prediction algorithms that use protein features, conservation, nucleotide composition and similar\ signals to determine if a genome variant is pathogenic or not.
\ \BayesDel is a deleteriousness meta-score for coding and \ non-coding variants, single nucleotide\ variants, and small insertion/deletions. The range of the score is from -1.29334 to 0.75731.\ The higher the score, the more likely the variant is pathogenic.
\\ MaxAF stands for maximum allele frequency. The old ACMG (American College of Medical Genetics and\ Genomics) rules utilize allele frequency to classify variants, so the "BayesDel without MaxAF"\ tracks were created to avoid double-dipping. However, new ACMG rules will not include allele\ frequency, so it is okay to use the "BayesDel with MaxAF" for variant classification in the future.\ For gene discovery research, it is better to use BayesDel with MaxAF.
\\ For gene discovery research, a universal cutoff value (0.0692655 with MaxAF, -0.0570105 without\ MaxAF) was obtained by maximizing sensitivity and specificity in classifying ClinVar variants;\ Version 1 (build date 2017-08-24).
\\ For clinical variant classification, Bayesdel thresholds have been calculated for a variant to\ reach various levels of evidence; please refer to Pejaver et al. 2022 for general application\ of these scores in clinical applications.\
\ \\ Interpretation: The authors define that at an M-CAP score > 0.025, 5% of \ pathogenic variants are misclassified as benign. 0.025 is the recommended cutoff.\
\ \\ The Mendelian Clinically Applicable Pathogenicity (M-CAP)\ score (Jagadeesh et al, Nat Genetics 2016) is a\ pathogenicity likelihood score that aims to misclassify no more than 5% of\ pathogenic variants while aggressively reducing the list of variants of\ uncertain significance. Much like allele frequency, M-CAP is readily\ interpreted; if it classifies a variant as benign, then that variant can be\ trusted to be benign with high confidence.
\ \\ At an M-CAP score > 0.025, 5% of pathogenic variants are misclassified as benign.\ The score varies from 0.0 - 1.0, following a geometric distribution with a mean of 0.09.\
\ \\ Interpretation: The authors defined the thresholds <0.140 for a variant\ to be benign, and > 0.730 for pathogenic with 95% confidence.
\\ The within-gene clustering of pathogenic and benign DNA changes is an important\ feature of the human exome.\ MutScore\ score (Quinodoz, AJHG 2022) integrates qualitative features of\ DNA substitutions with new additional information derived from \ positional clustering. Variants of unknown significance that are scored\ as benign by other algorithms but located close to known pathogenic variants\ should be weighted more pathogenic by MutScore. The score ranges from 0.0-1.0, resembles\ a negative binomial distribution with a maximum ~0.05, depending on the nucleotide.\ MutScore was seen to outperform other scores by papers Porretta et al and Brock et al.\
\ \\ Interpretation: Scores range from 0 to 1, with higher values indicating greater\ predicted pathogenicity. The authors suggest a clinical threshold of 0.821 for distinguishing\ pathogenic from benign missense variants. 75% of all possible missense variants are classified\ as benign, 25% as pathogenic.\
\\ PrimateAI-3D\ (Gao et al, Science 2023) is a semi-supervised 3D convolutional neural network trained on\ 4.5 million benign missense variants from 233 primate species and common human variants.\ It operates on voxelized protein structures at 2 Å resolution (from AlphaFold or\ homology models) combined with multiple sequence alignments from 592 species. The track\ contains pre-computed scores for all 70.7 million possible single nucleotide missense\ variants.\ Pathogenic variants are shown in red,\ benign in blue.\ Items can be filtered by prediction and by percentile score.\
\ \\ Interpretation: Scores range from -1 to 1. Positive scores indicate predicted\ disruption of promoter function, negative scores indicate the variant is tolerated.\
\\ PromoterAI\ predicts the impact of single nucleotide variants in gene\ promoter regions, scoring all possible substitutions within 500 bp of annotated\ transcription start sites. The track contains four bigWig subtracks (one per alternate\ allele) covering 39.5 million positions, plus a bigBed track for the 3.8% of positions\ where overlapping transcripts produce different scores.\
\ \\ Interpretation: Scores range from 0 to 1, with higher values indicating greater\ predicted likelihood of pathogenicity. The authors recommend a threshold of ≥ 0.5 to\ flag variants as likely disease-relevant.\
\\ ClinPred\ (Alirezaie et al, AJHG 2018) is a machine-learning predictor for nonsynonymous\ (missense) single-nucleotide variants. It combines existing pathogenicity scores\ with population allele frequency from gnomAD, and was trained on confidently\ annotated disease-causing and benign variants from ClinVar. The track contains\ four bigWig subtracks (one per alternate allele) with pre-computed scores for\ all possible human missense variants in the exome.\ Pathogenic variants are shown in red,\ benign in blue.\
\ \\ Interpretation: EVE scores range from 0 (benign) to 1 (pathogenic) and are\ normalized within each protein, so they are not directly comparable across proteins. A\ Class25 label assigns each variant to benign, uncertain, or pathogenic using a 25%\ uncertainty threshold.\
\\ EVE\ (Frazer et al, Nature 2021) is a deep generative model (a Bayesian variational\ autoencoder) trained per protein on evolutionary sequence alignments, without using\ clinical labels. The track shows scores for all possible missense substitutions in\ 2,949 disease-associated proteins as a heatmap (rows = amino acids, columns = protein\ positions), colored from benign (blue) through\ uncertain (white) to pathogenic (red).\
\ \\ Interpretation: popEVE scores are a continuous, proteome-wide measure of\ deleteriousness and, unlike most missense scores, are calibrated to be comparable across\ genes; lower (more negative) scores are more deleterious. The authors define a\ high-confidence severe threshold at −5.056 and a moderate threshold at −4.617.\
\\ popEVE\ (Orenbuch et al, Nature Genetics 2025) builds on EVE and the ESM-1v protein language model,\ calibrating their scores against human population variation (UK Biobank) with a Gaussian\ process to place variants across the whole proteome on a single scale. The track shows\ scores for all single-nucleotide-reachable missense substitutions across roughly 18,000\ proteins as a heatmap, colored on a global gradient from\ deleterious (red) to\ tolerated (blue).\
\ \There are eight subtracks for the BayesDel track: four include pre-computed MaxAF-integrated BayesDel\ scores for missense variants, one for each base. The other four are of the same format, but scores\ are not MaxAF-integrated.
\ \For SNVs, at each genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing the reference allele,\ (e.g. A to A) is always set to zero.
\ \Note: There are cases in which a genomic position will have one value missing.\
\ \When using this track, zoom in until you can see every base pair at the top of the display.\ Otherwise, there are several nucleotides per pixel under your mouse cursor and instead of an actual\ score, the tooltip text will show the average score of all nucleotides under the cursor. This is\ indicated by the prefix "~" in the mouseover.\
\ \\
Details on suggested ranges for BayesDel can be found in Bergquist et al Genet Med 2025, Table 2:\
\
There are four subtracks: one for each nucleotide.
\ \There are four subtracks: one for each alternate nucleotide. Each shows the\ ClinPred score for variants from the reference base to that nucleotide. Reference\ and synonymous alternates are set to 0; positions with no exome coverage appear as\ gaps. The track is colored at each position by the recommended threshold\ (≥ 0.5 = pathogenic, < 0.5 = benign).
\ \A single bigBed track containing all possible missense variants. Items are\ colored by prediction (red = pathogenic, blue = benign) and can be filtered by\ prediction or percentile score. See the per-track description page for details.
\ \Four bigWig subtracks (one per alternate nucleotide) covering positions within\ 500 bp of annotated transcription start sites, plus a bigBed track for\ positions where overlapping transcripts produce different scores.
\ \Each is a single bigBed track displayed as a heatmap: one column per amino acid\ position (placed at the codon's genomic coordinate) and one row per amino acid. Hover\ over a cell to see the substitution and its score. EVE is colored per protein from blue\ (benign) to red (pathogenic); popEVE uses a single global gradient (red = deleterious,\ blue = tolerated) so that cells are comparable across genes. These tracks are best viewed\ zoomed in to a single gene or exon.
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server, there is one subdirectory per score.\ The files for this track are called usually called by their alternate allele, e.g. mcapA.bw and mutScoreA.bw. Individual\ regions or the whole genome annotation can be obtained using our tool bigWigToBedGraph\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool\ can also be used to obtain only features within a given range, e.g. \ bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/hg19/mcap/mcapA.bw -chrom=chr21 -start=0 -end=100000000 stdout
\ \ \The original BayesDel files are available at the\ BayesDel website.\
The other algorithms also have their own download formats, on the\ M-CAP website and the MutScore Website.\ \
BayesDel data was converted from the files provided on the\ BayesDel_170824 Database.\ The number 170824 is the date (2017-08-24) the scores were created. Both sets of BayesDel scores are\ available in this database, one integrated MaxAF (named BayesDel_170824_addAF) and one without\ (named BayesDel_170824_noAF). Data conversion was performed using\ \ custom Python scripts.\
\ \M-CAP data was converted using a custom Python script and converted to\ bigWig, as documented in the our makeDoc\ text file. MutScore was already available in bigWig format to download.
\ \Thanks to the BayesDel, MutScore, M-CAP, ClinPred, PrimateAI-3D and PromoterAI teams for\ providing precomputed data, and to Tiana Pereira, Christopher Lee, Gerardo Perez, and Anna\ Benet-Pages of the Genome Browser team.
\ \\ Alirezaie N, Kernohan KD, Hartley T, Majewski J, Hocking TD.\ \ ClinPred: Prediction Tool to Identify Disease-Relevant Nonsynonymous Single-Nucleotide Variants.\ Am J Hum Genet. 2018 Oct 4;103(4):474-483.\ PMID: 30220433; PMC: PMC6174354\
\ \\ Bergquist T, Stenton SL, Nadeau EAW, Byrne AB, Greenblatt MS, Harrison SM, Tavtigian SV,\ O'Donnell-Luria A, Biesecker LG, Radivojac P et al.\ \ Calibration of additional computational tools expands ClinGen recommendation options for variant\ classification with PP3/BP4 criteria.\ Genet Med. 2025 Mar 10;27(6):101402.\ PMID: 40084623\
\ \\ Feng BJ.\ \ PERCH: A Unified Framework for Disease Gene Prioritization.\ Hum Mutat. 2017 Mar;38(3):243-251.\ PMID: 27995669; PMC: PMC5299048\
\ \\ Gao H, Hamp T, Ede J, Schraiber JG, McRae J, Singer-Berk M, Yang Y, Dietrich ASD,\ Fiziev PP, Kuderna LFK et al.\ \ The landscape of tolerated genetic variation in humans and primates.\ Science. 2023 Jun 2;380(6648):eabn8197.\ PMID: 37262156; PMC: PMC10187174\
\ \\ Jagadeesh KA, Wenger AM, Berger MJ, Guturu H, Stenson PD, Cooper DN, Bernstein JA, Bejerano G.\ \ M-CAP eliminates a majority of variants of uncertain significance in clinical exomes at high\ sensitivity.\ Nat Genet. 2016 Dec;48(12):1581-1586.\ PMID: 27776117\
\ \\ Pejaver V, Byrne AB, Feng BJ, Pagel KA, Mooney SD, Karchin R, O'Donnell-Luria A, Harrison SM,\ Tavtigian SV, Greenblatt MS et al.\ \ Calibration of computational tools for missense variant pathogenicity classification and ClinGen\ recommendations for PP3/BP4 criteria.\ Am J Hum Genet. 2022 Dec 1;109(12):2163-2177.\ PMID: 36413997; PMC: PMC9748256\
\ \\ Quinodoz M, Peter VG, Cisarova K, Royer-Bertrand B, Stenson PD, Cooper DN, Unger S, Superti-Furga A,\ Rivolta C.\ \ Analysis of missense variants in the human genome reveals widespread gene-specific clustering and\ improves prediction of pathogenicity.\ Am J Hum Genet. 2022 Mar 3;109(3):457-470.\ PMID: 35120630; PMC: PMC8948164\
\ \\ Sundaram L, Gao H, Padigepati SR, McRae JF, Li Y, Kosmicki JA, Fritzilas N, Hakenberg J,\ Dutta A, Shon J et al.\ \ Predicting the clinical impact of human mutation with deep neural networks.\ Nat Genet. 2018 Aug;50(8):1161-1170.\ PMID: 30038395; PMC: PMC6237276\
\ \\ Tian Y, Pesaran T, Chamberlin A, Fenwick RB, Li S, Gau CL, Chao EC, Lu HM, Black MH, Qian D.\ \ REVEL and BayesDel outperform other in silico meta-predictors for clinical variant\ classification.\ Sci Rep. 2019 Sep 4;9(1):12752.\ PMID: 31484976; PMC: PMC6726608\
\ \ phenDis 0 group phenDis\ longLabel Variant Deleteriousness / Variant Impact Prediction Scores\ pennantIcon Updated red ../goldenPath/newsarch.html#050126 "Two new tracks added May 1, 2026: PrimateAI-3D and PromoterAI"\ shortLabel Deleteriousness Predictions\ superTrack on hide\ track predictionScoresSuper\ visibility hide\ cnvDevDelay Development Delay gvf Copy Number Variation Morbidity Map of Developmental Delay 0 100 0 0 0 127 127 127 0 0 0\ Enrichment of large copy number variants (CNVs) has been linked to severe pediatric disease\ including developmental delay, intellectual disability and autism spectrum disorder. The\ association of individual loci with specific disorders, however, has still been problematic.\
\ \\ This track shows CNVs from cases of developmental delay along with healthy control sets from two\ separate studies. The study by Cooper et al. (2011) analyzed samples from 15,767 children\ with various developmental disabilities and compared them with samples from 8,329 adult controls to\ produce a detailed genome-wide morbidity map of developmental delay and congenital birth defects.\ The study by Coe et al. (2014) further expanded the morbidity map by analyzing 13,318 new\ case samples along with 11,255 new controls.\
\ \\ This is a composite track consisting of a Case subtrack and a Control subtrack. To turn a subtrack\ on or off, toggle the checkbox to the left of the subtrack name in the track controls at the top of\ the track description page.\
\ \\ Items in this track are colored red for copy number loss and\ blue for copy number gain.\
\ \\ The samples were analyzed using nine different CGH platforms with initial CNV calls filtered as\ described in Coe et al. (2014).\
\ \\ Final CNV calls were decoupled from identifying information and submitted to dbVar as\ nstd54 and\ nstd100\ for unrestricted release.\
\ \\ The 15,767 case individuals from the Cooper study comprise nstd54 sampleset 1, while the 8,329\ control individuals from the Cooper study comprise nstd54 samplesets 2-12. The 13,318 case\ individuals from the Coe study were combined with the Cooper case individuals to comprise nstd100\ sampleset 1. The 11,255 control individuals from the Coe study comprise nsdt100 samplesets 2 and 3.\
\ \\ The Case subtrack was constructed using nstd100 sampleset 1. The Control subtrack was constructed by\ combining nstd100 samplesets 2 and 3 with nstd54 samplesets 2-12.\
\ \\ We would like to thank Gregory Cooper, Brad Coe and the\ Eichler Lab at the University of\ Washington for providing the data for this track.\
\ \\ Coe BP, Witherspoon K, Rosenfeld JA, van Bon BW, Vulto-van Silfhout AT, Bosco P, Friend KL, Baker C,\ Buono S, Vissers LE et al.\ \ Refining analyses of copy number variation identifies specific genes associated with developmental\ delay.\ Nat Genet. 2014 Oct;46(10):1063-71.\ PMID: 25217958; PMC: PMC4177294\
\ \\ Cooper GM, Coe BP, Girirajan S, Rosenfeld JA, Vu TH, Baker C, Williams C, Stalker H, Hamid R, Hannig\ V et al.\ \ A copy number variation morbidity map of developmental delay.\ Nat Genet. 2011 Aug 14;43(9):838-46.\ PMID: 21841781; PMC: PMC3171215\
\ phenDis 1 compositeTrack on\ group phenDis\ longLabel Copy Number Variation Morbidity Map of Developmental Delay\ noScoreFilter .\ shortLabel Development Delay\ track cnvDevDelay\ type gvf\ visibility hide\ dgvGold DGV Gold Standard bigBed 12 + Database of Genomic Variants: Gold Standard Variants 0 100 0 0 0 127 127 127 0 0 0 http://dgv.tcag.ca/gb2/gbrowse_details/dgv2_hg38?ref=$S;start=${;end=$};name=$$;class=Sequence varRep 1 bigDataUrl /gbdb/hg38/dgv/dgvGold.bb\ longLabel Database of Genomic Variants: Gold Standard Variants\ mouseOver ID: $name\ This track displays copy number variants (CNVs), insertions/deletions (InDels),\ inversions and inversion breakpoints annotated by the\ Database of Genomic Variants (DGV), which\ contains genomic variations observed in healthy individuals.\ DGV focuses on structural variation, defined as\ genomic alterations that involve segments of DNA that are larger than\ 1000 bp. Insertions/deletions of 50 bp or larger are also included.\
\ \\ This track contains three subtracks:\
\
\ Color is used in both subtracks to indicate the type of variation:\
\ The DGV Gold Standard subtrack utilizes a boxplot-like display to represent the \ merging of records as explained in the Methods section below. In this track, the \ middle box (where applicable), represents the high confidence location of the CNV, \ while the thin lines and end boxes represent the possible range of the CNV.\
\\ Clicking on a variant leads to a page with detailed information about the variant, \ such as the study reference and PubMed abstract link, the study's method and any\ genes overlapping the variant. Also listed, if available, are the sequencing or array platform\ used for the study, a sample cohort description, sample size, sample ID(s) in which\ the variant was observed, observed gains and observed losses.\ If the particular variant is a merged variant, links to genome browser views of \ the supporting variants are listed. If the particular variant is a supporting variant,\ a link to the genome browser view of its merged variant is displayed.\ A link to DGV's Variant Details page for each variant is also provided.\
\\ For most variants, DGV uses accessions from peer archives of structural variation\ (dbVar\ at NCBI or DGVa at EBI).\ These accessions begin with either "essv",\ "esv", "nssv", or "nsv", followed by a number.\ Variant submissions processed by EBI begin with "e"\ and those processed by NCBI begin with "n".\
\\ Accessions with ssv are for variant calls on a particular sample, and if they\ are copy number variants, they generally indicate whether the change is a gain\ or loss. \ In a few studies the ssv represents the variant called by a single\ algorithm. If multiple algorithms were used, overlapping ssv's from\ the same individual would be combined to generate a sample level\ sv. \
\\ If there are many samples analyzed in a study, and if there are many\ samples which have the same variant, there will be multiple ssv's with\ the same start and end coordinates.\ These sample level variants are then merged and combined to form a\ representative variant that highlights the common variant found in\ that study. The result is called a structural variant (sv) record.\ Accessions with sv are for regions asserted by submitters to contain\ structural variants, and often span ssv elements for both losses and\ gains. dbVar and DGVa do not record numbers of losses and gains\ encompassed within sv regions.\
\\ DGV merges clusters of variants that share at least 70% reciprocal\ overlap in size/location, and assigns an accession beginning with\ "dgv", followed by an internal variant serial number,\ followed by an abbreviated study id. For example,\ the first merged variant from the Shaikh et al. 2009 study (study\ accession=nstd21) would be dgv1n21. The second merged variant would be\ dgv2n21 and so forth.\ Since in this case there is an additional level of clustering,\ it is possible for an "sv" variant to be both a merged\ variant and a supporting variant.\
\\ For most sv and dgv variants, DGV displays the total number of\ sample-level gains and/or losses at the bottom of their variant detail\ page. Since each ssv variant is for one sample, its total is 1.\
\ \\ Published structural variants are imported from peer archives\ dbVar and\ DGVa.\ DGV then applies quality filters and merges overlapping variants.\
\\ For data sets where the variation calls are reported at a\ sample-by-sample level, DGV merges calls with similar boundaries\ across the sample\ set. Only variants of the same type (i.e. CNVs, Indels, inversions)\ are merged, and gains and losses are merged separately.\ Sample level calls that overlap by ≥ 70% are merged in this\ process.\
\\ The initial criteria for the Gold Standard set require that a variant \ is found in at least two different studies and found in at least two different \ samples. After filtering out low-quality variants, the remaining variants are \ clustered according to 50% minimum overlap, and then merged into a single \ record. Gains and losses are merged separately.
\\ The highest ranking variant in the cluster defines the inner box, while the \ outer lines define the maximum possible start and stop coordinates of the CNV. \ In this way, the inner box forms a high-confidence CNV location and the \ thin connecting lines indicate confidence intervals for the location of CNV.
\ \\
The raw data can be explored interactively with the Table Browser, or\
the Data Integrator. For automated access, this track, like all\
others, is available via our API. However, for bulk\
processing, it is recommended to download the dataset. The genome annotation is stored in a bigBed\
file that can be downloaded from the\
download server.\
The exact filenames can be found in the track configuration file. Annotations can be converted to\
ASCII text by our tool bigBedToBed which can be compiled from the source code or\
downloaded as a precompiled binary for your system. Instructions for downloading source code and\
binaries can be found\
here. The tool can\
also be used to obtain only features within a given range, for example:
\ bigBedToBed https://hgdownload.soe.ucsc.edu/gbdb/hg38/dgv/dgvMerged.bb -chrom=chr6 -start=0 -end=1000000 stdout\\ \
\ Thanks to the Database of Genomic Variants for providing these data.\ In citing the Database of Genomic Variants please refer to MacDonald\ et al.\
\ \\ Iafrate AJ, Feuk L, Rivera MN, Listewnik ML, Donahoe PK, Qi Y, Scherer SW, Lee C.\ \ Detection of large-scale variation in the human genome.\ Nat Genet. 2004 Sep;36(9):949-51.\ PMID: 15286789\
\ \\ MacDonald JR, Ziman R, Yuen RK, Feuk L, Scherer SW.\ \ The Database of Genomic Variants: a curated collection of structural variation in the human\ genome.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D986-92.\ PMID: 24174537; PMC: PMC3965079\
\ \\ Zhang J, Feuk L, Duggan GE, Khaja R, Scherer SW.\ \ Development of bioinformatics resources for display and analysis of copy number and other structural\ variants in the human genome.\ Cytogenet Genome Res. 2006;115(3-4):205-14.\ PMID: 17124402\
\ \ varRep 1 compositeTrack on\ coriellUrlBase http://ccr.coriell.org/Sections/Search/Sample_Detail.aspx?Ref=\ dataVersion 2020-02-25\ exonArrows off\ exonNumbers off\ group varRep\ itemRgb on\ longLabel Database of Genomic Variants: Structural Variation (CNV, Inversion, In/del)\ noScoreFilter .\ shortLabel DGV Struct Var\ track dgvPlus\ type bed 9 +\ url http://dgv.tcag.ca/dgv/app/variant?id=$$&ref=$D\ urlLabel DGV Browser and Report:\ visibility hide\ dosageSensitivity Dosage Sensitivity bigBed 9 + 2 pHaplo and pTriplo dosage sensitivity map from Collins et al 2022 0 100 0 0 0 127 127 127 0 0 0\ This container track represents dosage sensitivity map data from Collins et al 2022. There are\ two tracks, one corresponding to the probability of haploinsufficiency (pHaplo) and \ one to the probability of triplosensitivity (pTriplo).
\\ Rare copy-number variants (rCNVs) include deletions and duplications that occur \ infrequently in the global human population and can confer substantial risk for \ disease. Collins et al aimed to quantify the properties of haploinsufficiency (i.e., \ deletion intolerance) and triplosensitivity (i.e., duplication intolerance) throughout \ the human genome by analyzing rCNVs from nearly one million individuals to construct a \ genome-wide catalog of dosage sensitivity across 54 disorders, which defined 163 dosage \ sensitive segments associated with at least one disorder. These segments were typically \ gene-dense and often harbored dominant dosage sensitive driver genes. An ensemble \ machine learning model was built to predict dosage sensitivity probabilities (pHaplo & \ pTriplo) for all autosomal genes, which identified 2,987 haploinsufficient and 1,559 \ triplosensitive genes, including 648 that were uniquely triplosensitive.\
\ \\ Each of the tracks is displayed with a distinct item (bed track) covering the entire gene locus wherever \ a score was available. Clicking on an item provides a link to DECIPHER which contains the sensitivity scores as well as\ additional information. Mousing over the items will display the gene symbol, the ESNG ID for that gene, \ and the respective sensitivity score for the track rounded to two decimal places. Filters are \ also available to specify specific score thresholds to display for each of the tracks.
\ \\
\ Each of the tracks is colored based on standardized cutoffs for pHaplo and pTriplo as described by the\ authors:
\\ pHaplo scores ≥0.86 indicate that the average effect sizes of deletions are as strong as \ the loss-of-function of genes known to be constrained against protein truncating variants (average OR≥2.7)\ (Karczewski et al., 2020). \ pHaplo scores ≥0.55 indicate an odds ratio ≥2.
\\ pTriplo scores ≥0.94 indicate that the average effect sizes of deletions are as strong as\ the loss-of-function of genes known to be constrained against protein truncating variants (average OR≥2.7)\ (Karczewski et al., 2020).\ pHaplo scores ≥0.68 indicate an odds ratio ≥2.
\\ Applying these cutoffs defined 2,987 haploinsufficient (pHaplo≥0.86) and 1,559\ triplosensitive (pTriplo≥0.94) genes with rCNV effect sizes comparable to loss-of-function\ of gold-standard PTV-constrained genes.
\\
See below for a summary of the color scheme:
\ \\ The data were downloaded from Zenodo which consisted of a 3-column file with\ gene symbols, pHaplo, and pTriplo scores. Since the data were created using\ GENCODEv19 models, the hg19 data was mapped using those coordinates by picking the earliest\ transcription start site of all of the respective gene transcripts and the furthest \ transcription end site. This leads to some gene boundaries that are not representative of a real\ transcript, but since the data are for gene loci annotations this maximum coverage was used.\ Finally, both scores were rounded to two decimal points for easier interpretation.
\\ For hg38, we attempted to use updated gene positions using a few different datasets since \ gene symbols have been updated many times since GENCODEv19. A summary of the workflow\ can be seen below, with each subsequent step being used only for genes where mapping failed:
\\ In summary, the hg19 track was mapped using the original GENCODEv19 mappings, and a series\ of steps were taken to map the hg38 gene symbols with updated coordinates. 19/18641 items\ could not be mapped and are missing from the hg38 tracks.
\\ The complete \ makeDoc can be found online. This includes all of the track creation steps.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tool \
bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/dosageSensitivityCollins2022/pHaploDosageSensitivity.bb stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ Thanks to DECIPHER for their support and assistance with the data. We would also like to \ thank Anna Benet-Pagès for suggesting and assisting in track development and interpretation.\
\ \\ Collins RL, Glessner JT, Porcu E, Lepamets M, Brandon R, Lauricella C, Han L, Morley T, Niestroj LM,\ Ulirsch J et al.\ \ A cross-disorder dosage sensitivity map of the human genome.\ Cell. 2022 Aug 4;185(16):3041-3055.e25.\ PMID: 35917817; PMC: PMC9742861\
\ phenDis 1 compositeTrack on\ group phenDis\ html dosageSensitivityCollins2022\ itemRgb on\ longLabel pHaplo and pTriplo dosage sensitivity map from Collins et al 2022\ noParentConfig on\ shortLabel Dosage Sensitivity\ track dosageSensitivity\ type bigBed 9 + 2\ visibility hide\ cons470wayViewphastcons Element Conservation (phastCons) bed 4 Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group, 0 100 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group,\ parent cons470way\ shortLabel Element Conservation (phastCons)\ track cons470wayViewphastcons\ view phastcons\ visibility hide\ embryo_brain_models Emb Brain models bigBed 12 + Embryonic Brain transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-EmbBrain.bb\ longLabel Embryonic Brain transcript models\ parent sample_models_view on\ shortLabel Emb Brain models\ subGroups view=sample_models_view sample=embryo_brain type=models\ track embryo_brain_models\ type bigBed 12 +\ visibility squish\ embryo_brain_ont_post_models Emb Brain ONT post models bigBed 12 + Embryonic Brain ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_EmbBrain01Rep1.bb\ itemRgb on\ longLabel Embryonic Brain ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Brain ONT post models\ subGroups view=per_expr_models_view sample=embryo_brain type=post_capture_ont_models\ track embryo_brain_ont_post_models\ type bigBed 12 +\ visibility hide\ embryo_brain_ont_post_reads Emb Brain ONT post reads bam Embryonic Brain ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_EmbBrain01Rep1.bam\ longLabel Embryonic Brain ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Brain ONT post reads\ subGroups view=per_expr_reads_view sample=embryo_brain type=post_capture_ont_reads\ track embryo_brain_ont_post_reads\ type bam\ visibility hide\ embryo_brain_ont_pre_models Emb Brain ONT pre models bigBed 12 + Embryonic Brain ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_EmbBrain01Rep1.bb\ itemRgb on\ longLabel Embryonic Brain ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Brain ONT pre models\ subGroups view=per_expr_models_view sample=embryo_brain type=pre_capture_ont_models\ track embryo_brain_ont_pre_models\ type bigBed 12 +\ visibility hide\ embryo_brain_ont_pre_reads Emb Brain ONT pre reads bam Embryonic Brain ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_EmbBrain01Rep1.bam\ longLabel Embryonic Brain ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Brain ONT pre reads\ subGroups view=per_expr_reads_view sample=embryo_brain type=pre_capture_ont_reads\ track embryo_brain_ont_pre_reads\ type bam\ visibility hide\ embryo_brain_pacbio_post_models Emb Brain PB post models bigBed 12 + Embryonic Brain PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_EmbBrain01Rep1.bb\ itemRgb on\ longLabel Embryonic Brain PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Brain PB post models\ subGroups view=per_expr_models_view sample=embryo_brain type=post_capture_pacbio_models\ track embryo_brain_pacbio_post_models\ type bigBed 12 +\ visibility hide\ embryo_brain_pacbio_post_reads Emb Brain PB post reads bam Embryonic Brain PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_EmbBrain01Rep1.bam\ longLabel Embryonic Brain PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Brain PB post reads\ subGroups view=per_expr_reads_view sample=embryo_brain type=post_capture_pacbio_reads\ track embryo_brain_pacbio_post_reads\ type bam\ visibility hide\ embryo_brain_pacbio_pre_models Emb Brain PB pre models bigBed 12 + Embryonic Brain PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_EmbBrain01Rep1.bb\ itemRgb on\ longLabel Embryonic Brain PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Brain PB pre models\ subGroups view=per_expr_models_view sample=embryo_brain type=pre_capture_pacbio_models\ track embryo_brain_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ embryo_brain_pacbio_pre_reads Emb Brain PB pre reads bam Embryonic Brain PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_EmbBrain01Rep1.bam\ longLabel Embryonic Brain PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Brain PB pre reads\ subGroups view=per_expr_reads_view sample=embryo_brain type=pre_capture_pacbio_reads\ track embryo_brain_pacbio_pre_reads\ type bam\ visibility hide\ embryo_heart_models Emb Heart models bigBed 12 + Embryonic Heart transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-EmbHeart.bb\ longLabel Embryonic Heart transcript models\ parent sample_models_view on\ shortLabel Emb Heart models\ subGroups view=sample_models_view sample=embryo_heart type=models\ track embryo_heart_models\ type bigBed 12 +\ visibility squish\ embryo_heart_ont_post_models Emb Heart ONT post models bigBed 12 + Embryonic Heart ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_EmbHeart01Rep1.bb\ itemRgb on\ longLabel Embryonic Heart ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Heart ONT post models\ subGroups view=per_expr_models_view sample=embryo_heart type=post_capture_ont_models\ track embryo_heart_ont_post_models\ type bigBed 12 +\ visibility hide\ embryo_heart_ont_post_reads Emb Heart ONT post reads bam Embryonic Heart ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_EmbHeart01Rep1.bam\ longLabel Embryonic Heart ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Heart ONT post reads\ subGroups view=per_expr_reads_view sample=embryo_heart type=post_capture_ont_reads\ track embryo_heart_ont_post_reads\ type bam\ visibility hide\ embryo_heart_ont_pre_models Emb Heart ONT pre models bigBed 12 + Embryonic Heart ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_EmbHeart01Rep1.bb\ itemRgb on\ longLabel Embryonic Heart ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Heart ONT pre models\ subGroups view=per_expr_models_view sample=embryo_heart type=pre_capture_ont_models\ track embryo_heart_ont_pre_models\ type bigBed 12 +\ visibility hide\ embryo_heart_ont_pre_reads Emb Heart ONT pre reads bam Embryonic Heart ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_EmbHeart01Rep1.bam\ longLabel Embryonic Heart ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Heart ONT pre reads\ subGroups view=per_expr_reads_view sample=embryo_heart type=pre_capture_ont_reads\ track embryo_heart_ont_pre_reads\ type bam\ visibility hide\ embryo_heart_pacbio_post_models Emb Heart PB post models bigBed 12 + Embryonic Heart PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_EmbHeart01Rep1.bb\ itemRgb on\ longLabel Embryonic Heart PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Heart PB post models\ subGroups view=per_expr_models_view sample=embryo_heart type=post_capture_pacbio_models\ track embryo_heart_pacbio_post_models\ type bigBed 12 +\ visibility hide\ embryo_heart_pacbio_post_reads Emb Heart PB post reads bam Embryonic Heart PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_EmbHeart01Rep1.bam\ longLabel Embryonic Heart PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Heart PB post reads\ subGroups view=per_expr_reads_view sample=embryo_heart type=post_capture_pacbio_reads\ track embryo_heart_pacbio_post_reads\ type bam\ visibility hide\ embryo_heart_pacbio_pre_models Emb Heart PB pre models bigBed 12 + Embryonic Heart PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_EmbHeart01Rep1.bb\ itemRgb on\ longLabel Embryonic Heart PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Heart PB pre models\ subGroups view=per_expr_models_view sample=embryo_heart type=pre_capture_pacbio_models\ track embryo_heart_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ embryo_heart_pacbio_pre_reads Emb Heart PB pre reads bam Embryonic Heart PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_EmbHeart01Rep1.bam\ longLabel Embryonic Heart PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Heart PB pre reads\ subGroups view=per_expr_reads_view sample=embryo_heart type=pre_capture_pacbio_reads\ track embryo_heart_pacbio_pre_reads\ type bam\ visibility hide\ embryo_ipsc_models Emb iPSC models bigBed 12 + Embryonic iPSC transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-EmbiPSC.bb\ longLabel Embryonic iPSC transcript models\ parent sample_models_view on\ shortLabel Emb iPSC models\ subGroups view=sample_models_view sample=embryo_ipsc type=models\ track embryo_ipsc_models\ type bigBed 12 +\ visibility squish\ embryo_ipsc_ont_post_models Emb iPSC ONT post models bigBed 12 + Embryonic iPSC ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_iPSC01Rep1.bb\ itemRgb on\ longLabel Embryonic iPSC ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb iPSC ONT post models\ subGroups view=per_expr_models_view sample=embryo_ipsc type=post_capture_ont_models\ track embryo_ipsc_ont_post_models\ type bigBed 12 +\ visibility hide\ embryo_ipsc_ont_post_reads Emb iPSC ONT post reads bam Embryonic iPSC ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_iPSC01Rep1.bam\ longLabel Embryonic iPSC ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb iPSC ONT post reads\ subGroups view=per_expr_reads_view sample=embryo_ipsc type=post_capture_ont_reads\ track embryo_ipsc_ont_post_reads\ type bam\ visibility hide\ embryo_ipsc_ont_pre_models Emb iPSC ONT pre models bigBed 12 + Embryonic iPSC ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_iPSC01Rep1.bb\ itemRgb on\ longLabel Embryonic iPSC ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb iPSC ONT pre models\ subGroups view=per_expr_models_view sample=embryo_ipsc type=pre_capture_ont_models\ track embryo_ipsc_ont_pre_models\ type bigBed 12 +\ visibility hide\ embryo_ipsc_ont_pre_reads Emb iPSC ONT pre reads bam Embryonic iPSC ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_iPSC01Rep1.bam\ longLabel Embryonic iPSC ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb iPSC ONT pre reads\ subGroups view=per_expr_reads_view sample=embryo_ipsc type=pre_capture_ont_reads\ track embryo_ipsc_ont_pre_reads\ type bam\ visibility hide\ embryo_ipsc_pacbio_post_models Emb iPSC PB post models bigBed 12 + Embryonic iPSC PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_iPSC01Rep1.bb\ itemRgb on\ longLabel Embryonic iPSC PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb iPSC PB post models\ subGroups view=per_expr_models_view sample=embryo_ipsc type=post_capture_pacbio_models\ track embryo_ipsc_pacbio_post_models\ type bigBed 12 +\ visibility hide\ embryo_ipsc_pacbio_post_reads Emb iPSC PB post reads bam Embryonic iPSC PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_iPSC01Rep1.bam\ longLabel Embryonic iPSC PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb iPSC PB post reads\ subGroups view=per_expr_reads_view sample=embryo_ipsc type=post_capture_pacbio_reads\ track embryo_ipsc_pacbio_post_reads\ type bam\ visibility hide\ embryo_ipsc_pacbio_pre_models Emb iPSC PB pre models bigBed 12 + Embryonic iPSC PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_iPSC01Rep1.bb\ itemRgb on\ longLabel Embryonic iPSC PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb iPSC PB pre models\ subGroups view=per_expr_models_view sample=embryo_ipsc type=pre_capture_pacbio_models\ track embryo_ipsc_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ embryo_ipsc_pacbio_pre_reads Emb iPSC PB pre reads bam Embryonic iPSC PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_iPSC01Rep1.bam\ longLabel Embryonic iPSC PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb iPSC PB pre reads\ subGroups view=per_expr_reads_view sample=embryo_ipsc type=pre_capture_pacbio_reads\ track embryo_ipsc_pacbio_pre_reads\ type bam\ visibility hide\ embryo_liver_models Emb Liver models bigBed 12 + Embryonic Liver transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-EmbLiver.bb\ longLabel Embryonic Liver transcript models\ parent sample_models_view on\ shortLabel Emb Liver models\ subGroups view=sample_models_view sample=embryo_liver type=models\ track embryo_liver_models\ type bigBed 12 +\ visibility squish\ embryo_liver_ont_post_models Emb Liver ONT post models bigBed 12 + Embryonic Liver ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_EmbLiver01Rep1.bb\ itemRgb on\ longLabel Embryonic Liver ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Liver ONT post models\ subGroups view=per_expr_models_view sample=embryo_liver type=post_capture_ont_models\ track embryo_liver_ont_post_models\ type bigBed 12 +\ visibility hide\ embryo_liver_ont_post_reads Emb Liver ONT post reads bam Embryonic Liver ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_EmbLiver01Rep1.bam\ longLabel Embryonic Liver ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Liver ONT post reads\ subGroups view=per_expr_reads_view sample=embryo_liver type=post_capture_ont_reads\ track embryo_liver_ont_post_reads\ type bam\ visibility hide\ embryo_liver_ont_pre_models Emb Liver ONT pre models bigBed 12 + Embryonic Liver ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_EmbLiver01Rep1.bb\ itemRgb on\ longLabel Embryonic Liver ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Liver ONT pre models\ subGroups view=per_expr_models_view sample=embryo_liver type=pre_capture_ont_models\ track embryo_liver_ont_pre_models\ type bigBed 12 +\ visibility hide\ embryo_liver_ont_pre_reads Emb Liver ONT pre reads bam Embryonic Liver ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_EmbLiver01Rep1.bam\ longLabel Embryonic Liver ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Liver ONT pre reads\ subGroups view=per_expr_reads_view sample=embryo_liver type=pre_capture_ont_reads\ track embryo_liver_ont_pre_reads\ type bam\ visibility hide\ embryo_liver_pacbio_post_models Emb Liver PB post models bigBed 12 + Embryonic Liver PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_EmbLiver01Rep1.bb\ itemRgb on\ longLabel Embryonic Liver PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Liver PB post models\ subGroups view=per_expr_models_view sample=embryo_liver type=post_capture_pacbio_models\ track embryo_liver_pacbio_post_models\ type bigBed 12 +\ visibility hide\ embryo_liver_pacbio_post_reads Emb Liver PB post reads bam Embryonic Liver PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_EmbLiver01Rep1.bam\ longLabel Embryonic Liver PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Liver PB post reads\ subGroups view=per_expr_reads_view sample=embryo_liver type=post_capture_pacbio_reads\ track embryo_liver_pacbio_post_reads\ type bam\ visibility hide\ embryo_liver_pacbio_pre_models Emb Liver PB pre models bigBed 12 + Embryonic Liver PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_EmbLiver01Rep1.bb\ itemRgb on\ longLabel Embryonic Liver PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Emb Liver PB pre models\ subGroups view=per_expr_models_view sample=embryo_liver type=pre_capture_pacbio_models\ track embryo_liver_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ embryo_liver_pacbio_pre_reads Emb Liver PB pre reads bam Embryonic Liver PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_EmbLiver01Rep1.bam\ longLabel Embryonic Liver PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Emb Liver PB pre reads\ subGroups view=per_expr_reads_view sample=embryo_liver type=pre_capture_pacbio_reads\ track embryo_liver_pacbio_pre_reads\ type bam\ visibility hide\ ENCFF431JDU_ENCFF964OOU_ENCFF787LMI_ENCFF388PVO ENCFF431JDU_ENCFF964OOU_ENCFF787LMI_ENCFF388PVO bigBed 9 + 5 HCT116: (1) cCREs 4 100 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF431JDU_ENCFF964OOU_ENCFF787LMI_ENCFF388PVO.bb\ longLabel HCT116: (1) cCREs\ mouseOver ID: ${name}\ The ENCODE4 long-read RNA-seq collection annotates trancripts using numerical triplets representing \ the identity of the start site, exon junction chain, and transcript end site of each transcript. \ This method reveals how promoter selection, splice pattern, and 3’ processing are deployed across \ human tissues.\
\ \\ Transcript names include a triplet annotation that represents transcript start site, exon junction \ chain, and transcript end site. For example, if transcript A has the label [1,2,3] and transcript B\ is labeled [1,1,3], then those transcripts share start and end sites but have a different combination\ of exons. Here is an exmaple drawn from hg38 at the INSIG1 locus:
\ \
\
\
\ In this example, the first two transcripts marked by arrows have the same start\ site ("1") and the same set of exons ("8"), but they have different end sites\ ("2" vs "1"). Similarly, the second two marked transcripts have the same start\ site ("1"), but a different set of exons ("8" vs "9") and a different end site\ ("1" vs "2").\
\ \ \\ GENCODE V29 and V40 were used as reference data; any transcript not present in either of these is\ colored blue.
\\ Mouseover on transcripts shows their ENCODE gene ID and the tissue or cell line where it’s most highly\ expressed and its TPM in that sample.\
\ \\ The data underlying this track is available in the file\ encode4LongRna.bb.\ Individual regions or the whole genome annotation can be obtained using our\ tool bigBedToBed, which is available on our\ download server.\ For example, to extract only annotations in a given region, you could use the following command:\
\ \\ bigBedToBed -chrom=chr1 -start=100000 -end=100500 https://hgdownload.gi.ucsc.edu/gbdb/hg38/encode4LongRna.bb stdout\\ \
\ Please refer to our\ mailing list archives\ for questions, or our\ Data Access FAQ\ for more information.\
\ \\ Data were retrieved from https://zenodo.org/records/15116042.\ The human_ucsc_transcripts.gtf was converted to BED format, and expression and CDS data\ added from the relevant files using a custom script.\
\ \\ Thanks to Fairlie Reese for providing data access and for helpful feedback.\
\ \\ Reese F, Williams B, Balderrama-Gutierrez G, Wyman D, Çelik MH, Rebboah E, Rezaie N, Trout D,\ Razavi-Mohseni M, Jiang Y et al.\ \ The ENCODE4 long-read RNA-seq collection reveals distinct classes of transcript structure\ diversity.\ bioRxiv. 2023 May 16;.\ PMID: 37292896; PMC: PMC10245583\
\ rna 1 bigDataUrl /gbdb/hg38/encode4/encode4LongRnaTranscripts.bb\ defaultLabelFields transcript_name\ filter.maxScore 0:93180.9\ filterByRange.maxScore on\ filterLabel.maxScore Filter by counts per million\ filterLimits.maxScore 0:93180.9\ html encode4LongRnaTranscripts.html\ itemRgb on\ labelFields name,transcript_name\ longLabel ENCODE4 Long Read Transcripts\ mouseOver $nameNOTE:
VarChat is an open platform \
powered by enGenome, and registration is free of charge. VarChat is intended for research\
use and may provide inaccurate answers. It is advisable to verify critical information independently.
\ VarChat is an open platform that leverages \ the power of generative artificial intelligence to support the genomic variant interpretation process \ by searching the available scientific literature for each variant and condensing it into a brief\ yet informative text. Each query quickly scans the latest scientific literature to provide\ up-to-date variant information.\
\ \\ VarChat is a generative AI-based system and each answer is generated live, so you may obtain slightly \ different answers at each iteration. A literature search will be performed and the total number of \ identified publications will be shown. Only a subset of them will be reported and used to generate your answer.\
\ \\ If you would like to stay updated on the latest developments, you may register for updates on the\ VarChat website. For data questions, VarChat\ can be contacted at varchat@engenome.com.
\ \\ Genomic locations of variants are labeled with the nucleotide change.\ Mousing over the items will show how many papers the variant was observed in, its gene,\ its HGVS nomenclature, and dbSNP rsID.\ Clicking on any item will provide a link directly to VarChat with additional information.
\ \\ The items are colored based on the amount of literature support as described on the table below:\
\ \\
| Color | \Level of literature support | \
|---|---|
| \ | High: at least 25 papers mention the variant | \
| \ | Medium: between 10 and 24 papers mention the variant | \
| \ | Low: fewer than 10 papers mention the variant | \
\
VarChat software is powered by enGenome.
\
enGenome, an accredited spin-off from the University of Pavia founded in 2016, combines bioinformatics, biotechnology, and software development expertise to enhance genetic disease diagnosis and treatment through advanced AI and bioinformatics tools, supported by a multidisciplinary team of engineers, biotechnologists, and developers.\
\ For every queried variant, VarChat produces concise and coherent summaries through an LLM model.\ Relevant references are identified through a modified BM25 ranking algorithm. More weight is\ given to papers that cite the variant in the abstract and were published in the last two years,\ while papers that report the variant only in the supplementary are penalized.
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be accessed from scripts through our\ API, the track name is "varChat".
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ The file for this track is called varChat.bb. Individual\ regions or the whole genome annotation can be obtained using our tool bigBedToBed\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system.
\\
Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g.\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/varChat.bb -chrom=chr21 -start=0 -end=10000000 stdout
\ De Paoli F, Berardelli S, Limongelli I, Rizzo E, Zucca S. VarChat: the generative AI\ assistant for the interpretation of human genomic variations. Bioinformatics.\ 2024Mar29;40(4). PMID: 38579245; PMC: PMC11055464
\ phenDis 1 bedNameLabel Nucleotide Change\ bigDataUrl /gbdb/hg38/bbi/varChat.bb\ dataVersion /gbdb/$D/bbi/varChatVersion.txt\ detailsDynamicTable VariantDetails|Variant Details\ exonNumbers off\ html varChat.html\ longLabel enGenome VarChat: Literature match and variant's summary\ maxItems 1000000\ maxWindowCoverage 40000\ mouseOverField _mouseOver\ noScoreFilter on\ parent varsInPubs pack\ shortLabel enGenome VarChat\ skipFields Variant,VariantUrl,score\ track varChat\ type bigBed 9 + 4\ url https://varchat.engenome.com/search?source=ucsc&text=$\ These tracks represent the experimentally validated promoters generated by \ the Eukaryotic Promoter Database.\
\ \\ Each item in the track is a representation of the promoter sequence identified by EPD. The\ "thin" part of the element represents the 49 bp upstream of the annotated transcription\ start site (TSS) whereas the "thick" part represents the TSS plus 10 bp downstream. The\ relative position of the thick and thin parts define the orientation of the promoter.
\\ Note that the EPD team has created a public track hub containing\ promoter and supporting annotations for human, mouse, and other vertebrate and model organism\ genomes.
\ \\ Briefly, gene transcript coordinates were obtained from multiple sources (HGNC, GENCODE, Ensembl,\ RefSeq) and validated using data from CAGE and RAMPAGE experimental studies obtained from FANTOM 5,\ UCSC, and ENCODE. Peak calling, clustering and filtering based on relative expression were applied\ to identify the most expressed promoters and those present in the largest number of samples.
\\ For the methodology and principles used by EPD to predict TSSs, refer to Dreos et al.\ (2013) in the References section below. A more detailed description of how this data was\ generated can be found at the following links:\ \
\ Data was generated by the EPD team at the \ Swiss Institute of Bioinformatics. \ For inquiries, contact the EPD team using this on-line form \ or email \ \ philipp.\ bucher@epfl.\ ch\ \ .\
\ \\ Dreos R, Ambrosini G, Perier RC, Bucher P.\ \ EPD and EPDnew, high-quality promoter resources in the\ next-generation sequencing era. Nucleic Acids\ Res. 2013 Jan 1;41(D1):D157-64. PMID: 23193273.\
\ \ expression 1 bedNameLabel Promoter ID\ compositeTrack on\ exonArrows on\ group expression\ html ../../epdNewPromoter\ longLabel Promoters from EPDnew\ shortLabel EPDnew Promoters\ track epdNew\ type bigBed 8\ urlLabel EPDnew link:\ visibility hide\ gnomADPextEsophagus_GastroesophagealJunction Esophagus-Gastroesophageal Junction bigWig 0 1 gnomAD pext Esophagus-Gastroesophageal Junction 0 100 139 115 85 197 185 170 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Esophagus_GastroesophagealJunction.bw\ color 139,115,85\ longLabel gnomAD pext Esophagus-Gastroesophageal Junction\ parent gnomadPext off\ shortLabel Esophagus-Gastroesophageal Junction\ track gnomADPextEsophagus_GastroesophagealJunction\ visibility hide\ gnomADPextEsophagus_Mucosa Esophagus-Mucosa bigWig 0 1 gnomAD pext Esophagus-Mucosa 0 100 85 34 0 170 144 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Esophagus_Mucosa.bw\ color 85,34,0\ longLabel gnomAD pext Esophagus-Mucosa\ parent gnomadPext off\ shortLabel Esophagus-Mucosa\ track gnomADPextEsophagus_Mucosa\ visibility hide\ gnomADPextEsophagus_Muscularis Esophagus-Muscularis bigWig 0 1 gnomAD pext Esophagus-Muscularis 0 100 187 153 136 221 204 195 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Esophagus_Muscularis.bw\ color 187,153,136\ longLabel gnomAD pext Esophagus-Muscularis\ parent gnomadPext off\ shortLabel Esophagus-Muscularis\ track gnomADPextEsophagus_Muscularis\ visibility hide\ exomeProbesets Exome Probesets bigBed Exome Capture Probesets and Targeted Region 0 100 0 0 0 127 127 127 0 0 0\ This set of tracks shows the genomic positions of probes and targets from a full \ suite of in-solution-capture target enrichment exome kits for Next Generation Sequencing (NGS)\ applications. Also known as exome sequencing or whole exome sequencing (WES), \ this technique allows high-throughput parallel sequencing of all exons (e.g., coding regions of genes \ which affect protein function), constituting about 1% of the human genome, or approximately 30 \ million base pairs.\
\\ The tracks are intended to show the major differences in target genomic regions between the \ different exome capture kits from the major players in the NGS sequencing market:\ Illumina Inc., \ Roche NimbleGen Inc., \ Agilent Technologies Inc.,\ MGI Tech,\ Twist Bioscience, and\ Integrated DNA Technologies Inc..\
\ \\ Items are shaded according to manufacturing company:\
\ Tracks labeled as Probes (P) indicate the footprint of the oligonucleotide probes\ mapped to the human genome. This is the technically relevant targeted region by the assay. However, \ the sequenced region will be bigger than this since flanking sequences are sequenced as well. \ Tracks labeled as Target Regions (T) indicate the genomic regions targeted by the\ assay. This is the biologically relevant target region. Not all targeted regions\ will necessarily be sequenced perfectly; there might be some capture bias at certain locations.\ The Target\ Regions are those normally used for coverage analysis. \
\ \Note that most exome probesets are available on hg19 only. If you are working with hg38 and cannot find\ a particular probeset there, try to go to hg19, configure the same track, and\ see if it exists there. If you cannot find an array, do not hesitate to send us\ an email with the name of the manufacturer website with the probe file. If\ an array is available on hg19 but not on hg38 and you need it for your work, we\ can lift the locations. Our mailing list can be reached at genome@soe.ucsc.edu.\
\ \\ The capture of the genomic regions of interest using in-solution capture, is achieved \ through the hybridization of a set of probes (oligonucleotides) with a sample of fragmented genomic \ DNA in a solution environment. The probes hybridize selectively to the genomic regions of interest \ which, after a process of exclusion of the non-selective DNA material, can be pulled down and \ sequenced, enabling selective DNA sequencing of the genomic regions of interest (e.g., exons).\ In-solution capture sequencing is a sensitive method to detect single nucleotide variants, \ insertions and deletions, and copy number variations.\
\ \ \ \
| Kit | \Targeted Region | \Databases Used for Design | \Year of Release | \
|---|---|---|---|
| IDT - xGen Exome Research Panel V1.0 | \ \ \ \ \39 Mb | \Coding sequences from RefSeq (19,396 genes) | \2015 | \
| IDT - xGen Exome Research Panel V2.0 | \34 Mb | \Coding sequences from RefSeq 109 (19,433 genes) | \2020 | \
| Twist - RefSeq Exome Panel | \3.6 Mb | \Curated subset of protein coding genes from CCDS | \N/A | \
| Twist - Core Exome Panel | \33 Mb | \Protein coding genes from CCDS | \N/A | \
| Twist - Comprehensive Exome Panel | \36.8 Mb | \Protein coding genes from RefSeq, CCDS, and GENCODE | \2020 | \
| Twist - Exome Panel 2.0 | \36.4 Mb | \Protein coding genes from RefSeq, CCDS, and GENCODE | \2021 | \
| MGI - Easy Exome Capture V4 | \59 Mb | \CCDS, GENCODE, RefSeq, and miRBase | \N/A | \
| MGI - Easy Exome Capture V5 | \69 Mb | \CCDS, GENCODE, RefSeq, miRBase, and MGI Clinical Database | \N/A | \
| Agilent - SureSelect Clinical Research Exome | \54 Mb | \Disease-associated regions from OMIM, HGMD, and ClinVar | \2014 | \
| Agilent - SureSelect Clinical Research Exome V2 | \63.7 Mb | \Disease-associated regions from OMIM, HGMD, ClinVar, and ACMG | \2017 | \
| Agilent - SureSelect Focused Exome | \12 Mb | \Disease-associated regions from HGMD, OMIM and ClinVar | \2016 | \
| Agilent - SureSelect All Exon V4 | \51 Mb | \Coding regions from CCDS, RefSeq, and GENCODE v6, miRBase v17, TCGA v6, and UCSC known genes | \2011 | \
| Agilent - SureSelect All Exon V4 + UTRs | \71 Mb | \Coding regions and 5' and 3' UTR sequences from CCDS, RefSeq, and GENCODE v6, regions from miRBase v17, TCGA v6, and UCSC known genes | \2011 | \
| Agilent - SureSelect All Exon V5 | \50 Mb | \Coding regions from Refseq, GENCODE, UCSC, TCGA, CCDS, and miRBase (21.522 genes) | \2012 | \
| Agilent - SureSelect All Exon V5 + UTRs | \74 Mb | \Coding regions and 5' and 3' UTR sequences from Refseq, GENCODE, UCSC, TCGA, CCDS, and miRBase (21.522 genes) | \2012 | \
| Agilent - SureSelect All Exon V6 r2 | \60 Mb | \Coding regions from RefSeq, CCDS, GENCODE, HGMD, and OMIM | \2016 | \
| Agilent - SureSelect All Exon V6 + COSMIC r2 | \66 Mb | \Coding regions from RefSeq, CCDS, GENCODE, HGMD, and OMIM, and targets from both TCGA and COSMIC | \2016 | \
| Agilent - SureSelect All Exon V6 + UTR r2 | \75 Mb | \Coding regions and 5' and 3' UTR sequences from RefSeq, GENCODE, CCDS, and UCSC known genes,and miRNAs and lncRNA sequences | \2016 | \
| Agilent - SureSelect All Exon V7 | \35.7 Mb | \Coding regions from RefSeq, CCDS, GENCODE, and UCSC known genes | \2018 | \
| Roche - KAPA HyperExome | \43Mb | \Coding regions from CCDS, RefSeq, Ensembl, GENCODE,and variants from ClinVar | \2020 | \
| Roche - SeqCap EZ Exome V3 | \64 Mb | \Coding regions from RefSeq RefGene CDS, CCDS, and miRBase v14 databases, plus coverage of 97% Vega, 97% Gencode, and 99% Ensembl | \2018 | \
| Roche - SeqCap EZ Exome V3 + UTR | \92 Mb | \Coding sequences from RefSeq RefGene, CCDS, and miRBase v14, plus coverage of 97% Vega, 97% Gencode, and 99% Ensembl and UTRs from RefSeq RefGene table from UCSC GRCh37/hg19 March 2012 and Ensembl (GRCh37 v64) | \2018 | \
| Roche - SeqCap EZ MedExome | \47 Mb | \Coding sequences from CCDS 17, RefSeq, Ensembl 76, VEGA 56, GENCODE 20, miRBase 21, and disease-associated regions from GeneTests, ClinVar, and based on customer input | \2014 | \
| Roche - SeqCap EZ MedExome + Mito | \47 Mb | \Coding sequences and mitochondrial genes from CCDS 17, RefSeq, Ensembl 76, VEGA 56, GENCODE 20 and miRBase 21, disease-associated regions from GeneTests, ClinVar, and based on customer input | \2014 | \
| Illumina - Nextera DNA Exome V1.2 | \45 Mb | \Coding regions from RefSeq, CCDS, Ensembl, and GENCODE v19 | \2015 | \
| Illumina - Nextera Rapid Capture Exome | \37 Mb | \212,158 targeted exonic regions with start and stop chromosome locations in GRCh37/hg19 | \2013 | \
| Illumina - Nextera Rapid Capture Exome V1.2 | \37 Mb | \Coding regions from RefSeq, CCDS, Ensembl, and GENCODE v12 | \2014 | \
| Illumina - Nextera Rapid Capture Expanded Exome | \66 Mb | \Coding regions from RefSeq, CCDS, Ensembl, and GENCODE v12 | \2013 | \
| Illumina - TruSeq DNA Exome V1.2 | \45 Mb | \Coding regions from RefSeq, CCDS, and Ensembl | \2017 | \
| Illumina - TruSeq Rapid Exome V1.2 | \45 Mb | \Coding regions from RefSeq, CCDS, Ensembl, and GENECODE v19 | \2015 | \
| Illumina - TruSight ONE V1.1 | \12 Mb | \Coding regions of 6700 genes from HGMD, OMIM, and GeneTest | \2017 | \
| Illumina - TruSight Exome | \7 Mb | \Disease-causing mutations as curated by HGMD | \2017 | \
| Illumina - AmpliSeq Exome Panel | \N/A | \CCDS coding regions | \2019 | \
\ The raw data can be explored interactively with the Table Browser\ or cross-referenced with Data Integrator. The data can be\ accessed from scripts through our API, with track names\ found in the Table Schema page for each subtrack after "Primary Table:".\ \
\ For downloading the data, the annotations are stored in bigBed files that\ can be accessed at\ \ our download directory. \ Regional or the whole genome text annotations can be obtained using our utility \ bigBedToBed. Instructions for downloading utilities can be found\ here.\
\ \\ Thanks to Illumina (U.S.), Roche NimbleGen, Inc. (U.S.), Agilent Technologies (U.S.), MGI Tech\ (Beijing Genomics Institute, China), Twist Bioscience (U.S.), and Integrated DNA Technologies (IDT),\ Inc. (U.S.), and Bionano Genomics (U.S.) for making these data available and to Tiana Pereira, Pranav Muthuraman, Began Nguy\ and Anna Benet-Pages for enginering these tracks.\
\ \ \ \ map 1 allButtonPair on\ compositeTrack on\ group map\ longLabel Exome Capture Probesets and Targeted Region\ shortLabel Exome Probesets\ track exomeProbesets\ type bigBed\ visibility hide\ fantom5 FANTOM5 FANTOM5: Mapped transcription start sites (TSS) and their usage 0 100 0 0 0 127 127 127 0 0 0\ The FANTOM5 track shows mapped transcription start sites (TSS) and their usage in primary cells,\ cell lines, and tissues to produce a comprehensive overview of gene expression across the human\ body by using single molecule sequencing.\
\ \Items in this track are colored according to their strand orientation. Blue\ indicates alignment to the negative strand, and red indicates\ alignment to the positive strand.\
\ \Individual biological states are profiled by HeliScopeCAGE, which is a variation of the CAGE\ (Cap Analysis Gene Expression) protocol based on a single molecule sequencer. The standard protocol\ requiring 5 µg of total RNA as a starting material is referred to as hCAGE, and an\ optimized version for a lower quantity (~ 100 ng) is referred to as LQhCAGE (Kanamori-Katyama\ et al. 2011).\
Transcription start sites (TSSs) were mapped and their usage in human and mouse primary cells,\ cell lines, and tissues was to produce a comprehensive overview of mammalian gene expression across the\ human body. 5′-end of the mapped CAGE reads are counted at a single base pair resolution\ (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the\ sample. Individual samples shown in "TSS activity" tracks are grouped as below.\
TSS (CAGE) peaks across the panel of the biological states (samples) are identified by DPI\ (decomposition based peak identification, Forrest et al. 2014), where each of the peaks consists of\ neighboring and related TSSs. The peaks are used as anchors to define promoters and units of\ promoter-level expression analysis. Two subsets of the peaks are defined based on evidence of read\ counts, depending on scopes of subsequent analyses, and the first subset (referred as a\ robust set of the peaks, thresholded for expression analysis is shown as TSS peaks. They are\ named "p#@GENE_SYMBOL" if associated with 5'-end of known genes, or "p@CHROM:START..END,STRAND"\ otherwise. The summary tracks consist of the TSS (CAGE) peaks and summary profiles of TSS\ activities (total and maximum values). The summary track consists of the following tracks.\
\ 5′-end of the mapped CAGE reads are counted at a single base pair resolution (CTSS, CAGE tag starting sites) on the genomic coordinates, which represent TSS activities in the sample. The read counts tracks indicate raw counts of CAGE reads, and the TPM tracks indicate normalized counts as TPM (tags per million).\
\ \\ FANTOM5 data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ The FANTOM5 reprocessed data can be found and downloaded on the FANTOM website.
\ \\ Thanks to the FANTOM5 consortium,\ the Large Scale Data Managing Unit and Preventive Medicine and\ Applied Genomics Unit, the Center for Integrative Medical Sciences (IMS), and\ RIKEN for providing this data\ and its analysis.
\ \\ FANTOM Consortium and the RIKEN PMI and CLST (DGT), Forrest AR, Kawaji H, Rehli M, Baillie JK, de\ Hoon MJ, Haberle V, Lassmann T, Kulakovskiy IV, Lizio M et al.\ \ A promoter-level mammalian expression atlas.\ Nature. 2014 Mar 27;507(7493):462-70.\ PMID: 24670764; PMC: PMC4529748\
\ \\ Kanamori-Katayama M, Itoh M, Kawaji H, Lassmann T, Katayama S, Kojima M, Bertin N, Kaiho A, Ninomiya\ N, Daub CO et al.\ \ Unamplified cap analysis of gene expression on a single-molecule sequencer.\ Genome Res. 2011 Jul;21(7):1150-9.\ PMID: 21596820; PMC: PMC3129257\
\ \\ Lizio M, Harshbarger J, Shimoji H, Severin J, Kasukawa T, Sahin S, Abugessaisa I, Fukuda S, Hori F,\ Ishikawa-Kato S et al.\ \ Gateways to the FANTOM5 promoter level mammalian expression atlas.\ Genome Biol. 2015 Jan 5;16(1):22.\ PMID: 25723102; PMC: PMC4310165\
\ regulation 0 group regulation\ html fantom5.html\ longLabel FANTOM5: Mapped transcription start sites (TSS) and their usage\ shortLabel FANTOM5\ superTrack on\ track fantom5\ visibility hide\ fetalGeneAtlasAssay Fetal Assay bigBarChart Fetal Gene Atlas binned by assay (cell/nucleus) from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars Cell Nuclei\ barChartColors #4c758b #e5b909\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/Assay.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/Assay.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by assay (cell/nucleus) from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal Assay\ track fetalGeneAtlasAssay\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ fetalGeneAtlasCellType Fetal Cells bigBarChart Fetal Gene Atlas binned by cell type from Cao et al 2020 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars mixed_AFP+_ALB+_cell acinar_cell adrenocortical_cell amacrine_cell antigen_presenting_cell astrocyte bipolar_neuron bronchiolar/alveolar_epithelial_cell pancreas_CCL19+_CCL21+_cell heart_CLC+_IL5RA+_cell mixed_CSH1+_CSH2+_cell cardiomyocyte chromaffin_cell ciliated_epithelial_cell corneal/conjunctival_epithelial_cell ductal_cell heart_ELF3+_AGBL2+_cell enteric_nervous_system_(ENS)_glial_cell enteric_nervous_system_(ENS)_neuron endocardial_cell epicardial_adipose_cell erythroblast excitatory_neuron extravillous_trophoblast ganglion_cell goblet_cell granule_neuron hematopoietic_stem_cell hepatoblast horizontal_cell placenta_IGFBP1+_DKK1+_cell inhibitory_interneuron inhibitory_neuron intestinal_epithelial_cell islet_endocrine_cell lens_fibre_cell limbic_system_neuron lymphatic_endothelial_cell lymphoid_cell stomach_MUC13+_DMBT1+_cell megakaryocyte mesangial_cell mesothelial_cell metanephric_cell microglial_cell myeloid_cell neuroendocrine_cell oligodendrocyte placenta_PAEP+_MECOM+_cell eye_PDE11A+_FAM19A2+_cell stomach_PDE1C+_ACSM3+_cell parietal_and_chief_cell photoreceptor_cell Purkinje_neuron retinal_pigment_cell retinal_progenitor/Muller_glial_cell heart_SATB2+_LRRC7+_cell brain_SKOR2+_NPSR1+_cell brain_SLC24A4+_PEX5L+_cell adrenal_gland_SLC26A4+_PAEP+_cell spleen_STC2+_TLX1+_cell satellite_cell Schwann_cell skeletal_muscle_cell smooth_muscle_cell squamous_epithelial_cell stellate_cell stromal_cell sympathoblasts syncytiotrophoblast_and_villous_cytotrophoblast thymic_epithelial_cell thymocyte trophoblast_giant_cell unipolar_brush_cell ureteric_bud_cell vascular_endothelial_cell visceral_neuron\ barChartColors #c75cc6 #3259c7 #7d8952 #d3ac19 #de201f #adb119 #be9c2d #577881 #a4a096 #b787ac #9275da #af1ea8 #aa973d #477f92 #65b5cb #2f5cc6 #c471c0 #80c709 #cba81f #489338 #fe8839 #8a7352 #e1b60c #5f37bb #ddb311 #305cc5 #deb410 #ad4e3b #b001af #b99b2f #7c7062 #deb40f #e7ba08 #536a95 #3f61b4 #ad9f9a #e1b60d #0aba08 #d02b29 #4766a4 #8b6651 #82953b #d07f49 #8c9840 #d92422 #e31b1b #6c7676 #bca424 #756d72 #b39635 #999eaa #2b59cd #ae9537 #dcb212 #88775c #b09f2b #dbc46b #dcb212 #dab014 #c9c6b4 #618237 #8d656b #80c60a #b80db6 #8d5675 #2889a7 #838546 #809836 #958951 #79785f #87a9b4 #b5443b #5425d7 #d7b015 #507093 #12b50d #c9a721\ barChartFacets organ,cell_class,cell_type\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/cell_type.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by cell type from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal Cells\ track fetalGeneAtlasCellType\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ fetalGeneAtlasDonor Fetal Donor ID bigBarChart Fetal Gene Atlas binned by donor ID from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars H26350 H26547 H27058 H27098 H27295 H27423 H27431 H27432 H27458 H27464 H27471 H27472 H27473 H27474 H27477 H27552 H27620 H27634 H27771 H27772 H27798 H27799 H27870 H27876 H27909 H27913 H27915 H27948\ barChartColors #647e66 #8a933b #e2b60c #92953b #ae20a5 #c8a91d #e5b909 #dfb40f #d6af15 #e3b80b #e3b80a #deb50e #a5199f #e6ba08 #e4b80a #9d9935 #cdaa1d #e6ba08 #859d34 #70904e #85973d #70835e #1a58dc #2359d2 #3f69a4 #779052 #64846d #557f72\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/donor.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by donor ID from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal Donor ID\ track fetalGeneAtlasDonor\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ fetalGeneAtlasExperiment Fetal Exp bigBarChart Fetal Gene Atlas binned by experiment id from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars exp1 exp2 exp3 exp4 exp5 exp6 exp7\ barChartColors #c9ab1b #dfb50e #d0ae18 #d4b114 #e8bb07 #5e836a #406ea0\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/Experiment_batch.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/Experiment_batch.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by experiment id from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal Exp\ track fetalGeneAtlasExperiment\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ fetalGeneAtlas Fetal Gene Atlas bigBarChart Fetal Gene Atlas from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 group singleCell\ longLabel Fetal Gene Atlas from Cao et al 2020\ pennantIcon 19.jpg liftover.html "lifted from hg19"\ shortLabel Fetal Gene Atlas\ superTrack on\ track fetalGeneAtlas\ type bigBarChart\ visibility hide\ fetalGeneAtlasOrganCellLineage Fetal Lineage bigBarChart Fetal Gene Atlas binned by cell lineage and organ from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars Adrenal-Adrenocortical_cells Adrenal-CSH1_CSH2_positive_cells Adrenal-Chromaffin_cells Adrenal-Erythroblasts Adrenal-Lymphoid_cells Adrenal-Megakaryocytes Adrenal-Myeloid_cells Adrenal-SLC26A4_PAEP_positive_cells Adrenal-Schwann_cells Adrenal-Stromal_cells Adrenal-Sympathoblasts Adrenal-Vascular_endothelial_cells Cerebellum-Astrocytes Cerebellum-Granule_neurons Cerebellum-Inhibitory_interneurons Cerebellum-Microglia Cerebellum-Oligodendrocytes Cerebellum-Purkinje_neurons Cerebellum-SLC24A4_PEX5L_positive_cells Cerebellum-Unipolar_brush_cells Cerebellum-Vascular_endothelial_cells Cerebrum-Astrocytes Cerebrum-Excitatory_neurons Cerebrum-Inhibitory_neurons Cerebrum-Limbic_system_neurons Cerebrum-Megakaryocytes Cerebrum-Microglia Cerebrum-Oligodendrocytes Cerebrum-SKOR2_NPSR1_positive_cells Cerebrum-Vascular_endothelial_cells Eye-Amacrine_cells Eye-Astrocytes Eye-Bipolar_cells Eye-Corneal_and_conjunctival_epithelial_cells Eye-Ganglion_cells Eye-Horizontal_cells Eye-Lens_fibre_cells Eye-Microglia Eye-PDE11A_FAM19A2_positive_cells Eye-Photoreceptor_cells Eye-Retinal_pigment_cells Eye-Retinal_progenitors_and_Muller_glia Eye-Skeletal_muscle_cells Eye-Smooth_muscle_cells Eye-Stromal_cells Eye-Vascular_endothelial_cells Heart-CLC_IL5RA_positive_cells Heart-Cardiomyocytes Heart-ELF3_AGBL2_positive_cells Heart-Endocardial_cells Heart-Epicardial_fat_cells Heart-Erythroblasts Heart-Lymphatic_endothelial_cells Heart-Lymphoid_cells Heart-Megakaryocytes Heart-Myeloid_cells Heart-SATB2_LRRC7_positive_cells Heart-Schwann_cells Heart-Smooth_muscle_cells Heart-Stromal_cells Heart-Vascular_endothelial_cells Heart-Visceral_neurons Intestine-Chromaffin_cells Intestine-ENS_glia Intestine-ENS_neurons Intestine-Erythroblasts Intestine-Intestinal_epithelial_cells Intestine-Lymphatic_endothelial_cells Intestine-Lymphoid_cells Intestine-Mesothelial_cells Intestine-Myeloid_cells Intestine-Smooth_muscle_cells Intestine-Stromal_cells Intestine-Vascular_endothelial_cells Kidney-Erythroblasts Kidney-Lymphoid_cells Kidney-Megakaryocytes Kidney-Mesangial_cells Kidney-Metanephric_cells Kidney-Myeloid_cells Kidney-Stromal_cells Kidney-Ureteric_bud_cells Kidney-Vascular_endothelial_cells Liver-Erythroblasts Liver-Hematopoietic_stem_cells Liver-Hepatoblasts Liver-Lymphoid_cells Liver-Megakaryocytes Liver-Mesothelial_cells Liver-Myeloid_cells Liver-Stellate_cells Liver-Vascular_endothelial_cells Lung-Bronchiolar_and_alveolar_epithelial_cells Lung-CSH1_CSH2_positive_cells Lung-Ciliated_epithelial_cells Lung-Lymphatic_endothelial_cells Lung-Lymphoid_cells Lung-Megakaryocytes Lung-Mesothelial_cells Lung-Myeloid_cells Lung-Neuroendocrine_cells Lung-Squamous_epithelial_cells Lung-Stromal_cells Lung-Vascular_endothelial_cells Lung-Visceral_neurons Muscle-Erythroblasts Muscle-Lymphatic_endothelial_cells Muscle-Lymphoid_cells Muscle-Megakaryocytes Muscle-Myeloid_cells Muscle-Satellite_cells Muscle-Schwann_cells Muscle-Skeletal_muscle_cells Muscle-Smooth_muscle_cells Muscle-Stromal_cells Muscle-Vascular_endothelial_cells Pancreas-Acinar_cells Pancreas-CCL19_CCL21_positive_cells Pancreas-Ductal_cells Pancreas-ENS_glia Pancreas-ENS_neurons Pancreas-Erythroblasts Pancreas-Islet_endocrine_cells Pancreas-Lymphatic_endothelial_cells Pancreas-Lymphoid_cells Pancreas-Mesothelial_cells Pancreas-Myeloid_cells Pancreas-Smooth_muscle_cells Pancreas-Stromal_cells Pancreas-Vascular_endothelial_cells Placenta-AFP_ALB_positive_cells Placenta-Extravillous_trophoblasts Placenta-IGFBP1_DKK1_positive_cells Placenta-Lymphoid_cells Placenta-Megakaryocytes Placenta-Myeloid_cells Placenta-PAEP_MECOM_positive_cells Placenta-Smooth_muscle_cells Placenta-Stromal_cells Placenta-Syncytiotrophoblasts_and_villous_cytotrophoblasts Placenta-Trophoblast_giant_cells Placenta-Vascular_endothelial_cells Spleen-AFP_ALB_positive_cells Spleen-Erythroblasts Spleen-Lymphoid_cells Spleen-Megakaryocytes Spleen-Mesothelial_cells Spleen-Myeloid_cells Spleen-STC2_TLX1_positive_cells Spleen-Stromal_cells Spleen-Vascular_endothelial_cells Stomach-Ciliated_epithelial_cells Stomach-ENS_glia Stomach-ENS_neurons Stomach-Erythroblasts Stomach-Goblet_cells Stomach-Lymphatic_endothelial_cells Stomach-Lymphoid_cells Stomach-MUC13_DMBT1_positive_cells Stomach-Mesothelial_cells Stomach-Myeloid_cells Stomach-Neuroendocrine_cells Stomach-PDE1C_ACSM3_positive_cells Stomach-Parietal_and_chief_cells Stomach-Squamous_epithelial_cells Stomach-Stromal_cells Stomach-Vascular_endothelial_cells Thymus-Antigen_presenting_cells Thymus-Stromal_cells Thymus-Thymic_epithelial_cells Thymus-Thymocytes Thymus-Vascular_endothelial_cells\ barChartColors #7d8952 #9478d5 #aa963d #826c5f #c03833 #856855 #d82423 #c9c6b4 #80c50c #72923c #958951 #22ab19 #abb219 #deb410 #deb40f #d82524 #b9a227 #dcb212 #dab014 #d7b015 #2fa323 #b5ac1e #e1b60c #e7ba08 #e1b60d #ddd4cb #d92423 #bea523 #dcb212 #2da521 #d3ac19 #bdbb78 #be9c2d #65b5cb #ddb311 #b99b2f #ad9f9a #e17170 #b39635 #ae9537 #88775c #b09f2b #c96ac4 #dad6cf #85844c #7bbf6f #b787ac #af1ea8 #c471c0 #489338 #fe8839 #e8e0e0 #0fb60c #cd2d2a #8b6354 #e11d1d #dbc46b #80c20e #8e5377 #82814e #22ab1a #bb9d2f #6c717d #81c10f #cca91e #af9a98 #536a95 #22ac18 #c23732 #c5a98d #db2121 #8c666a #7f933d #24aa1a #ad9999 #c53330 #cdb9ad #82953b #8c9840 #dd201f #7b884b #507093 #20ad17 #8b7450 #ad4e3b #b001af #c13a33 #8f6550 #d2a78b #cf2c2a #838546 #409a2b #577881 #cfc2ee #487f91 #0bb909 #cf2c29 #846859 #cf7f4a #e11d1d #707770 #4c7a91 #889934 #1ead16 #dcc56a #dbd2d2 #65cb61 #d07f7b #dbd2cf #e36f6e #8d656b #abd164 #b80db6 #876467 #837c53 #22aa1a #3259c7 #a4a096 #2f5cc6 #80c30e #a29142 #81646c #3f61b4 #6ac766 #c33333 #c1a593 #d92323 #845e72 #7a7f54 #21aa1b #c65ec5 #5f37bb #7c7062 #9f534f #dcd2ce #cd2d2c #756d72 #86715c #827d51 #79785f #5425d7 #4d9138 #c65fc4 #946446 #c53530 #b3998b #dba988 #d62624 #618237 #849336 #25aa19 #477f92 #aac86e #c3b580 #ae9999 #305cc5 #7abe71 #c43531 #4766a4 #bda693 #e17170 #989ead #999eaa #2b59cd #2d87a3 #7b8b47 #74c26c #de201f #aea28c #87a9b4 #b5443b #86b876\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/Organ_cell_lineage.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/Organ_cell_lineage.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by cell lineage and organ from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal Lineage\ track fetalGeneAtlasOrganCellLineage\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ fetalGeneAtlasOrgan Fetal Organ bigBarChart Fetal Gene Atlas binned by organ from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars Adrenal Cerebellum Cerebrum Eye Heart Intestine Kidney Liver Lung Muscle Pancreas Placenta Spleen Stomach Thymus\ barChartColors #7c8e4a #e6ba08 #e5b909 #d6b015 #ae20a6 #5f7577 #849c3a #aa0ea6 #619841 #b90db6 #2359d2 #6f637a #836824 #1e5ad9 #b94138\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/Organ.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/Organ.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by organ from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal Organ\ track fetalGeneAtlasOrgan\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ fetalGeneAtlasRtGroup Fetal RT Group bigBarChart Fetal Gene Atlas binned by RT group from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars Adrenal_H26350 Adrenal_H26547 Adrenal_H27098 Adrenal_H27471 Adrenal_H27472 Adrenal_H27474 Adrenal_H27552 Cerebellum_H27471 Cerebellum_H27472 Cerebellum_H27474 Cerebellum_H27477 Cerebellum_H27634 Cerebrum_H27058 Cerebrum_H27098 Cerebrum_H27423 Cerebrum_H27431 Cerebrum_H27432 Cerebrum_H27464 Cerebrum_H27471 Cerebrum_H27474 Eye_H27458 Eye_H27472 Eye_H27552 Eye_H27620 Eye_H27634 Heart_H26547 Heart_H27098 Heart_H27295 Heart_H27423 Heart_H27431 Heart_H27464 Heart_H27471 Heart_H27472 Heart_H27473 Intestine_H27771 Intestine_H27772 Intestine_H27798 Intestine_H27799 Intestine_H27876 Intestine_H27909 Intestine_H27913 Intestine_H27915 Intestine_H27948 Kidney_H27771 Kidney_H27772 Kidney_H27798 Kidney_H27870 Kidney_H27876 Kidney_H27909 Kidney_H27913 Kidney_H27915 Kidney_H27948 Liver_H27058 Liver_H27098 Liver_H27423 Liver_H27431 Liver_H27464 Liver_H27471 Liver_H27472 Liver_H27473 Liver_H27474 Lung_H26350 Lung_H26547 Lung_H27058 Lung_H27098 Lung_H27423 Lung_H27431 Lung_H27464 Lung_H27471 Lung_H27472 Lung_H27474 Lung_H27477 Muscle_H27098 Muscle_H27431 Muscle_H27471 Muscle_H27472 Muscle_H27473 Muscle_H27474 Muscle_H27477 Muscle_H27634 Pancreas_H27870 Pancreas_H27876 Pancreas_H27948 Placenta_H26350 Placenta_H26547 Placenta_H27058 Placenta_H27098 Placenta_H27423 Placenta_H27431 Placenta_H27464 Placenta_H27471 Placenta_H27472 Placenta_H27473 Placenta_H27474 Spleen_H26350 Spleen_H26547 Spleen_H27431 Spleen_H27464 Spleen_H27471 Spleen_H27472 Spleen_H27552 Spleen_H27634 Stomach_H27870 Stomach_H27876 Stomach_H27909 Stomach_H27948 Thymus_H26547 Thymus_H27423 Thymus_H27431 Thymus_H27471 Thymus_H27552 Thymus_H27634\ barChartColors #6a7d64 #92933f #71855a #6a7e64 #9e9936 #798e46 #879241 #e6b908 #e6b909 #e5b909 #e4b909 #e6ba08 #e2b60c #dab311 #e2b60c #e5b909 #dfb40f #e4b80a #e1b60d #e7ba08 #d6af15 #d8b114 #d4ae17 #cdaa1d #d1ad18 #ad21a5 #ab24a2 #ae20a5 #ae1ea8 #ae1fa6 #ab24a1 #a23394 #ae1fa7 #ac23a3 #616e83 #5d6d86 #717a5c #70835e #566d8a #656e65 #d6d3d6 #577189 #e6e1e3 #869d33 #72914c #999d33 #909a39 #819643 #8a993f #789051 #7f9646 #859743 #a616a1 #a22297 #a22099 #ab0ba9 #a713a3 #a715a1 #a813a3 #ab0ca8 #ac09aa #537f72 #629545 #549541 #588d57 #5e9b34 #4c8c58 #568f4e #7d993d #579e2f #6aa030 #548666 #b611b2 #b90cb6 #b80eb4 #b80eb5 #b611b2 #ac24a2 #b612b1 #b90cb6 #2059d5 #2858cd #68723e #5425d7 #9b97a8 #5c34c1 #737555 #d3c9e3 #6c5c85 #8f71e1 #81824e #746d6c #6d7959 #757f54 #b1997a #7f6d2f #ab9b79 #7e6a33 #687e28 #5e8638 #658027 #a05427 #1c58db #2c86a5 #1e57db #2b73b6 #b6433a #bc3d36 #e7e1e2 #e3ccca #e5ccc8 #bb3f37\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/RT_group.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/RT_group.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by RT group from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal RT Group\ track fetalGeneAtlasRtGroup\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ fetalGeneAtlasSex Fetal Sex bigBarChart Fetal Gene Atlas binned by sex from Cao et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\
This group of tracks shows data from \
A human cell atlas of fetal gene expression. This is a collection of\
single cell and single nucleus combinatorial indexing-based RNA-seq data covering 4 million\
cells from 15 organs obtained during mid-gestation. The cells were sequenced in\
a highly multiplexed fashion and then clustered with annotations as described\
in Cao et al., 2020.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The Fetal Cells subtrack contains the \ data organized by cell type, with RNA signals from all cells of a given type pooled \ and averaged into one bar for each cell type. The \ Fetal Lineage subtrack shows \ similar data, but with the cell types subdivided more finely and by organ. Additional \ bar chart subtracks pool the cell by other characteristics such as by sex \ (Fetal Sex), assay \ (FetalAssay), donor \ (Fetal Donor ID), experiment \ (Fetal Exp), organ \ (Fetal Organ), and reverse transcription group \ (Fetal RT Group).
\ \\ Please see descartes.brotmanbaty.org for\ further interactive displays and additional data.
\ \\ The cell types are colored by which class they belong to according to the following table.\ The coloring algorithm allows cells that show some blended characteristics to show blended\ colors so there will be some color variation within a class. The colors will be purest in\ the Fetal Cells subtrack, where the bars \ represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.\ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia |
\ Three-level single-cell combinatorial indexing (sci-RNAseq3) as described in\ Cao et al., 2020 was used on 121 samples from 28 fetuses estimated 72\ to 129 days post-conception. This included samples from 15 organs. and\ resulted in RNA profiles for 4 million cells. The samples were flash-frozen for\ majority of the experiments and then nuclei extracted for sequencing. Samples\ from tissues from the kidney and digestive system were fixed after\ disassociation to deactivate endogenous RNases and proteases.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\ matrixToBarChart, and bedToBigBed were used to transform these into a bar chart\ format bigBed file that can be visualized. The coloring was done by defining\ colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types.\ The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \ \ \Thanks to the many authors who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \ \ singleCell 1 barChartBars F M\ barChartColors #dbb410 #e6ba08\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/fetalGeneAtlas/sex.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/fetalGeneAtlas/sex.bb\ defaultLabelFields name2\ html fetalGeneAtlas\ labelFields name,name2\ longLabel Fetal Gene Atlas binned by sex from Cao et al 2020\ parent fetalGeneAtlas\ shortLabel Fetal Sex\ track fetalGeneAtlasSex\ transformFunc NONE\ url https://cells.ucsc.edu/?ds=fetal-gene-atlas+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ fishClones FISH Clones bed 5 + Clones Placed on Cytogenetic Map Using FISH 0 100 0 150 0 127 202 127 0 0 0\ This track shows the location of fluorescent in situ hybridization \ (FISH)-mapped clones along the assembly sequence. The locations of\ these clones were obtained from the NCBI Human BAC Resource\ here. Earlier versions of this track obtained this\ information directly from the paper Cheung, et al. (2001).\
\ \\ More information about the BAC clones, including how they may be obtained, \ can be found at the \ Human BAC Resource and the \ Clone Registry web sites hosted by \ NCBI.\ To view Clone Registry information for a clone, click on the clone name at \ the top of the details page for that item.
\ \\ This track has a filter that can be used to change the color or \ include/exclude the display of a dataset from an individual lab. This is \ helpful when many items are shown in the track display, especially when only \ some are relevant to the current task. The filter is located at the top of \ the track description page, which is accessed via the small button to the \ left of the track's graphical display or through the link on the track's \ control menu. To use the filter:\
\ When you have finished configuring the filter, click the Submit \ button.
\ \\ We would like to thank all of the labs that have contributed to this resource:\
\ Cheung VG, Nowak N, Jang W, Kirsch IR, Zhao S, Chen XN, Furey TS, Kim UJ, Kuo WL, Olivier M et\ al.\ \ Integration of cytogenetic landmarks into the draft sequence of the human genome.\ Nature. 2001 Feb 15;409(6822):953-8.\ PMID: 11237021\
\ map 1 color 0,150,0,\ group map\ longLabel Clones Placed on Cytogenetic Map Using FISH\ origAssembly hg18\ pennantIcon 18.jpg ../goldenPath/help/liftOver.html "lifted from hg18"\ shortLabel FISH Clones\ superTrack assemblyContainer pack\ track fishClones\ type bed 5 +\ visibility hide\ g2p G2P Project bigBed 9 + Gene2Phenotype Project 0 100 0 0 0 127 127 127 0 0 0\ This track displays detailed, evidence-based gene-disease models, curated from the literature by\ experts. The track can be used to filter genomic sequencing data from individuals with genetic\ disorders to identify likely causative variants and accelerate diagnosis. More information about\ the G2P project can be found on the\ Gene2Phenotype\ website.\
\ \\ For each track, items are colored according to the likelihood that the gene-disease\ association is true:
\Each mouseover tooltip provides the following information:
\\ Expert-curated gene disease models released by the Gene2Phenotype project were imported and\ processed to create a BED-based track annotating genomic regions reported to be associated with\ disease in the literature. Standard genome assembly coordinates and gene annotations were used to\ map entries to the browser.\
\ \\ For more information on the Gene2Phenotype project, please contact \ \ g2p-help@ebi.\ ac.\ uk\ \
\ \\ Source data for these tracks are available directly from\ Gene2Phenotype. \
\ \\ Thormann A, Halachev M, McLaren W, Moore DJ, Svinti V, Campbell A, Kerr SM, Tischkowitz M, Hunt SE,\ Dunlop MG et al.\ \ Flexible and scalable diagnostic filtering of genomic variants using G2P with Ensembl VEP.\ Nat Commun. 2019 May 30;10(1):2373.\ DOI: 10.1038/s41467-019-10016-3; PMID: 31147538; PMC: PMC6542828\
\\ Yates TM, Ansari M, Thompson L, Hunt SE, Uhalte EC, Hobson RJ, Marsh JA, Wright CF, Firth HV.\ \ Curating genomic disease-gene relationships with Gene2Phenotype (G2P).\ Genome Med. 2024 Nov 6;16(1):127.\ DOI: 10.1186/s13073-024-01398-1; PMID: 39506859; PMC: PMC11539801\
\ phenDis 1 bigDataUrl /gbdb/hg38/g2p/g2p.bb\ cartVersion 9\ group phenDis\ html g2p.html\ longLabel Gene2Phenotype Project\ mouseOver G2P ID: ${g2p_id}\ This track shows structural variants (SVs) identified by PacBio HiFi long-read\ sequencing of probands and their families enrolled in the Genomic Answers for\ Kids (GA4K) program at Children's Mercy Research Institute. GA4K is a\ longitudinal pediatric genomics initiative that aims to enroll 30,000 children\ with suspected rare genetic disorders, together with their parents, to build\ a large-scale resource of clinical and genomic data.\
\\ The callset contains 115,554 SVs (52,564 deletions, 58,219 insertions, 4,408\ duplications, 363 inversions) from 502 sequenced samples. Variants are\ site-level (no per-sample genotypes) and each SV has been replicated, meaning\ that it was either observed in two or more unrelated GA4K individuals, or\ matched an SV from an external long-read reference set (Decode or the Human\ Pangenome Reference Consortium).\
\ \\ Items are colored by SV type:\
\ Insertions are placed at the insertion site with a width of 1 bp; deletions,\ duplications and inversions span the affected interval. Filters are available\ for SV type, SV length, carrier-sample count and allele frequency. The detail\ page also shows the total number of samples genotyped at each site.\
\ \\ The Genomic Answers for Kids (GA4K) program at Children's Mercy Research\ Institute is a longitudinal pediatric rare-disease initiative described in\ Cohen et al. 2022. GA4K probands and their families are sequenced with\ PacBio HiFi long reads (Revio and Sequel II), and the 502-sample GA4K\ PacBio SV release (pb_joint_merged.sv.vcf.gz) is produced by\ running \ pbsv per sample and merging with\ JASMINE\ v1.1.4 (--output-genotypes). The merged site-level VCF is\ filtered to SVs replicated in at least two independent observations\ (either matching a second unrelated CMH individual in the same Jasmine\ cluster, or matching an SV in the deCODE Icelandic or HPRC callsets via\ \ svpack match). The released catalog contains 115,554 replicated SVs\ (52,564 deletions, 58,219 insertions, 4,408 duplications and 363\ inversions) with recomputed carrier counts (SVC), total sample counts\ (SVN) and allele frequencies (SVF = SVC/SVN).\
\\ The source VCF was cloned from the Children's Mercy Research Institute\ GA4K GitHub repository,\ \ github.com/ChildrensMercyResearchInstitute/GA4K\ (pacbio_sv_vcf/pb_joint_merged.sv.vcf.gz).\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator and exported from there\ to spreadsheet or tab-sep tables. From scripts, the data can be accessed\ through our API, track=ga4kSv.\
\\ For automated download and analysis, the annotation is stored in a bigBed file\ that can be downloaded from\ our\ download server. The file for this track is called ga4kSv.bb.\ Individual regions or the whole annotation can be obtained using the\ bigBedToBed utility, available as a precompiled binary or from source\ as described on our\ utilities\ page.\ Example:\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/ga4kSv.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\\ The original VCF is available from the Children's Mercy Research Institute\ GA4K data release at\ \ github.com/ChildrensMercyResearchInstitute/GA4K.\
\ \\ Thanks to the Children's Mercy Research Institute and the Genomic Answers\ for Kids participants and their families for making this dataset publicly\ available.\
\ \\ Cohen ASA, Farrow EG, Abdelmoity AT, Alaimo JT, Amudhavalli SM, Anderson JT, Bansal L, Bartik L,\ Baybayan P, Belden B et al.\ \ Genomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes.\ Genet Med. 2022 Jun;24(6):1336-1348.\ PMID: 35305867\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/ga4kSv.bb\ filter.AC 0:996\ filter.alleleFreq 0:1\ filter.carrierCount 1:498\ filter.insLen 0:14923\ filter.svLen 0:809711\ filterByRange.AC on\ filterByRange.alleleFreq on\ filterByRange.carrierCount on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Allele Count (approx)\ filterLabel.alleleFreq Allele Frequency\ filterLabel.carrierCount Number of Carrier Samples\ filterLabel.insLen Insertion Length\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.alleleFreq 0:1\ filterType.svType multipleListOr\ filterValues.svType DEL,INS,DUP,INV\ itemRgb on\ longLabel Structural Variants from 502 Children's Mercy GA4K Probands (PacBio HiFi)\ mouseOver Var: $name ($svType)\ This track shows the gaps in the GRCh38 (hg38) genome assembly defined in the\ AGP file delivered with the sequence. These gaps are being closed during the \ finishing process on the human genome. For information on the AGP file format, see the NCBI \ AGP Specification. The NCBI website also provides an \ overview of genome assembly procedures, as well as \ specific information about the hg38 assembly.\
\\ Gaps are represented as black boxes in this track.\ If the relative order and orientation of the contigs on either side\ of the gap is supported by read pair data, \ it is a bridged gap and a white line is drawn \ through the black box representing the gap. \
\This assembly contains the following principal types of gaps:\
\ The GC percent track shows the percentage of G (guanine) and C (cytosine) bases\ in 5-base windows. High GC content is typically associated with\ gene-rich areas.\
\\ This track may be configured in a variety of ways to highlight different\ apsects of the displayed information. Click the\ "Graph configuration help"\ link for an explanation of the configuration options.\ \
The data and presentation of this graph were prepared by\ Hiram Clawson.\
\ \ map 0 altColor 128,128,128\ autoScale Off\ color 0,0,0\ graphTypeDefault Bar\ gridDefault OFF\ group map\ html gc5Base\ longLabel GC Percent in 5-Base Windows\ maxHeightPixels 128:36:16\ shortLabel GC Percent\ track gc5BaseBw\ type bigWig 0 100\ viewLimits 30:70\ visibility hide\ windowingFunction Mean\ genCC GenCC bigBed 9 + 34 GenCC: The Gene Curation Coalition Annotations 0 100 0 0 0 127 127 127 0 0 0 https://search.thegencc.org/genes/$\ This track shows annotations from The Gene Curation Coalition (GenCC).\ The GenCC provides information pertaining to the validity of gene-disease relationships, \ with a current focus on Mendelian diseases. Curated gene-disease relationships are submitted \ by GenCC member organizations that currently provide online resources (e.g. ClinGen, DECIPHER, \ Orphanet, etc.), as well as diagnostic laboratories that have committed to sharing their internal \ curated gene-level knowledge (e.g. Ambry Genetics, Illumina, Invitae, etc.).
\\ The GenCC aims to clarify overlap between gene curation efforts and develop\ consistent terminology for validity, allelic requirement and mechanism\ of disease. Each item on this track corresponds with a gene, and contains\ a large number of information such as associated disease, evidence classification,\ specific submission notes and identifiers from different databases. In cases where\ multiple annotations exist for the same gene, multiple items are displayed.
\ \\ Each item displayed represents a submission to the GenCC database. The displayed \ name is a combination of the gene symbol and the disease's original submission ID. \ This submission ID is either the OMIM#, MONDO# or Orphanet#. Clicking\ on any item will display the complete meta data for that item, including\ linkouts to the GenCC, NCBI, Ensembl, HGNC, GeneCards, Pombase (MONDO),\ and Human Phenotype Ontology (HPO). Mousing over any item will display the\ associated disease title, the classification title, and the mode of inheritance\ title.
\ \\ Items are colored based on the GenCC classification, or validation, of the\ evidence in the color scheme seen in the table below. \ For more information on this process, see the GenCC\ validity terms FAQ. A filter for the track is also available\ to display a subset of the items based on their classification.
\ \\
| Color | \Evidence classification | \
|---|---|
| Definitive | |
| Strong | |
| Moderate | |
| Supportive | |
| Limited | |
| Disputed Evidence | |
| Refuted Evidence | |
| No Known Disease Relationship |
\ Limitations: Most entries include both NM_ accessions as well as ENST and ENSG identifiers.\ From the original file, which contains no coordinates, two genes were not mapped\ to the hg38 genome, SLCO1B7 and ATXN8. This means that the hg38 track has 2 fewer items\ than what can be found in the GenCC download file. For hg19, one additional\ gene was not mapped, KCNJ18. In addition to this, the GenCC data in the Genome\ Browser does not include OMIM data due to licensing restrictions. For more\ information, see the Methods section below.
\ \\ The source data can be explored in \ GenCC database. The source files can also be found on the GenCC downloads page.
\ \\ The GenCC data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored at UCSC in bigBed\ files that can be downloaded from\ our download server.\ The data may also be explored interactively using our\ REST API.
\ \\
The file for this track may also be locally explored using our tools bigBedToBed \
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/genCC.bb stdout
\ The data were downloaded from the GenCC downloads page in tsv format. Manual\ curation was performed on the file to remove newline characters and tab characters present in \ the submission notes, in total fewer than 20 manual edits were made.
\\ The track was first built on hg38 by associating the gene symbols with the NCBI MANE 1.0 \ release transcripts. These coordinates were added to the items as well as the NM_ accession,\ ENST ID and ENSG ID. For items where there was no gene symbol match in MANE (~130), the gene\ symbols were queried against GENCODEv40 comprehensive set release. In places where multiple\ transcript matches were found, the earliest transcription start and latest end site was used\ from among the transcripts to encompass the entire gene coordinates. Two genes were not able\ to be mapped for hg38, SLCO1B7 and ATXN8, resulting in two missing submissions in the Genome\ Browser when compared to the raw file. Lastly, the items were colored according to their\ evidence classification as seen on the GenCC database.
\\ For hg19, the hg38 NM_ accessions were used to convert the item coordinates according to the\ latest hg19 refseq release. For items that failed to convert, the gene symbols were queried\ using the GENCODEv40 hg19 lift comprehensive set. One additional gene symbol failed to map in\ hg19, KCNJ18, leading to 3 fewer items on this track when compared to the raw file.
\\ For both assemblies, GenCC OMIM data is excluded do to data restrictions.\ For complete documentation of the processing of these tracks, read the\ \ GenCC MakeDoc.
\ \\ Thanks to the entire GenCC\ committee for creating these annotations and making them available.
\ \\ DiStefano MT, Goehringer S, Babb L, Alkuraya FS, Amberger J, Amin M, Austin-Tse C, Balzotti M, Berg\ JS, Birney E et al.\ \ The Gene Curation Coalition: A global effort to harmonize gene-disease evidence resources.\ Genet Med. 2022 May 4;.\ PMID: 35507016\
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/genCC.bb\ filterLabel.classification_title evidence classification\ filterValues.classification_title Supportive,Strong,Definitive,Limited,Moderate,No Known Disease Relationship,Disputed Evidence,Refuted Evidence\ group phenDis\ itemRgb on\ longLabel GenCC: The Gene Curation Coalition Annotations\ mouseOver Disease title: $disease_title\ This super track contains previous versions of the GENCODE primary gene set.\ genes 1 group genes\ html ../../knownGeneArchive\ longLabel GENCODE Archive\ shortLabel GENCODE Archive\ superTrack on\ track knownGeneArchive\ type bed 6 +\ gencNcOrfsComprehensive GENCODE Phase II ncORFs Compr bigGenePred ncORFs: GENCODE Phase II non-canonical ORFs - comprehensive 3 100 0 0 0 127 127 127 0 0 0
\ The three Gencode ncORF tracks in the non-canonical ORF track container show \ non-canonical translated open reading frames (ncORFs) identified\ from ribosome profiling (Ribo-seq) data and mapped to the GENCODE annotation by the\ GENCODE / TransCODE consortium.\ The data is available in two phases:\
\ \\ The Phase I catalog contains 7,264 unique human ncORFs called from Ribo-seq data\ across seven publications and mapped to GENCODE v35. Only translations of 16 codons or above\ and initiating from ATG start codons were incorporated. Redundant sense-overlapping ORFs were\ merged. Of these, 3,085 ORFs were found by more than one publication, providing independent\ replication evidence. This catalog was developed as part of an effort to standardize the\ annotation of translated ORFs across reference databases including Ensembl/GENCODE, HGNC,\ UniProtKB, and PeptideAtlas.\
\ \\ The Phase II catalog nearly quadruples the Phase I set, defining 28,359 ncORFs in the\ Comprehensive set, mapped to GENCODE v45. Compared to Phase I, additional published\ Ribo-seq datasets were incorporated and the restrictions on ORF size and initiation codon\ were lifted.\
\ \\ Two subsets are provided for the Phase II data:\
\\ All three GENCODE ncORF tracks are displayed in bigGenePred format and labeled with their\ ORF identifier. The default color scheme and available filter controls differ by track.\
\ \\ The Phase I track colors items by Kozak consensus strength by default.\ Two alternative color schemes can be selected from the track controls page\ (Color by dropdown): Evidence type and HLA class (see below).\
\ \| \ | Golden amber — Strong Kozak context. Both position −3 (A/G) and\ position +4 (G) match the consensus. | \
|---|---|
| \ | Steel blue — Moderate Kozak context. One of the two positions matches. | \
| \ | Gray — Weak Kozak context. Neither position matches. | \
| \ | Black — Non-ATG start codon (Kozak rule does not apply) or context\ unavailable. | \
\ Select Color by: Evidence type to highlight peptide evidence from\ Deutsch et al. (see References). ORFs with no mass spectrometry evidence are gray.\
\ \| \ | Gold — TransCODE peptidein (628 ORFs). Confirmed as\ confidently translated by PeptideAtlas; candidate for peptidein annotation in\ reference databases. | \
|---|---|
| \ | Steel blue — HLA immunopeptidomics evidence only (1,373 ORFs). | \
| \ | Forest green — Non-HLA (whole-cell tryptic) evidence only (66 ORFs). | \
| \ | Orange — Both HLA and non-HLA evidence (35 ORFs). | \
| \ | Gray — No peptide evidence (5,114 ORFs) or in the peptidein set\ based on binding predictions only with no direct MS sequences (48 ORFs). | \
\ Select Color by: HLA class to color items by the HLA class in which peptides were\ detected:\
\ \| \ | Steel blue — Class I only (1,632 ORFs). | \
|---|---|
| \ | Crimson — Class II only (10 ORFs). | \
| \ | Orange — Both class I and class II (143 ORFs). | \
| \ | Gray — No HLA data (5,479 ORFs). | \
\ The Phase I track can be filtered by: start codon, Kozak strength, Kozak TE, replicated\ status, and — using the peptide evidence fields — peptidein status\ (isPeptidein), HLA class (hlaClass), HLA evidence tier\ (hlaFinalTier), HPP guideline category (hlaHppCategory), and Ribo-seq\ quality (riboseqQuality).\
\ \\ Mouseover for Phase I shows ORF name, host gene, Kozak strength and TE,\ replicated status, peptidein flag, HLA evidence tier, and HLA peptide count.\
\ \\ The Phase II Primary and Comprehensive tracks color items by Kozak strength using the\ same scheme as Phase I. Peptide evidence fields are not included in the Phase II tracks.\ Common filters: start codon, Kozak strength, Kozak TE.\
\ \\ Each Phase I item carries the following peptide evidence fields from Deutsch et al.\ (2026), accessible via the details page and Table Browser:\
\ \| Field | Description |
|---|---|
| isPeptidein | yes/no: ORF is in the PeptideAtlas peptidein set (Table S12) |
| hlaClass | HLA class(es) detected: I, II, or Both |
| hlaFinalTier | HLA evidence tier (Tier 1B = numerous peptides; Tier 2B = one peptide) |
| hlaHppCategory | HPP guideline category (HPP+, 1PepCandidate, Insufficient) |
| hlaNPeptides | Number of distinct HLA peptide sequences detected |
| riboseqQuality | Manual quality of Ribo-seq evidence (Excellent/Sufficient/Insufficient) |
| hlaIPeptides | HLA class I peptide sequences (comma-separated) |
| hlaIIPeptides | HLA class II peptide sequences (comma-separated) |
| nonHlaFinalTier | Non-HLA (tryptic proteomics) evidence tier |
| nonHlaHppCategory | Non-HLA HPP guideline category |
| nonHlaNPeptides | Number of distinct non-HLA peptide sequences |
| nonHlaPeptides | Non-HLA tryptic peptide sequences (comma-separated) |
\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. The data can be accessed from\ scripts through our API; the track names are\ "gencNcOrfs" (Phase I), "gencNcOrfsPrimary" (Phase II Primary),\ and "gencNcOrfsComprehensive" (Phase II Comprehensive).\
\ \\ For automated download and analysis, the genome annotations are stored in bigBed files that\ can be downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\
\ \\ Mudge et al. (2022, see References) consolidated translation evidence from seven published\ ribosome profiling datasets that used harringtonine or lactimidomycin treatment to enrich\ for translation initiation sites. Ribo-seq reads were mapped to the GENCODE v35 annotation\ on GRCh38. Only ATG-initiated ORFs of at least 16 codons were retained, and redundant\ sense-overlapping ORFs were merged by taking the longest representative, yielding 7,264\ ncORFs across five biotype classes: upstream ORFs (uORFs), downstream ORFs (dORFs),\ intronic ORFs (intORFs), pseudogenic translations (PT), and lncRNA-embedded ORFs.\ The catalog was developed as part of a reference-database coordination effort involving\ Ensembl/GENCODE, HGNC, UniProtKB, and PeptideAtlas.\
\ \\ Chothani et al. (2025, see References) expanded the catalog by incorporating additional\ Ribo-seq datasets across more cell types and tissues and mapping to GENCODE v45. The\ ATG-start codon and 16-codon length restrictions were lifted to capture near-cognate\ initiations and micropeptides. A data-driven scoring framework using ribosome occupancy\ uniformity and P-site in-frame fraction identified a Primary subset of 10,127 ncORFs with\ translation signatures comparable to canonical coding genes; the Comprehensive set contains\ all 28,359 mapped ORFs.\
\ \\ Each ORF was annotated with its Kozak consensus strength by fetching the 11-base genomic\ context around the start codon from hg38.2bit and classifying positions −3 and +4\ relative to the A of the start codon: both matching (A/G at −3 and G at +4) =\ Strong; one matching = Moderate; neither = Weak; non-ATG start = non-ATG. A numeric\ translational efficiency (TE) score was also assigned by looking up the 11-base context in\ the Noderer 2014 TE table (Mol Syst Biol 10:748, PMID 25170020).\
\ \\ Deutsch et al. (2026, see References) queried the 7,264 Phase I ORFs against two independent\ PeptideAtlas mass spectrometry repositories. The HLA immunopeptidomics build was constructed\ from HLA-I and HLA-II peptidomes across more than 100 HLA-typed donors spanning multiple\ tissue types and cancer cell lines; peptides were enriched by affinity purification and\ identified by tandem mass spectrometry. The whole-cell tryptic proteomics build used\ conventional shotgun proteomics from a broad range of cell lines and tissues. Spectra\ were manually reviewed and classified according to the Prensner et al. tier system (Tier 1B =\ numerous HPP-quality HLA peptides; Tier 2B = a single qualifying HLA peptide; Tier 1A/2A =\ additional non-HLA evidence). The study introduced the term peptidein for a\ translation product detectable by mass spectrometry but not yet annotatable as a protein due\ to absent functional evidence. Of the 7,264 Phase I ORFs, 628 passed PeptideAtlas curation\ as peptideins (Table S12); a further 1,522 have HLA or tryptic peptide evidence below the\ peptidein threshold.\
\ \\
The supplementary data tables (Tables S2, S3, S6, S7, and S12) from Deutsch et al. were\
downloaded from the paper's supplementary materials at\
\
https://doi.org/10.1038/s41586-026-10459-x.\
Each table was joined to the Phase I bigGenePred by the short ORF identifier (e.g.,\
c14riboseqorf80) using the script\
addPeptideEvidence.py.\
The script appended 14 new fields to all 7,264 Phase I items; the 5,114 ORFs without peptide\
evidence receive default empty values so they remain visible in the track and filterable on\
isPeptidein and related fields. Non-HLA peptides that map to known proteins or\
are too short to be informative were excluded (Tables S2, exclude column).\
The complete build procedure is documented in\
ncOrfs.txt.\
\ Thanks to Jonathan Mudge, Jorge Ruiz-Orera, John Prensner, Sebastiaan van Heesch, and the\ GENCODE / TransCODE consortium for creating and maintaining these annotations.\
\ \\ Deutsch EW, Kok LW, Mudge JM, Valls CF, Jungreis I, Ruiz-Orera J, Sun Z, Kusebauch U, Fierro-Monti\ I, Abelin JG et al.\ \ Expanding the human proteome with microproteins and peptideins.\ Nature. 2026 May 6;.\ PMID: 42092140\
\ \\ Chothani S, Ruiz-Orera J, Tierney JAS, Clauwaert J, Deutsch EW, Alba MM, Aspden JL, Baranov PV,\ Bazzini AA, Bruford EA et al.\ \ An expanded reference catalog of translated open reading frames for biomedical research.\ bioRxiv. 2025 Jul 7;.\ PMID: 40672165; PMC: PMC12265627\
\ \\ Mudge JM, Ruiz-Orera J, Prensner JR, Brunet MA, Calvet F, Jungreis I, Gonzalez JM, Magrane M,\ Martinez TF, Schulz JF et al.\ \ Standardized annotation of translated open reading frames.\ Nat Biotechnol. 2022 Jul;40(7):994-999.\ PMID: 35831657; PMC: PMC9757701\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/ncOrfs/gencNcOrf/Ribo-seq_ORFs.comprehensive.kozak.bb\ filter.kozakTE -1:1.5\ filterByRange.kozakTE on\ filterLimits.kozakTE -1:1.5\ filterType.kozakStrength multipleListOr\ filterType.startCodon multipleListOr\ filterValues.kozakStrength Strong,Moderate,Weak,non-ATG,None\ filterValues.startCodon ATG,CTG,GTG,TTG,ACG,other,none\ html gencNcOrfs\ itemRgb on\ longLabel ncORFs: GENCODE Phase II non-canonical ORFs - comprehensive\ mouseOver $name in $geneName2 ($geneType)\ The three Gencode ncORF tracks in the non-canonical ORF track container show \ non-canonical translated open reading frames (ncORFs) identified\ from ribosome profiling (Ribo-seq) data and mapped to the GENCODE annotation by the\ GENCODE / TransCODE consortium.\ The data is available in two phases:\
\ \\ The Phase I catalog contains 7,264 unique human ncORFs called from Ribo-seq data\ across seven publications and mapped to GENCODE v35. Only translations of 16 codons or above\ and initiating from ATG start codons were incorporated. Redundant sense-overlapping ORFs were\ merged. Of these, 3,085 ORFs were found by more than one publication, providing independent\ replication evidence. This catalog was developed as part of an effort to standardize the\ annotation of translated ORFs across reference databases including Ensembl/GENCODE, HGNC,\ UniProtKB, and PeptideAtlas.\
\ \\ The Phase II catalog nearly quadruples the Phase I set, defining 28,359 ncORFs in the\ Comprehensive set, mapped to GENCODE v45. Compared to Phase I, additional published\ Ribo-seq datasets were incorporated and the restrictions on ORF size and initiation codon\ were lifted.\
\ \\ Two subsets are provided for the Phase II data:\
\\ All three GENCODE ncORF tracks are displayed in bigGenePred format and labeled with their\ ORF identifier. The default color scheme and available filter controls differ by track.\
\ \\ The Phase I track colors items by Kozak consensus strength by default.\ Two alternative color schemes can be selected from the track controls page\ (Color by dropdown): Evidence type and HLA class (see below).\
\ \| \ | Golden amber — Strong Kozak context. Both position −3 (A/G) and\ position +4 (G) match the consensus. | \
|---|---|
| \ | Steel blue — Moderate Kozak context. One of the two positions matches. | \
| \ | Gray — Weak Kozak context. Neither position matches. | \
| \ | Black — Non-ATG start codon (Kozak rule does not apply) or context\ unavailable. | \
\ Select Color by: Evidence type to highlight peptide evidence from\ Deutsch et al. (see References). ORFs with no mass spectrometry evidence are gray.\
\ \| \ | Gold — TransCODE peptidein (628 ORFs). Confirmed as\ confidently translated by PeptideAtlas; candidate for peptidein annotation in\ reference databases. | \
|---|---|
| \ | Steel blue — HLA immunopeptidomics evidence only (1,373 ORFs). | \
| \ | Forest green — Non-HLA (whole-cell tryptic) evidence only (66 ORFs). | \
| \ | Orange — Both HLA and non-HLA evidence (35 ORFs). | \
| \ | Gray — No peptide evidence (5,114 ORFs) or in the peptidein set\ based on binding predictions only with no direct MS sequences (48 ORFs). | \
\ Select Color by: HLA class to color items by the HLA class in which peptides were\ detected:\
\ \| \ | Steel blue — Class I only (1,632 ORFs). | \
|---|---|
| \ | Crimson — Class II only (10 ORFs). | \
| \ | Orange — Both class I and class II (143 ORFs). | \
| \ | Gray — No HLA data (5,479 ORFs). | \
\ The Phase I track can be filtered by: start codon, Kozak strength, Kozak TE, replicated\ status, and — using the peptide evidence fields — peptidein status\ (isPeptidein), HLA class (hlaClass), HLA evidence tier\ (hlaFinalTier), HPP guideline category (hlaHppCategory), and Ribo-seq\ quality (riboseqQuality).\
\ \\ Mouseover for Phase I shows ORF name, host gene, Kozak strength and TE,\ replicated status, peptidein flag, HLA evidence tier, and HLA peptide count.\
\ \\ The Phase II Primary and Comprehensive tracks color items by Kozak strength using the\ same scheme as Phase I. Peptide evidence fields are not included in the Phase II tracks.\ Common filters: start codon, Kozak strength, Kozak TE.\
\ \\ Each Phase I item carries the following peptide evidence fields from Deutsch et al.\ (2026), accessible via the details page and Table Browser:\
\ \| Field | Description |
|---|---|
| isPeptidein | yes/no: ORF is in the PeptideAtlas peptidein set (Table S12) |
| hlaClass | HLA class(es) detected: I, II, or Both |
| hlaFinalTier | HLA evidence tier (Tier 1B = numerous peptides; Tier 2B = one peptide) |
| hlaHppCategory | HPP guideline category (HPP+, 1PepCandidate, Insufficient) |
| hlaNPeptides | Number of distinct HLA peptide sequences detected |
| riboseqQuality | Manual quality of Ribo-seq evidence (Excellent/Sufficient/Insufficient) |
| hlaIPeptides | HLA class I peptide sequences (comma-separated) |
| hlaIIPeptides | HLA class II peptide sequences (comma-separated) |
| nonHlaFinalTier | Non-HLA (tryptic proteomics) evidence tier |
| nonHlaHppCategory | Non-HLA HPP guideline category |
| nonHlaNPeptides | Number of distinct non-HLA peptide sequences |
| nonHlaPeptides | Non-HLA tryptic peptide sequences (comma-separated) |
\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. The data can be accessed from\ scripts through our API; the track names are\ "gencNcOrfs" (Phase I), "gencNcOrfsPrimary" (Phase II Primary),\ and "gencNcOrfsComprehensive" (Phase II Comprehensive).\
\ \\ For automated download and analysis, the genome annotations are stored in bigBed files that\ can be downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\
\ \\ Mudge et al. (2022, see References) consolidated translation evidence from seven published\ ribosome profiling datasets that used harringtonine or lactimidomycin treatment to enrich\ for translation initiation sites. Ribo-seq reads were mapped to the GENCODE v35 annotation\ on GRCh38. Only ATG-initiated ORFs of at least 16 codons were retained, and redundant\ sense-overlapping ORFs were merged by taking the longest representative, yielding 7,264\ ncORFs across five biotype classes: upstream ORFs (uORFs), downstream ORFs (dORFs),\ intronic ORFs (intORFs), pseudogenic translations (PT), and lncRNA-embedded ORFs.\ The catalog was developed as part of a reference-database coordination effort involving\ Ensembl/GENCODE, HGNC, UniProtKB, and PeptideAtlas.\
\ \\ Chothani et al. (2025, see References) expanded the catalog by incorporating additional\ Ribo-seq datasets across more cell types and tissues and mapping to GENCODE v45. The\ ATG-start codon and 16-codon length restrictions were lifted to capture near-cognate\ initiations and micropeptides. A data-driven scoring framework using ribosome occupancy\ uniformity and P-site in-frame fraction identified a Primary subset of 10,127 ncORFs with\ translation signatures comparable to canonical coding genes; the Comprehensive set contains\ all 28,359 mapped ORFs.\
\ \\ Each ORF was annotated with its Kozak consensus strength by fetching the 11-base genomic\ context around the start codon from hg38.2bit and classifying positions −3 and +4\ relative to the A of the start codon: both matching (A/G at −3 and G at +4) =\ Strong; one matching = Moderate; neither = Weak; non-ATG start = non-ATG. A numeric\ translational efficiency (TE) score was also assigned by looking up the 11-base context in\ the Noderer 2014 TE table (Mol Syst Biol 10:748, PMID 25170020).\
\ \\ Deutsch et al. (2026, see References) queried the 7,264 Phase I ORFs against two independent\ PeptideAtlas mass spectrometry repositories. The HLA immunopeptidomics build was constructed\ from HLA-I and HLA-II peptidomes across more than 100 HLA-typed donors spanning multiple\ tissue types and cancer cell lines; peptides were enriched by affinity purification and\ identified by tandem mass spectrometry. The whole-cell tryptic proteomics build used\ conventional shotgun proteomics from a broad range of cell lines and tissues. Spectra\ were manually reviewed and classified according to the Prensner et al. tier system (Tier 1B =\ numerous HPP-quality HLA peptides; Tier 2B = a single qualifying HLA peptide; Tier 1A/2A =\ additional non-HLA evidence). The study introduced the term peptidein for a\ translation product detectable by mass spectrometry but not yet annotatable as a protein due\ to absent functional evidence. Of the 7,264 Phase I ORFs, 628 passed PeptideAtlas curation\ as peptideins (Table S12); a further 1,522 have HLA or tryptic peptide evidence below the\ peptidein threshold.\
\ \\
The supplementary data tables (Tables S2, S3, S6, S7, and S12) from Deutsch et al. were\
downloaded from the paper's supplementary materials at\
\
https://doi.org/10.1038/s41586-026-10459-x.\
Each table was joined to the Phase I bigGenePred by the short ORF identifier (e.g.,\
c14riboseqorf80) using the script\
addPeptideEvidence.py.\
The script appended 14 new fields to all 7,264 Phase I items; the 5,114 ORFs without peptide\
evidence receive default empty values so they remain visible in the track and filterable on\
isPeptidein and related fields. Non-HLA peptides that map to known proteins or\
are too short to be informative were excluded (Tables S2, exclude column).\
The complete build procedure is documented in\
ncOrfs.txt.\
\ Thanks to Jonathan Mudge, Jorge Ruiz-Orera, John Prensner, Sebastiaan van Heesch, and the\ GENCODE / TransCODE consortium for creating and maintaining these annotations.\
\ \\ Deutsch EW, Kok LW, Mudge JM, Valls CF, Jungreis I, Ruiz-Orera J, Sun Z, Kusebauch U, Fierro-Monti\ I, Abelin JG et al.\ \ Expanding the human proteome with microproteins and peptideins.\ Nature. 2026 May 6;.\ PMID: 42092140\
\ \\ Chothani S, Ruiz-Orera J, Tierney JAS, Clauwaert J, Deutsch EW, Alba MM, Aspden JL, Baranov PV,\ Bazzini AA, Bruford EA et al.\ \ An expanded reference catalog of translated open reading frames for biomedical research.\ bioRxiv. 2025 Jul 7;.\ PMID: 40672165; PMC: PMC12265627\
\ \\ Mudge JM, Ruiz-Orera J, Prensner JR, Brunet MA, Calvet F, Jungreis I, Gonzalez JM, Magrane M,\ Martinez TF, Schulz JF et al.\ \ Standardized annotation of translated open reading frames.\ Nat Biotechnol. 2022 Jul;40(7):994-999.\ PMID: 35831657; PMC: PMC9757701\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/ncOrfs/gencNcOrf/Ribo-seq_ORFs.primary.kozak.bb\ filter.kozakTE -1:1.5\ filterByRange.kozakTE on\ filterLimits.kozakTE -1:1.5\ filterType.kozakStrength multipleListOr\ filterType.startCodon multipleListOr\ filterValues.kozakStrength Strong,Moderate,Weak,non-ATG,None\ filterValues.startCodon ATG,CTG,GTG,TTG,ACG,other,none\ html gencNcOrfs\ itemRgb on\ longLabel ncORFs: GENCODE Phase II non-canonical ORFs - primary\ mouseOver $name in $geneName2 ($geneType)\ The aim of the GENCODE \ Genes project (Harrow et al., 2006) is to produce a set of \ highly accurate annotations of evidence-based gene features on the human reference genome.\ This includes the identification of all protein-coding loci with associated\ alternative splice variants, non-coding with transcript evidence in the public \ databases (NCBI/EMBL/DDBJ) and pseudogenes. A high quality set of gene\ structures is necessary for many research studies such as comparative or \ evolutionary analyses, or for experimental design and interpretation of the \ results.
\\ The GENCODE Genes tracks display the high-quality manual annotations merged \ with evidence-based automated annotations across the entire\ human genome. The GENCODE gene set presents a full merge\ between HAVANA manual annotation and Ensembl automatic annotation.\ Priority is given to the manually curated HAVANA annotation using predicted\ Ensembl annotations when there are no corresponding manual annotations. With \ each release, there is an increase in the number of annotations that have undergone\ manual curation. \ This annotation was carried out on the GRCh38 (hg38) genome assembly.\
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ These are multi-view composite tracks that contain differing data sets\ (views). Instructions for configuring multi-view tracks are\ here.\ Only some subtracks are shown by default. The user can select which subtracks\ are displayed via the display controls on the track details pages.\ Further details on display conventions and data interpretation are available in the track descriptions.
\ \\ GENCODE Genes and its associated tables can be explored interactively using the\ REST API, the\ Table Browser or the\ Data Integrator.\ The GENCODE data files for hg38 are available in our\ \ downloads directory as wgEncodeGencode* files in genePred format.\ All the tables can also be queried directly from our public MySQL\ servers, with instructions on this method available on our\ MySQL help page as well as on\ our blog.
\ \\ GENCODE version 49\ corresponds to Ensembl 115.\
\\ GENCODE version 48\ corresponds to Ensembl 114.\
\ GENCODE version 47\ corresponds to Ensembl 113.\ \\ GENCODE version 46\ corresponds to Ensembl 112.\
\\ GENCODE version 45\ corresponds to Ensembl 111.\
\\ GENCODE version 44\ corresponds to Ensembl 110.\
\\ GENCODE version 43\ corresponds to Ensembl 109.\
\\ GENCODE version 42\ corresponds to Ensembl 108.\
\\ GENCODE version 41\ corresponds to Ensembl 107.\
\\ GENCODE version 40\ corresponds to Ensembl 106.\
\\ GENCODE version 39\ corresponds to Ensembl 105.\
\\ GENCODE version 38\ corresponds to Ensembl 104.\
\\ GENCODE version 37\ corresponds to Ensembl 103.\
\\ GENCODE version 36\ corresponds to Ensembl 102.\
\\ GENCODE version 35\ corresponds to Ensembl 101.\
\\ GENCODE version 34\ corresponds to Ensembl 100.\
\\ GENCODE version 33\ corresponds to Ensembl 99.\
\\ GENCODE version 30\ corresponds to Ensembl 96.\
\\ GENCODE version 29\ corresponds to Ensembl 94.\
\\ GENCODE version 28\ corresponds to Ensembl 92.\
\\ GENCODE version 27\ corresponds to Ensembl 90.\
\\ GENCODE version 26\ corresponds to Ensembl 88.\
\\ GENCODE version 24\ corresponds to Ensembl 84.\
\ GENCODE version 23\ corresponds to Ensembl 81.\ \ GENCODE version 22\ corresponds to Ensembl 79.\ \ GENCODE version 20\ corresponds to Ensembl 76.\ \\ See also: The GENCODE Project Release History.\
\ \The GENCODE project is an international collaboration funded by NIH/NHGRI\ grant U41HG007234. More information is available\ at www.gencodegenes.org,\ Participating GENCODE institutions and personnel can be found\ \ here.\
\ \\ Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong\ J, Barnes I et al.\ \ GENCODE 2021.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D916-D923.\ PMID: 33270111;\ PMC: PMC7778937;\ DOI: 10.1093/nar/gkaa1087\
\ \ \A full list of GENCODE publications are available\ at The GENCODE\ Project web site.\
\ \GENCODE data are available for use without restrictions.
\ \ genes 0 group genes\ longLabel Container of all new and previous GENCODE releases\ shortLabel GENCODE Versions\ superTrack on\ track wgEncodeGencodeSuper\ trackHandler wgEncodeGencode\ interactions Gene Interactions bigBed 9 Protein Interactions from Curated Databases and Text-Mining 0 100 0 0 0 127 127 127 0 0 0\ The Pathways and Gene Interactions track shows a summary of gene interaction and pathway data\ collected from two sources: curated pathway/protein-interaction databases and interactions found\ through text mining of PubMed abstracts.
\ \\ The track features a single item for each gene loci in the genome. On the item itself, the gene\ symbol for the loci is displayed followed by the top gene interactions noted by their gene symbol.\ Clicking an item will take you a\ gene interaction graph\ that includes detailed information on the support for the various interactions.
\ \\ Items are colored based on the number of documents supporting the interactions of a\ particular gene. Genes with >100 supporting documents are colored\ black, genes with >10 but <100\ supporting documents are colored dark blue, and\ those with >10 supporting documents are colored\ light blue.
\ \\ See the\ help documentation\ accompanying this gene interaction graph for more information on its configuration.
\ \\ The pathways and gene interactions were imported from a number of databases and mined from\ millions of PubMed abstracts. More information can be found in the\ "Data Sources\ and Methods"\ section of the help page for the gene interaction graph.
\ \\ The underlying data for this track can be accessed interactively through the\ Table Browser or\ Data Integrator. \ The data for this track is spread across a number of relational tables. The best way to \ export or analyze the data is using our public MySQL server.\ The list of tables and how they are linked together are described in the \ documentation \ linked at the bottom of the gene interaction viewer.\
\ \\ The genome annotation is just a summary of the actual interactions database and therefore often not \ of interest to most users. It is stored in a bigBed file that can be obtained\ from the\ download server.\ \ The data underlying the\ graphical display is in bigBed\ formatted file named interactions.bb. Individual regions or the whole genome annotation\ can be obtained using our tool bigBedToBed. Instructions\ for downloading source code and precompiled binaries can be found\ here. The tool can also\ be used to obtain only features within a given range, for example:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/interactions.bb\ -chrom=chr6 -start=0 -end=1000000 stdout\
\ \\ The text-mined data for the gene interactions and pathways were generated by Chris Quirk and\ Hoifung Poon as part of\ Microsoft Research, Project\ Hanover.
\ \\ Pathway data was provided by the databases listed under\ "Data Sources\ and Methods"\ section of the help page for the gene interaction graph.\ In particular, thank you to Ian Donaldson from IRef for his\ unique collection of interaction databases.
\ \\ The short gene descriptions are a merge of the HPRD\ and PantherDB gene/molecule classifications. Thanks to Arun Patil from\ HPRD for making them available as a download.
\ \\ The track display and gene interaction graph\ were developed at the UCSC Genome Browser by Max Haeussler.
\ \\ Poon H, Quirk C, DeZiel C, Heckerman D.\ Literome: PubMed-scale genomic knowledge base in the cloud\ Bioinformatics. 2014 Oct;30(19):2840-2.\ PMID: 24939151
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/interactions.bb\ directUrl hgGeneGraph?db=hg38&gene=%s\ exonNumbers off\ group phenDis\ hgsid on\ itemRgb on\ labelOnFeature on\ linkIdInName on\ longLabel Protein Interactions from Curated Databases and Text-Mining\ noScoreFilter on\ shortLabel Gene Interactions\ track interactions\ type bigBed 9\ visibility hide\ ghGeneTss Gene TSS bigBed 9 GeneHancer Regulatory Elements and Gene Interactions 3 100 0 0 0 127 127 127 0 0 0 http://www.genecards.org/cgi-bin/carddisp.pl?gene=$$ regulation 1 itemRgb on\ longLabel GeneHancer Regulatory Elements and Gene Interactions\ parent geneHancer\ searchIndex name\ shortLabel Gene TSS\ track ghGeneTss\ type bigBed 9\ url http://www.genecards.org/cgi-bin/carddisp.pl?gene=$$\ urlLabel In GeneCards:\ view b_TSS\ visibility pack\ geneHancer GeneHancer bed 3 GeneHancer Regulatory Elements and Gene Interactions 0 100 0 0 0 127 127 127 0 0 0\ GeneHancer is a database of human regulatory elements (enhancers and promoters) \ and their inferred target genes, which is embedded \ in GeneCards, a human gene \ compendium.\ The GeneHancer database was created by integrating >1 million regulatory elements \ from multiple genome-wide databases. \ Associations between the regulatory elements and target genes\ were based on multiple sources of linking molecular data, along with distance,\ as described in Methods below.\
\\ The GeneHancer track set contains tracks representing:\
\ Each GeneHancer regulatory element is identified by a GeneHancer id. \ For example: GH0XJ101383 is located on chromosome X, with starting position of 101,383 kb\ (GRCh38/hg38 reference).\ Based on the id, one can obtain full GeneHancer information, as displayed in the Genomics \ section within the gene-centric web pages of GeneCards. Links to the GeneCards information pages\ are provided on the track details pages.
\ \\ For the interaction tracks (Clusters and Interactions) a slight offset can be noticed between \ the line endpoints. This helps to identify the start and end of the feature. In this case,\ the higher point is the source (enhancers) and the lower point is the target.
\ \\ Colors are used to distinguish promoters and enhancers and to indicate the GeneHancer element confidence score:
\\ Promoters: \ High\ Medium\ Low\
\\ Enhancers: \ High\ Medium\ Low\
\ \\ Colors are used to improve gene and interactions visibility. \ Successive genes are colored in different colors, and interactions of a gene have the same color.
\ \\ The Interactions view in Full mode shows GeneHancers and target genes connected by curves or \ half-rectangles (when one of the connected regions is off-screen). \ Configuration options are available to change the drawing style, and to limit the view to\ interactions with one or both connected items in the region.\ Interactions are identified on mouseover or clicked on for details at the end regions, or at \ the curve peak, which is marked with a gray ring shape. Interactions in the reverse direction\ (Gene TSS precedes GeneHancer on the genome) are drawn with a dashed line.\ \
\ The Clusters view groups interactions by target gene; the target gene and all GeneHancers \ associated with it are displayed in a single browser item. The gene TSS and associated GeneHancers \ are shown as blocks linked together, with the TSS drawn as a "tall" item, and the \ GeneHancers drawn "short". \ A user configuration option is provided to change the view to group by GeneHancer \ (with tall GeneHancer and short TSS's). \ Clusters composed of interactions with a single gene are colored to correspond to the gene, \ and those composed of interactions with multiple genes are colored dark gray.
\ \\ GeneHancer identifications were created from >1 million regulatory elements \ obtained from seven genome-wide databases:\
\ Employing an integration algorithm that removes redundancy, the GeneHancer pipeline\ identified ˜250k integrated candidate regulatory elements (GeneHancers).\ Each GeneHancer is assigned an annotation-derived confidence score. \ The GeneHancers that are derived from more than one information source are defined \ as "elite" GeneHancers.
\\ Gene-GeneHancer associations, and their likelihood-based scores, were generated \ using information that helps link regulatory elements to genes:\
\ Associations that are derived from more than one information source are defined \ as "elite" associations, which leads to the definition of the "double elite"\ dataset - elite gene associations of elite GeneHancers.
\\ More details are provided at the GeneCards\ \ information page.\ For a full description of the methods used, refer to the GeneHancer manuscript1.
\\ Source data for the GeneHancer version 4.8 was downloaded during May 2018.
\ \\ Due to our agreement with the Weizmann Institute, we cannot allow full genome \ queries from the Table Browser or share download files. You can still access \ data for individual chromosomes or positional data from the \ Table Browser.
\ \\ GeneHancer is the property of the Weizmann Institute of Science and \ is not available for download or mirroring by any third party \ without permission. Please contact the Weizmann Institute directly for \ data inquiries.
\ \\ Thanks to Simon Fishilevich, Marilyn Safran, Naomi Rosen, and Tsippi Iny Stein of the GeneCards \ group and Shifra Ben-Dor of the Bioinformatics Core group at the Weizmann Institute, \ for providing this data and documentation, creating track hub versions of these tracks \ as prototypes, and overall responsiveness during development of these tracks.
\\ Contact: \ simon.\ fishilevich@weizmann.\ ac.\ il\ \
\ Supported in part by a grant from LifeMap Sciences Inc.
\ \\ Fishilevich S., Nudel R., Rappaport N., Hadar R., Plaschkes I., Iny Stein T., Rosen N., Kohn A., Twik M., Safran M., Lancet D. and Cohen D. GeneHancer: genome-wide integration of enhancers and target genes in GeneCards, Database (Oxford) (2017), doi:10.1093/database/bax028. [PDF] PMID 28605766
\\ Stelzer G, Rosen R, Plaschkes I, Zimmerman S, Twik M, Fishilevich S, Iny Stein T, Nudel R, Lieder I, Mazor Y, Kaplan S, Dahary, D, Warshawsky D, Guan- Golan Y, Kohn A, Rappaport N, Safran M, and Lancet D. The GeneCards Suite: From Gene Data Mining to Disease Genome Sequence Analysis, Current Protocols in Bioinformatics (2016), 54:1.30.1-1.30.33. doi: 10.1002/cpbi.5. PMID 27322403
\ regulation 1 compositeTrack on\ dataVersion January 2019 (V2: Corrections to Experiment field)\ dimensions dimX=set dimY=view\ group regulation\ longLabel GeneHancer Regulatory Elements and Gene Interactions\ shortLabel GeneHancer\ sortOrder set=+ view=+\ subGroup1 view View a_GH=Regulatory_Elements b_TSS=Gene_TSS c_I=Interactions d_I=Clusters\ subGroup2 set Set a_ELITE=Double_Elite b_ALL=All\ tableBrowser noGenome\ track geneHancer\ type bed 3\ visibility hide\ geneid Geneid Genes genePred geneidPep Geneid Gene Predictions 0 100 0 90 100 127 172 177 0 0 0\ This track shows gene predictions from the\ geneid program developed by\ Roderic Guigó's Computational Biology of RNA Processing\ group which is part of the\ Centre de Regulació Genòmica\ (CRG) in Barcelona, Catalunya, Spain.\
\ \\ Geneid is a program to predict genes in anonymous genomic sequences designed\ with a hierarchical structure. In the first step, splice sites, start and stop\ codons are predicted and scored along the sequence using Position Weight Arrays\ (PWAs). Next, exons are built from the sites. Exons are scored as the sum of the\ scores of the defining sites, plus the log-likelihood ratio of a\ Markov Model for coding DNA. Finally, from the set of predicted exons, the gene\ structure is assembled, maximizing the sum of the scores of the assembled exons.\
\ \\ Thanks to Computational Biology of RNA Processing\ for providing these data.\ \
\ \\ Blanco E, Parra G, Guigó R.\ Using geneid to identify genes.\ Curr Protoc Bioinformatics. 2007 Jun;Chapter 4:Unit 4.3.\ PMID: 18428791\
\ \ \\ Parra G, Blanco E, Guigó R.\ \ GeneID in Drosophila.\ Genome Res. 2000 Apr;10(4):511-5.\ PMID: 10779490; PMC: PMC310871\
\ genes 1 color 0,90,100\ group genes\ html ../../geneid\ longLabel Geneid Gene Predictions\ parent genePredArchive\ shortLabel Geneid Genes\ track geneid\ type genePred geneidPep\ visibility hide\ geneReviews GeneReviews bigBed 9 + GeneReviews 0 100 0 80 0 127 167 127 0 0 0 https://www.ncbi.nlm.nih.gov/books/NBK1116/?term=$$\ GeneReviews is an online collection of expert-authored, peer-reviewed\ articles that describe specific gene-related diseases. GeneReviews articles are\ searchable by disease name, gene symbol, protein name, author, or title. GeneReviews\ is supported by the National Institutes of Health, hosted at NCBI as part of the\ \ Genetic Testing Registry (GTR). The GeneReviews data underlying this track will be updated frequently. \
\ \The GeneReviews track allows the user to locate the NCBI GeneReviews resource\ quickly from the Genome Browser. Hovering the mouse on track items shows the gene symbol and \ associated diseases. A condensed version of the GeneReviews article\ name and its related diseases are displayed on the item details page as links. Similar\ information, when available, is provided in the details page of items from the UCSC Genes,\ RefSeq Genes, and OMIM Genes tracks.\
\ \\ The raw data for the GeneReviews track can be explored interactively with the\ Table Browser. Cross-referencing can be done with\ Data Integrator. The complete source file,\ in bigBed format, \ can be downloaded from our\ downloads directory.\ For automated analysis,\ the data may be queried from our\ REST API.\
\ \\ Previous versions of this track can be found on our archive download server.\
\ \\ Pagon RA, Adam MP, Bird TD, et al., editors. GeneReviews® [Internet]. Seattle (WA): University of Washington, Seattle; 1993-2014. Available from: \ \ https://www.ncbi.nlm.nih.gov/books/NBK1116.\
\ \ phenDis 1 bigDataUrl /gbdb/hg38/geneReviews/geneReviews.bb\ color 0, 80, 0\ group phenDis\ html geneReviews\ longLabel GeneReviews\ mouseOver Gene: $name\ The tracks listed here contain data from\ The Genome in a\ Bottle Consortium (GIAB), an open, public consortium hosted by \ NIST. The priority of GIAB is to develop \ reference standards, reference methods, and reference data by authoritative characterization of \ human genomes for use in benchmarking, including analytical validation and technology \ development that will support translation of whole human genome sequencing to clinical practice. The\ sole purpose of this work is to provide validated variants and regions to enable technology and \ bioinformatics developers to benchmark and optimize their detection methods.\
\\ The Ashkenazim and the Chinese Trio tracks show benchmark SNV calls from two \ son/father/mother trios of Ashkenazi Jewish and Han Chinese ancestry from the \ Personal Genome Project, \ consented for commercial redistribution.\
\\ The Genome In a Bottle Structural Variants track shows benchmark SV calls (nssv) \ and variant regions (nsv) (5,262 insertions and 4,095 deletions, > 50 bp, in 2.51 Gb of \ the genome) from the son (HG002/NA24385) from the Ashkenazi Jewish trio.\
\\ Samples are disseminated as National Institute of Standards and Technology (NIST)\ Reference Materials.\
\\ Unlike a regular genome browser track, the Ashkenazim and the Chinese Trio tracks display \ the genome variants of each individual as two haplotypes; SNPs, small insertions and deletions\ are mapped to each haplotype based on the phasing information of the VCF file. The\ haplotype 1 and the haplotype 2 are displayed as two separate black lanes for the\ browser window region. Each variant is drawn as a vertical dash. Homozygous variants will\ show two identical dashes on both haplotype lanes. Phased heterozygous variants are placed on\ one of the haplotype lanes and unphased heterozygous variants are displayed in the area\ between the two haplotype lanes.\
\\ Predicted de novo variants and variants that are inconsistent with phasing in the trio son can be \ colored in red using the track Configuration options.\
\ \\ Benchmark VCF and BED files for small variants are available for GRCh37 and GRCh38 under each\ genome at NCBI FTP site. \ Structural variants are available for GRCh37 at dbVAR \ nst175.\
\ \\ Zook JM, McDaniel J, Olson ND, Wagner J, Parikh H, Heaton H, Irvine SA, Trigg L, Truty R, McLean CY\ et al.\ \ An open resource for accurately benchmarking small variant and reference calls.\ Nat Biotechnol. 2019 May;37(5):561-566.\ PMID: 30936564; PMC: PMC6500473\
\ \\ Zook JM, Hansen NF, Olson ND, Chapman L, Mullikin JC, Xiao C, Sherry S, Koren S, Phillippy AM,\ Boutros PC et al.\ \ A robust benchmark for detection of germline large deletions and insertions.\ Nat Biotechnol. 2020 Jun 15;.\ PMID: 32541955\
\ \ varRep 1 compositeTrack on\ group varRep\ html giab\ longLabel Genome In a Bottle Structural Variants and Trios\ shortLabel Genome In a Bottle\ subGroup1 view Views trios=Trios sv=Structural_Variants\ track giab\ type bed 3\ visibility hide\ triosView Genome In a Bottle Trios vcfPhasedTrio Genome in a Bottle Ashkenazim and Chinese Trios 0 100 0 0 0 127 127 127 0 0 0 varRep 0 longLabel Genome in a Bottle Ashkenazim and Chinese Trios\ parent giab\ shortLabel Genome In a Bottle Trios\ track triosView\ type vcfPhasedTrio\ view trios\ visibility hide\ genscan Genscan Genes genePred genscanPep Genscan Gene Predictions 0 100 170 100 0 212 177 127 0 0 0\ This track shows predictions from the\ Genscan program\ written by Chris Burge.\ The predictions are based on transcriptional, translational and donor/acceptor\ splicing signals as well as the length and compositional distributions of exons,\ introns and intergenic regions.\
\ \\ For more information on the different gene tracks, see our Genes FAQ.
\ \\ This track follows the display conventions for\ gene prediction\ tracks.\
\ \\ The track description page offers the following filter and configuration\ options:\
\ For a description of the Genscan program and the model that underlies it,\ refer to Burge and Karlin (1997) in the References section below.\ The splice site models used are described in more detail in Burge (1998)\ below.\
\ \\ Burge C.\ Modeling Dependencies in Pre-mRNA Splicing Signals.\ In: Salzberg S, Searls D, Kasif S, editors.\ Computational Methods in Molecular Biology.\ Amsterdam: Elsevier Science; 1998. p. 127-163.\
\ \\ Burge C, Karlin S.\ \ Prediction of complete gene structures in human genomic DNA.\ J. Mol. Biol. 1997 Apr 25;268(1):78-94.\ PMID: 9149143\
\ genes 1 color 170,100,0\ group genes\ html ../../genscan\ longLabel Genscan Gene Predictions\ parent genePredArchive\ shortLabel Genscan Genes\ track genscan\ type genePred genscanPep\ visibility hide\ encTfChipPkENCFF567NFS GM12878 CUX1 narrowPeak Transcription Factor ChIP-seq Peaks of CUX1 in GM12878 from ENCODE 3 (ENCFF567NFS) 0 100 85 152 255 170 203 255 0 0 0 regulation 1 color 85,152,255\ longLabel Transcription Factor ChIP-seq Peaks of CUX1 in GM12878 from ENCODE 3 (ENCFF567NFS)\ parent encTfChipPk off\ shortLabel GM12878 CUX1\ subGroups cellType=GM12878 factor=CUX1\ track encTfChipPkENCFF567NFS\ gnfAtlas2 GNF Atlas 2 expRatio GNF Expression Atlas 2 0 100 0 0 0 127 127 127 0 0 0This track shows expression data from the GNF Gene Expression\ Atlas 2. This contains two replicates each of 79 human\ tissues run over Affymetrix microarrays. \ By default, averages of related tissues are shown. Display all tissues\ by selecting "All Arrays" from the "Combine arrays" menu\ on the track settings page.\ As is standard with microarray data red indicates overexpression in the \ tissue, and green indicates underexpression. You may want to view gene\ expression with the Gene Sorter as well as the Genome Browser.
\ \\ Su AI, Wiltshire T, Batalov S, Lapp H, Ching KA, Block D, Zhang J, Soden R, Hayakawa M, Kreiman G\ et al.\ \ A gene atlas of the mouse and human protein-encoding transcriptomes.\ Proc Natl Acad Sci U S A. 2004 Apr 20;101(16):6062-7.\ PMID: 15075390; PMC: PMC395923\
\ expression 1 expDrawExons on\ expScale 4.0\ expStep 0.5\ expTable gnfHumanAtlas2MedianExps\ group expression\ groupings gnfHumanAtlas2Groups\ longLabel GNF Expression Atlas 2\ shortLabel GNF Atlas 2\ track gnfAtlas2\ type expRatio\ visibility hide\ gnomadVariants gnomAD Genome Aggregation Database (gnomAD) 0 100 0 0 0 127 127 127 0 0 0\ The Genome Aggregation Database\ (gnomAD) is a resource developed by an international coalition of investigators at the Broad\ Institute and collaborating institutions, with the goal of aggregating and harmonizing exome and\ whole-genome sequencing data from large-scale sequencing projects spanning disease-specific cohorts\ and population genetics studies. Individuals affected by severe pediatric diseases and first-degree\ relatives were excluded from the studies. However, some individuals with severe disease may still\ have remained in the datasets, although probably at an equivalent or lower frequency than observed\ in the general population. For each variant, gnomAD provides allele frequencies stratified by\ genetic ancestry group, alongside quality metrics such as depth of coverage and genotype quality\ scores. The database also supplies sequencing coverage, structural variants, CNVs, and short tandem\ repeats. Additionally, gnomAD provides non-coding constraint and gene-level constraint\ metrics — including pLI scores, observed/expected (oe) ratios, and LOEUF values —\ that quantify intolerance to loss-of-function variation and are widely used to prioritize\ candidate disease genes. The most\ current release on hg38 is v4.1, but the older v3 and v2 versions are also available.\
\ \\ The available data tracks are:\
\ For questions on the gnomAD data, also see the gnomAD FAQ.
\\ More details on the Variant type(s) can be found on the Sequence Ontology page.
\ \\ The raw data can be explored interactively with the \ Table Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API, and the genome annotations are stored in files that\ can be downloaded from our download server, subject\ to the conditions set forth by the gnomAD consortium (see below).
\ \\ The data can also be found directly from the gnomAD downloads page. Please refer to\ our mailing list archives for questions, or our Data Access FAQ for more information.
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the Creative Commons Zero Public Domain Dedication as described here.\
\ \\ Please note that some annotations within the provided files may have restrictions on usage. See here for more information.\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM, Ganna\ A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ Analysis of protein-coding\ genetic variation in 60,706 humans. Nature. 2016 Aug 17;536(7616):285-91.\ PMID: 27535533;\ PMC: PMC5018207\
\\ Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, Khera AV, Lowther C,\ Gauthier LD, Wang H et al.\ \ A structural variation reference for medical and population genetics.\ Nature. 2020 May;581(7809):444-451.\ PMID: 32461652; PMC: PMC7334194\
\\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alföldi J, Watts NA, Vittal C,\ Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664\
\ varRep 0 cartVersion 6\ group varRep\ html gnomad\ longLabel Genome Aggregation Database (gnomAD)\ pennantIcon New red ../goldenPath/newsarch.html#041026 "New gnomAD STR track added Apr. 10, 2026"\ shortLabel gnomAD\ superTrack on\ track gnomadVariants\ gnomadPLI gnomAD Constraint Metrics bigBed 12 Genome Aggregation Database (gnomAD) Predicted Constraint Metrics (LOEUF, pLI, and Z-scores) 0 100 0 0 0 127 127 127 0 0 0\ The Genome Aggregation Database (gnomAD) - Predicted Constraint Metrics track set contains\ metrics of pathogenicity per-gene as predicted for gnomAD v2.1.1, v4.0, or v4.1 and identifies genes subject to\ strong selection against various classes of mutation.\
\ \\ This track includes several subtracks of constraint metrics calculated at gene (canonical\ transcript) and transcript level. For more information see the following\ blog post.\ The metrics include:\
\ There are two "groups" of tracks in this set, and three gnomAD versions (v2.1.1, v4.0, and v4.1):\

\ Clicking the grey box to the left of the track, or right-clicking and choosing the Configure option,\ brings up the interface for filtering items based on their pLI score, or labeling the items\ based on their Ensembl identifier and/or Gene Name.\
\ \\ Please see the gnomAD browser help page and FAQ for further explanation of the topics below.
\ \\ Observed count: The number of unique single-nucleotide variants in each transcript/gene\ with 123 or fewer alternative alleles (MAF < 0.1%).\
\\ Expected count: A depth-corrected probability prediction model that takes into account\ sequence context, coverage, and methylation was used to predict expected\ variant counts. For more information please see Lek et al., 2016.\
\\ Variants found in exons with a median depth < 1 were removed from both counts.\
\ The O/E constraint score is the ratio of the observed/expected variants in that gene. Each item in\ this track shows the O/E ratio for three different types of variation: missense, synonymous, and\ loss-of-function. The O/E ratio is a continuous measurement of how tolerant a gene or\ transcript is to a certain class of variation. When a gene has a low O/E value, it is under stronger\ selection for that class of variation than a gene with a higher O/E value. Because Counts depend on\ gene size and sample size, the precision of the values varies a lot from one gene to the next. \ Therefore, the 90% confidence interval (CI) is also displayed along with the O/E ratio to better\ assist interpretation of the scores.\
\ When evaluating how constrained a gene is, it is essential to consider the CI when using O/E. In \ research and clinical interpretation of Mendelian cases, pLI > 0.9 has been widely used for \ filtering. Accordingly, the Gnomad team suggests using the upper bound of the O/E confidence interval\ LOEUF < 0.35 as a threshold if needed.\
\ Please see the Methods section below for more information about how the scores were calculated.\
\ \\ The pLI and Z-scores of the deviation of observed variant counts relative to the expected number \ are intended to measure how constrained or intolerant a gene or transcript is to a specific type of\ variation. Genes or transcripts that are particularly depleted of a specific class of variation\ (as observed in the gnomAD data set) are considered intolerant of that specific type of variation.\ Z-scores are available for the missense and synonymous categories and pLI scores are available for\ the loss-of-function variation.\
\\ Missense and Synonymous: Positive Z-scores indicate more constraint (fewer observed \ variants than expected), and negative scores indicate less constraint (more observed variants than\ expected). A greater Z-score indicates more intolerance to the class of variation. Z-scores\ were generated by a sequence-context-based mutational model that predicted the number of expected\ rare (< 1% MAF) variants per transcript. The square root of the chi-squared value of the \ deviation of observed counts from expected counts was multiplied by -1 if the observed count was\ greater than the expected and vice versa. For the synonymous score, each Z-score was corrected by\ dividing by the standard deviation of all synonymous Z-scores between -5 and 5. For the missense\ scores, a mirrored distribution of all Z-scores between -5 and 0 was created, and then all missense\ Z-scores were corrected by dividing by the standard deviation of the Z-score of the mirror\ distribution.\
\\ Loss-of-function: pLI closer to 1 indicates that the gene or transcript cannot tolerate\ protein truncating variation (nonsense, splice acceptor and splice donor variation). The gnomAD\ team recommends transcripts with a pLI >= 0.9 for the set of transcripts extremely intolerant\ to truncating variants. pLI is based on the idea that transcripts can be classified into three\ categories:\
\ Please see Samocha et al., 2014 and Lek et al., 2016 for further discussion of these metrics.\
\ \\ For version 2.1.1 only, the GENCODE transcripts were filtered according to the following criteria:\
\ For version v2.1.1, the gnomAD gene/transcript data is based on hg19. In order to map transcripts and genes to the hg38 genome the following steps were taken:\
\ For version v4.0 and v4.1, the gnomAD transcript data is based on hg38. In order to map the\ transcripts to hg38, the transcript version numbers in the gnomAD download file were joined with\ GENCODE V39 and NCBI RefSeq coordinates available at UCSC.\
\ \\ Per gene and per transcript data were downloaded from the gnomAD Google Storage bucket:\
\ gs://gnomad-public/release/2.1.1/constraint/gnomad.v2.1.1.lof_metrics.by_gene.txt.bgz\ gs://gnomad-public/release/2.1.1/constraint/gnomad.v2.1.1.lof_metrics.by_transcript.txt.bgz\\ These data were then joined to the Gencode set of genes/transcripts available at the UCSC\ Genome Browser (see previous section) and then transformed into a bigBed 12+5. For the full list of commands used to\ make this track please see the\ makedoc.\ \ \
\ Per gene and per transcript data were downloaded from the gnomAD Google Storage bucket:\
\ https://storage.googleapis.com/gcp-public-data--gnomad/release/4.0/constraint/gnomad.v4.0.constraint_metrics.tsv\\ These data were then joined to the Gencode/NCBI set of genes/transcripts available at the UCSC\ Genome Browser and then transformed into a bigBed 12+5. For the full list of commands used to\ make this track please see the\ makedoc.\ \ \
\ Per gene and per transcript data were downloaded from the gnomAD Google Storage bucket:\
\ https://storage.googleapis.com/gcp-public-data--gnomad/release/4.1/constraint/gnomad.v4.1.constraint_metrics.tsv\\ These data were then joined to the Gencode/NCBI set of genes/transcripts available at the UCSC\ Genome Browser and then transformed into a bigBed 12+5. For the full list of commands used to\ make this track please see the\ makedoc.\ \ \
\
The raw data can be explored interactively with the Table Browser, or\
the Data Integrator. For automated access, this track, like all \
others, is available via our API. However, for bulk \
processing, it is recommended to download the dataset. The genome annotation is stored in a bigBed \
file that can be downloaded from the\
download server. The exact\
filenames can be found in the track configuration file. Annotations can be converted to ASCII text\
by our tool bigBedToBed which can be compiled from the source code or downloaded as\
a precompiled binary for your system. Instructions for downloading source code and binaries can be\
found here. The tool\
can also be used to obtain only features within a given range, for example:
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gnomAD/pLI/pliByTranscript.bb -chrom=chr6 -start=0 -end=1000000 stdout\\
\ Please refer to our\ mailing list archives\ for questions and example queries, or our\ Data Access FAQ\ for more information.\
\ \\ More information about using and understanding the gnomAD data can be found in the\ gnomAD FAQ site.\
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the ODC Open Database License\ (ODbL) as described here.\
\ \\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ \ Analysis of protein-coding genetic variation in 60,706 humans.\ Nature. 2016 Aug 18;536(7616):285-91.\ PMID: 27535533; PMC: PMC5018207\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM,\ Ganna A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\ \\ Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, Khera AV, Lowther C,\ Gauthier LD, Wang H et al.\ \ A structural variation reference for medical and population genetics.\ Nature. 2020 May;581(7809):444-451.\ PMID: 32461652; PMC: PMC7334194\
\ \\ Cummings BB, Karczewski KJ, Kosmicki JA, Seaby EG, Watts NA, Singer-Berk M, Mudge JM, Karjalainen J,\ Satterstrom FK, O'Donnell-Luria AH et al.\ \ Transcript expression-aware annotation improves rare variant interpretation.\ Nature. 2020 May;581(7809):452-458.\ PMID: 32461655; PMC: PMC7334198\
\ \ varRep 1 compositeTrack On\ dataVersion Release v4.1 (April 19, 2024), Release v4 (November 2023), Release 2.1.1 (March 6, 2019)\ group varRep\ html gnomadPLI.html\ labelFields name,geneName\ longLabel Genome Aggregation Database (gnomAD) Predicted Constraint Metrics (LOEUF, pLI, and Z-scores)\ parent gnomadVariants\ shortLabel gnomAD Constraint Metrics\ subGroup1 view Views v2=constraintV2 v4=constraintV4 v4_1=constraintV4.1\ track gnomadPLI\ type bigBed 12\ visibility hide\ gnomadPext gnomAD pext bigWig 0 1 Genome Aggregation Database (gnomAD) Proportion Expression Across Transcript Scores (pext) 0 100 0 0 0 127 127 127 0 0 0\ The Genome Aggregation Database (gnomAD) Proportion Expression Across Transcript Scores (pext) track set displays isoform expression levels across 50\ tissues from the Genotype Tissue Expression (GTEx) v10 dataset; tissues with fewer than 50 samples\ were excluded (Fallopian Tube, Endocervix, Ectocervix, Kidney, Medulla).\
\ \\ The gnomAD pext tracks provide a comprehensive view of the expression of exons across a \ gene using the proportion expression across transcripts, or pext metric, a \ transcript-level annotation metric that quantifies isoform expression for variants. This metric \ was calculated by annotating each variant with the expression of all possible consequences across \ all transcripts for each tissue and normalizing the expression of the annotation to the total \ expression of the gene, which can be interpreted as a measure of the proportion of the total \ transcriptional output from a gene that would be affected by the variant annotation in question.\ More information can be found on the Broad institute's pext help page \
\ \\ Each of the subtracks shows the pext metric for a specific tissue, except the gnomAD \ pext Mean Proportion subtrack that shows the average pext metrics calculated from the 50 GTEx \ tissues.\
\ \\ The pext graphs display the mean expression at each base position for protein-coding (CDS) regions.\ While UTRs do have expression in transcriptome datasets, this information is not included\ for the visualization. The details page shows calculated sample percentages for the range of\ sequence within the browser window.\
\ \ \\ The pext values are derived from isoform quantifications using the RSEM tool. Detailed information about\ development and commands to create these files can be found here. Pext values were downloaded\ from the gnomAD website \ and transformed into bigWigs, one per tissue. For the full list of UCSC specific steps, please\ see the "gnomAD PEXT scores" section of the\ \ hg38 makedoc from our GitHub repository.\
\ \\ Note that isoform quantification tools can be imprecise, especially for longer genes with many\ annotated isoforms. Regions with low pext values might be enriched for annotation errors (ie. there\ may be edge cases for which an exon that is established to be critical for gene function may appear\ unexpressed with pext). Also note that the GTEx dataset is postmortem adult tissue, and thus\ the possibility that an exon may be development-specific or may be expressed in tissues not\ represented in GTEx can not be dismissed.\
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API.\ The data can also be found directly from the gnomAD downloads page.\
\ \\ Please refer to our\ mailing list archives\ for questions and example queries, or our\ Data Access FAQ\ for more information.
\ \\ More information about using and understanding the gnomAD data can be found in the\ gnomAD FAQ site.\
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the ODC Open Database License\ (ODbL) as described here.\
\ \\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ \ Analysis of protein-coding genetic variation in 60,706 humans.\ Nature. 2016 Aug 18;536(7616):285-91.\ PMID: 27535533; PMC: PMC5018207\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM,\ Ganna A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\ \\ Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, Khera AV, Lowther C,\ Gauthier LD, Wang H et al.\ \ A structural variation reference for medical and population genetics.\ Nature. 2020 May;581(7809):444-451.\ PMID: 32461652; PMC: PMC7334194\
\ \\ Cummings BB, Karczewski KJ, Kosmicki JA, Seaby EG, Watts NA, Singer-Berk M, Mudge JM, Karjalainen J,\ Satterstrom FK, O'Donnell-Luria AH et al.\ \ Transcript expression-aware annotation improves rare variant interpretation.\ Nature. 2020 May;581(7809):452-458.\ PMID: 32461655; PMC: PMC7334198\
\ \\ Cummings BB, Karczewski KJ, Kosmicki JA, Seaby EG, Watts NA, Singer-Berk M, Mudge JM, Karjalainen J,\ Satterstrom FK, O'Donnell-Luria AH et al.\ \ Transcript expression-aware annotation improves rare variant interpretation.\ Nature. 2020 May;581(7809):452-458.\ PMID: 32461655; PMC: PMC7334198\
\ varRep 0 compositeTrack on\ dataVersion Release 4.1\ html gnomadPext\ longLabel Genome Aggregation Database (gnomAD) Proportion Expression Across Transcript Scores (pext)\ maxHeightPixels 100:16:8\ parent gnomadVariants\ shortLabel gnomAD pext\ track gnomadPext\ type bigWig 0 1\ viewLimits 0:1\ visibility hide\ gnomadCopyNumberVariants gnomAD Rare CNV Variants bigBed 9 + Genome Aggregation Database (gnomAD) - Rare CNV variants (<1% overall site frequency) v4.1 0 100 0 0 0 127 127 127 0 0 0 https://gnomad.broadinstitute.org/variant/$$?dataset=gnomad_cnv_r4\ The Genome Aggregation Database (gnomAD) - Rare CNV variants (<1% overall site frequency) v4.1 track set shows rare autosomal coding copy number variants (CNVs) with an overall\ site frequency of less than 1%. These variants were identified from exome sequencing (ES) data of\ 464,297 individuals. The data can also be explored via the\ gnomAD browser.\ \
\ Items are colored by the type of variant:\
| Variant Type | \|
|---|---|
| Deletion (DEL) | \31939 | \
| Duplication (DUP) | \36760 | \
Mouseover on an item will display the position, size of variant, genes impacted by\ variant (>=10% CDS overlap by deletion or >=75% CDS overlap by duplication), and site\ frequency of non-neuro control samples. Item description pages include a linkout to\ the gnomAD browser showing additional genetic ancestry group information.
\ \ \\ To identify rare coding CNVs from the ES data of 464,297 individuals in gnomAD v4, the GATK-gCNV\ method was employed, as described in Babadi et al., Nat Genet, 2023.
\ \
\
\
The CNV discovery process started with collecting the number of reads mapped to 363,301 autosomal\ target intervals derived from protein-coding exons (Fig. 1a, b; Babadi et al.). These read counts\ were used to capture sample-level technical variability, such as differences in exome capture kits\ or sequencing centers, and generated 1,045 different batches of samples for parallel processing\ (Fig. 1c). For each of these batches, 200 random samples were selected for training GATK-gCNV in\ cohort mode,which can be thought of as the creation of a "panel of normals" (PoN). The resulting\ PoN models were then used to efficiently delineate CNV events on all of the samples of their\ respective cohorts using the GATK-gCNV case mode (Fig. 1d,e).\
\ \\ The raw, individual-level CNV calls produced by GATK-gCNV for all samples were then collated,\ and variants observed in multiple individuals were clustered using single-linkage clustering.\ Quality filtering followed the procedures outlined in Babadi et al., filtering CNVs based on\ sample-level (number of events per individual) and call-level (frequency, size, quality score) metrics\ Due to the significant increase in cohort size and heterogeneity compared to the datasets reported\ in Babadi et al., additional filters were applied. Samples with more than five chromosomes harboring\ rare CNVs, as well as those containing more than three rare terminal CNVs, were excluded. 1,049\ sites producing noisy normalized read-depth signals were masked. The final retained CNVs and sites\ were subsequently annotated for impacted genes and frequencies.
\ \\ More information can be found at the\ \ gnomAD site.
\ \\ The bed files was obtained from the gnomAD Google Storage bucket:
\ \\ https://storage.googleapis.com/gcp-public-data--gnomad/release/4.1/exome_cnv/gnomad.v4.1.cnv.all.bed\\ \ The data was then transformed into a bigBed track. For the full list of commands used to make this\ track please see the "gnomAD CNVs v4.1" section of the\ makedoc.\ \ \
\
The raw data can be explored interactively with the Table Browser, or\
the Data Integrator. For automated access, this track, like all \
others, is available via our API. However, for bulk \
processing, it is recommended to download the dataset. The genome annotation is stored in a bigBed \
file that can be downloaded from the\
download server.\
The exact filenames can be found in the track configuration file. Annotations can be converted to\
ASCII text by our tool bigBedToBed which can be compiled from the source code or\
downloaded as a precompiled binary for your system. Instructions for downloading source code and\
binaries can be found\
here. The tool can\
also be used to obtain only features within a given range, for example:
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gnomAD/v4/cnv/gnomad.v4.1.cnv.all.bb -chrom=chr6 -start=0 -end=1000000 stdout\\ \
\ Please refer to our\ mailing list archives\ for questions and example queries, or our\ Data Access FAQ\ for more information.
\ \\ More information about using and understanding the gnomAD data can be found in the\ gnomAD FAQ site.\
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the ODC Open Database License\ (ODbL) as described here.\
\ \ \\ Babadi M, Fu JM, Lee SK, Smirnov AN, Gauthier LD, Walker M, Benjamin DI, Zhao X, Karczewski KJ, Wong\ I et al.\ \ GATK-gCNV enables the discovery of rare copy number variants from exome sequencing data.\ Nat Genet. 2023 Sep;55(9):1589-1597.\ PMID: 37604963; PMC: PMC10904014\
\ \\ Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, Khera AV, Lowther C,\ Gauthier LD, Wang H et al.\ \ A structural variation reference for medical and population genetics.\ Nature. 2020 May;581(7809):444-451.\ PMID: 32461652; PMC: PMC7334194\
\ \\ Cummings BB, Karczewski KJ, Kosmicki JA, Seaby EG, Watts NA, Singer-Berk M, Mudge JM, Karjalainen J,\ Satterstrom FK, O'Donnell-Luria AH et al.\ \ Transcript expression-aware annotation improves rare variant interpretation.\ Nature. 2020 May;581(7809):452-458.\ PMID: 32461655; PMC: PMC7334198\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM,\ Ganna A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\ \\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ \ Analysis of protein-coding genetic variation in 60,706 humans.\ Nature. 2016 Aug 18;536(7616):285-91.\ PMID: 27535533; PMC: PMC5018207\
\ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/v4/cnv/gnomad.v4.1.cnv.all.bb\ dataVersion Release 4.1 (November 01, 2023)\ filterLabel.svtype Type of Variation\ filterValues.svtype DEL|Deletion,DUP|Duplication\ html gnomadCNV\ itemRgb on\ longLabel Genome Aggregation Database (gnomAD) - Rare CNV variants (<1% overall site frequency) v4.1\ mergeSpannedItems on\ mouseOver Position: $chrom:${chromStart}-${chromEnd}\ The gnomAD STR track displays short tandem repeat (STR) genotypes at 87\ disease-associated loci from the\ Genome Aggregation\ Database (gnomAD) v3.1.3. The data include individual-level STR genotypes from\ 18,511 whole-genome sequenced samples across 10 populations, aggregated\ into per-locus allele frequency distributions.
\ \\ These loci were selected because tandem repeat expansions at these sites have been\ reported to cause human genetic diseases, including Huntington disease (HTT),\ fragile X syndrome (FMR1), Friedreich ataxia (FXN), various\ spinocerebellar ataxias, myotonic dystrophies, and other neurological and\ neuromuscular disorders. Most loci (56) have motifs between 3–6 bp, while\ additional loci have longer motifs of 10–24 bp.
\ \\ The genotypes were generated using\ ExpansionHunter\ v5 on gnomAD v3.1 whole-genome sequencing data (150 bp read lengths). Of the\ samples, 64% were PCR-free, 13% PCR-plus, and 23% had unknown PCR protocol.\ ExpansionHunter was selected because it had the best accuracy among existing tools\ for detecting expansions at disease-associated loci. Results were generated without\ off-target regions to minimize overestimation of repeat sizes.\ For each locus, the data show the distribution of repeat allele sizes observed\ across the gnomAD population, providing a reference for normal and expanded allele\ ranges. For more details on the methods, see the\ gnomAD blog post on STR calls.
\ \\ Items are colored by the length of the repeat motif:
\\ Each item is labeled by the gene name. Hovering shows the repeat motif,\ gene, total sample count, and number passing quality filters. Clicking an item\ links to the corresponding gnomAD STR locus page with interactive allele\ frequency histograms and detailed population breakdowns.
\ \\ The detail page for each locus shows:
\\
The gnomAD STR genotype data file\
(gnomAD_STR_genotypes__2025_03_17.tsv.gz) was downloaded from the\
gnomAD downloads page. This file contains individual-level\
STR genotypes at 87 disease-associated loci generated using\
ExpansionHunter\
on gnomAD v3.1.3 whole-genome sequencing data.
\ For the UCSC Genome Browser track, the individual genotype records (~1.4 million rows)\ were aggregated per locus to produce summary statistics: total sample count,\ PASS-filter count, allele size frequency distributions, and per-population sample counts.\ Coordinates were used as provided (0-based). Some loci include genotypes for multiple\ motif patterns (e.g., complex repeat structures) and for adjacent repeats; these are\ represented as separate records.
\ \\ The 10 populations represented are: African/African American (afr),\ Admixed American/Latino (amr), Amish (ami), Ashkenazi Jewish (asj),\ East Asian (eas), Finnish (fin), Middle Eastern (mid), Non-Finnish European (nfe),\ South Asian (sas), and Other (oth).
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated\ analysis, the data may be queried from our\ REST API. The underlying bigBed\ file can be downloaded from our\ download\ server.
\ \\ The complete gnomAD STR dataset, including individual-level genotypes, is available\ from the gnomAD downloads page. Interactive locus-level views with\ allele frequency histograms are available at the\ gnomAD STR browser.
\ \\ Thanks to the gnomAD\ production team at the Broad Institute for generating and distributing this data.
\ \\ Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q,\ Alföldi J, Watts NA, Vittal C, Gauthier LD et al.\ \ A genomic mutational constraint map using variation in 76,156 human\ genomes.\ Nature. 2024 Jan;625(7993):92-100.\ PMID: 38057664; PMC: PMC11629659\
\ \\ Dolzhenko E, Deshpande V, Schlesinger F, Krusche P, Petrovski R,\ Chen S, Emig-Agius D, Gross A, Narzisi G, Bowman B\ et al.\ \ ExpansionHunter: a sequence-graph-based tool to analyze variation\ in short tandem repeat regions.\ Bioinformatics. 2019 Nov 1;35(22):4754-4756.\ PMID: 31134279; PMC: PMC6853681\
\ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/gnomadStr.bb\ dataVersion gnomAD v3.1.3 STR genotypes (March 2025)\ html gnomadStr\ itemRgb on\ longLabel Genome Aggregation Database (gnomAD) - Short Tandem Repeat Genotypes at Disease-Associated Loci\ mouseOver Gene: $geneNOTE: Only variants that have passed\
\ the quality filter are displayed by default.
\
\ The Genome Aggregation Database (gnomAD) - Structural Variants v4.1 track set shows structural variants calls (>=50 nucleotides) from the gnomAD v4.1\ release on 63,046 unrelated genomes. It mostly (but not entirely) overlaps with the genome set used\ for the gnomAD short variant release. For more information see the following blog post, \ \ Structural variants in gnomAD.
\ \\ Items are shaded according to variant type, mouseover on items indicates affected\ protein-coding genes, size of the variant (which may differ from the chromosomal coordinates in\ cases like insertions), variant type (insertion, duplication, etc), allele count, allele number,\ and allele frequency. When more than 2 genes are affected by a variant, the full list can be\ obtained by clicking on the item and reading the details page. A short summary is available in the\ below table:
\ \| Variant Type | \All SV's | \
|---|---|
| Breakend (BND) | \356035 | \
| Complex (CPX) | \15189 | \
| Translocation (CTX) | \99 | \
| Deletion (DEL) | \1206278 | \
| Duplication (DUP) | \269326 | \
| Insertion (INS) | \304645 | \
| Inversion (INV) | \2193 | \
| Copy number variants (CNV) | \721 | \
\ Detailed information on the CNV color code is described \ here. All tracks can be \ filtered according to the size of the variant and variant type, using the track Configure\ options.\
\ \\ Three filters are available for this track:\
\\ The bed files was obtained from the gnomAD Google Storage bucket:\ \
\ https://storage.googleapis.com/gcp-public-data--gnomad/release/4.1/genome_sv/gnomad.v4.1.sv.non_neuro_controls.sites.bed.gz\\ \ The data was then transformed into a bigBed track. For the full list of commands used to make this\ track please see the "gnomAD Structural Variants v4" section of the\ makedoc.\ \ \
\
The raw data can be explored interactively with the Table Browser, or\
the Data Integrator. For automated access, this track, like all \
others, is available via our API. However, for bulk \
processing, it is recommended to download the dataset. The genome annotation is stored in a bigBed \
file that can be downloaded from the\
download server.\
The exact filenames can be found in the track configuration file. Annotations can be converted to\
ASCII text by our tool bigBedToBed which can be compiled from the source code or\
downloaded as a precompiled binary for your system. Instructions for downloading source code and\
binaries can be found\
here. The tool can\
also be used to obtain only features within a given range, for example:
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gnomAD/v4/structuralVariants/gnomad.v4.1.sv.non_neuro_controls.sites.bb -chrom=chr6 -start=0 -end=1000000 stdout\\ \
\ Please refer to our\ mailing list archives\ for questions and example queries, or our\ Data Access FAQ\ for more information.
\ \\ More information about using and understanding the gnomAD data can be found in the\ gnomAD FAQ site.\
\ \\ Thanks to the Genome Aggregation\ Database Consortium for making these data available. The data are released under the ODC Open Database License\ (ODbL) as described here.\
\ \ \\ Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T, O'Donnell-Luria AH, Ware JS, Hill\ AJ, Cummings BB et al.\ \ Analysis of protein-coding genetic variation in 60,706 humans.\ Nature. 2016 Aug 18;536(7616):285-91.\ PMID: 27535533; PMC: PMC5018207\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, Collins RL, Laricchia KM,\ Ganna A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\ \\ Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, Khera AV, Lowther C,\ Gauthier LD, Wang H et al.\ \ A structural variation reference for medical and population genetics.\ Nature. 2020 May;581(7809):444-451.\ PMID: 32461652; PMC: PMC7334194\
\\ Cummings BB, Karczewski KJ, Kosmicki JA, Seaby EG, Watts NA, Singer-Berk M, Mudge JM, Karjalainen J,\ Satterstrom FK, O'Donnell-Luria AH et al.\ \ Transcript expression-aware annotation improves rare variant interpretation.\ Nature. 2020 May;581(7809):452-458.\ PMID: 32461655; PMC: PMC7334198\
\ \ varRep 1 bigDataUrl /gbdb/hg38/gnomAD/v4/structuralVariants/gnomad.v4.1.sv.non_neuro_controls.sites.bb\ dataVersion Release 4.1 (November 01, 2023)\ filter.af_controls 0:1\ filter.af_non_neuro 0:1\ filter.svlen 50:199840172\ filterByRange.af_controls on\ filterByRange.af_non_neuro on\ filterByRange.svlen on\ filterLabel.af_controls Filter by common disease control allele frequency\ filterLabel.af_non_neuro Filter by non-neurological allele frequency\ filterLabel.svlen Filter by Variant Size\ filterLabel.svtype Type of Variation\ filterLimits.af_controls 0:1\ filterLimits.af_non_neuro 0:1\ filterType.FILTER multipleListAnd\ filterValues.FILTER PASS,HIGH_NCR,IGH_MHC_OVERLAP,UNRESOLVED,REFERENCE_ARTIFACT\ filterValues.svtype BND|Breakend,CPX|Complex,CTX|Translocation,DEL|Deletion,DUP|Duplication,INS|Insertion,INV|Inversion,MCNV|Multi-allele CNV\ filterValuesDefault.FILTER PASS\ html gnomadSv.html\ itemRgb on\ longLabel Genome Aggregation Database (gnomAD) - Structural Variants v4.1\ mergeSpannedItems on\ mouseOverField _mouseOver\ parent gnomadVariants on\ shortLabel gnomAD Structural Variants\ track gnomadStructuralVariants\ type bigBed 9 +\ url https://gnomad.broadinstitute.org/variant/$$?dataset=gnomad_sv_r4\ urlLabel gnomAD Structural Variant Browser\ visibility hide\ ctgPos2 GRC Contigs ctgPos Genome Reference Consortium Contigs 3 100 0 0 0 127 127 127 0 0 24 chr1,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chrX,chrY, https://www.ncbi.nlm.nih.gov/nuccore/$$\ This track shows the names of the assembled supercontigs for the GRCh38 (hg38) assembly \ determined by the Genome Reference Consortium (GRC).\
\\ Data for this track were obtained from \ localId2acc files downloaded from GenBank.\
\ map 0 chromosomes chr1,chr3,chr4,chr5,chr6,chr7,chr8,chr9,chr10,chr11,chr12,chr13,chr14,chr15,chr16,chr17,chr18,chr19,chr2,chr20,chr21,chr22,chrX,chrY\ longLabel Genome Reference Consortium Contigs\ shortLabel GRC Contigs\ superTrack assemblyContainer pack\ track ctgPos2\ type ctgPos\ url https://www.ncbi.nlm.nih.gov/nuccore/$$\ grcIncidentDb GRC Incident bigBed 4 + GRC Incident Database 0 100 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/projects/genome/assembly/grc/issue_detail.cgi?id=$$\ This track shows locations in the human assembly where assembly\ problems have been noted or resolved, as reported by the\ Genome Reference Consortium (GRC). \
\\ If you would like to report an assembly problem, please use the GRC\ issue reporting system.\
\ \\ Data for this track are extracted from the GRC\ incident database from the specific species *_issues.gff3 file.\ The track is synchronized once daily to incorporate new updates. \
\ \The data and presentation of this track were prepared by\ Hiram Clawson.\
\ map 1 group map\ longLabel GRC Incident Database\ shortLabel GRC Incident\ track grcIncidentDb\ type bigBed 4 +\ url https://www.ncbi.nlm.nih.gov/projects/genome/assembly/grc/issue_detail.cgi?id=$$\ urlLabel GRC Incident:\ visibility hide\ patchesPsl GRC Patches psl GRC Patches: Alt Haplotypes and Fix Sequences 3 100 0 0 0 127 127 127 0 0 0\ These tracks show the two types of patch sequences from the Genome Reference Consortium\ (GRC) patch releases:
\ \\ This track shows alignments of fix patch sequences to\ main chromosome sequences in the reference genome assembly.\ When errors are corrected in the reference genome assembly, the\ Genome Reference Consortium\ (GRC) adds fix patch sequences containing the corrected regions.\ This strikes a balance between providing the most complete and correct genome\ sequence, while maintaining stable chromosome coordinates for the original assembly\ sequences.\
\\ Fix patches are often associated with incident reports displayed in the GRC Incidents\ track.\
\ \\ This track shows alignments of alternate locus (also known as "alternate haplotype")\ reference sequences to main chromosome sequences in the reference genome assembly.\ Some loci in the genome are highly variable, with sets of variants that tend\ to segregate into distinct haplotypes.\ Only one haplotype can be included in a reference assembly chromosome sequence.\ Instead of providing a separate complete chromosome sequence for each haplotype,\ which could cause confusion with divergent chromosome coordinates and\ ambiguity about which sequence is the official reference, the\ Genome Reference Consortium\ (GRC) adds alternate locus sequences, ranging from tens of thousands of bases\ up to low millions of bases in size, to represent the distinct haplotypes. \
\ \\ Both tracks follow the display conventions for\ \ PSL alignment tracks.\ Mismatching bases are highlighted in red.\ Several types of alignment gap may also be colored;\ for more information, see\ \ Alignment Insertion/Deletion Display Options.\
\\ By default, the tracks are only visible when there are items in the view window.\ This can be disabled by the checkbox Hide empty subtracks.
\ \\ The alignments were provided by NCBI as GFF files and translated into the PSL\ representation for browser display by UCSC.\
\ map 1 compositeTrack on\ group map\ hideEmptySubtracks on\ html patchesPsl\ indelDoubleInsert on\ indelQueryInsert on\ longLabel GRC Patches: Alt Haplotypes and Fix Sequences\ pennantIcon p14 black https://genome-blog.gi.ucsc.edu/blog/patches/ "Includes annotations on GRCh38.p14 patch sequences"\ shortLabel GRC Patches\ track patchesPsl\ type psl\ visibility pack\ gtexEqtlHighConf GTEx cis-eQTLs bigBed GTEx fine-mapped cis-eQTLs 0 100 0 0 0 127 127 127 0 0 0\ This track shows genetic variants likely affecting proximal gene expression in 49 human tissues\ from the\ Genotype-Tissue Expression (GTEx)\ V8 data release.\ \ The data items displayed are gene expression quantitative trait loci within 1MB\ of gene transcription start sites (cis-eQTLs), significantly associated with\ gene expression and in the credible set of variants for the gene at a high\ confidence level. The data can only be calculated for the autosomes,\ so no data is shown on chrX.\
\ \\ Both the CAVIAR and DAP-G tracks show gene/variant pairs for 49 GTEx tissues.\ Variants are linked to the genes they interact with by a line. Variants\ are represented by thicker-width, single-base items. Genes are represented as\ thinner-width items covering the length of the gene. The direction of the\ chevrons on the line indicate whether the variant is upstream or downstream of\ the gene with the chevrons always pointing from the variant to the gene. If a\ variant is internal to the gene, then the variant is shown as a thicker segment\ than the gene. Items in the track are colored according to their tissue, with\ the color matching those in the GTEx Gene V8 Track.\ \
\ Hovering over items in the track display will show the variant ID (often a\ dbSNP rsID), the target gene, tissue, and posterior probablity (Causal\ Posterior Probability (CPP) for CAVIAR; SNP Posterior Inclusion Probability\ (PIP) for DAP-G). Clicking an item will show the details of that interaction\ with link outs to view more details on the GTEx website.\
\ \\ Track configuration supports filtering by tissue, gene, or posterior probability.\
\ \\ Details on GTEx v8 analysis, including code, can be found in the\ GTEx GWAS Analysis Github.\
\ \\ Raw data for these analyses are available from the\ GTEx Portal.\
\ \\ The CAVIAR\ track at UCSC was created using the CAVIAR high-confidence set, which\ represents the high causal variants that have a causal posterior probability\ (CPP) of > 0.1.\
\ \\ The DAP-G track at\ UCSC was created using the DAP-G 95% credible set, which represents varaints\ with strong eQTLs signals, which are signal clusters with signal-level\ posterior inclusion probability (SPIP) > 0.95.\
\ \\ The raw data for this track can be accessed in multiple ways. It can be explored interactively \ using the Table Browser or \ Data Integrator. You can also access the data\ entries in JSON format through our \ JSON API.
\ \\ The data in this track are organized in bigBed file format. The underlying files\ can be obtained from our downloads server:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gtex/eQtl/gtexCaviar.bb\ -chrom=chr16 -start=34990190 -end=36727467 stdout
\ \\ GTEx Consortium.\ \ The GTEx Consortium atlas of genetic regulatory effects across human tissues.\ Science. 2020 Sep 11;369(6509):1318-1330.\ PMID: 32913098; PMC: PMC7737656\
\\ Lee Y, Luca F, Pique-Regi R, Wen X.\ \ Bayesian Multi-SNP Genetic Association Analysis: Control of FDR and Use of\ Summary Statistics.\ bioRxiv. 2018 May 8.\
\\ Wen X, Lee Y, Luca F, Pique-Regi R.\ \ Efficient Integrative Multi-SNP Association Analysis via Deterministic Approximation of\ Posteriors.\ Am J Hum Genet. 2016 Jun 2;98(6):1114-1129.\ PMID: 27236919; PMC: PMC4908152\
\\ Ongen H, Buil A, Brown AA, Dermitzakis ET, Delaneau O.\ \ Fast and efficient QTL mapper for thousands of molecular phenotypes.\ Bioinformatics. 2016 May 15;32(10):1479-85.\ PMID: 26708335; PMC: PMC4866519\
\\ Hormozdiari F, Kostem E, Kang EY, Pasaniuc B, Eskin E.\ \ Identifying causal variants at loci with multiple signals of association.\ Genetics. 2014 Oct;198(2):497-508.\ PMID: 25104515; PMC: PMC4196608\
\\ GTEx Consortium.\ \ The Genotype-Tissue Expression (GTEx) project.\ Nat Genet. 2013 Jun;45(6):580-5.\ PMID: 23715323; PMC: PMC4010069\
\\ \ GTEx Portal Documentation\
\ regulation 1 compositeTrack off\ group regulation\ itemRgb on\ longLabel GTEx fine-mapped cis-eQTLs\ shortLabel GTEx cis-eQTLs\ track gtexEqtlHighConf\ type bigBed\ visibility hide\ gtexGene GTEx Gene bed 6 + Gene Expression in 53 tissues from GTEx RNA-seq of 8555 samples (570 donors) 0 100 0 0 0 127 127 127 1 0 0\ The\ NIH Genotype-Tissue Expression (GTEx) project\ was created to establish a sample and data resource for studies on the relationship between \ genetic variation and gene expression in multiple human tissues. \ This track shows median gene expression levels in 51 tissues and 2 cell lines, \ based on RNA-seq data from the GTEx midpoint milestone data release (V6, October 2015).\ This release is based on data from 8555 tissue samples obtained from 570 adult post-mortem individuals.
\ \\
In Full and Pack display modes, expression for each gene is represented by a colored bargraph,\
where the height of each bar represents the median expression level across all samples for a \
tissue, and the bar color indicates the tissue.\
Tissue colors were assigned to conform to the GTEx Consortium publication conventions.\

\
The bargraph display has the same width and tissue order for all genes.\
Mouse hover over a bar will show the tissue and median expression level.\
The Squish display mode draws a rectangle for each gene, colored to indicate the tissue\
with highest expression level if it contributes more than 10% to the overall expression\
(and colored black if no tissue predominates).\
In Dense mode, the darkness of the grayscale rectangle displayed for the gene reflects the total\
median expression level across all tissues.
\ The GTEx transcript model used to quantify expression level is displayed below the graph,\ colored to indicate the transcript class \ (coding, \ noncoding, \ pseudogene, \ problem), \ following GENCODE conventions.\
\\ Click-through on a graph displays a boxplot of expression level quartiles with outliers, \ per tissue, along with a link to the corresponding gene page on the GTEx Portal.
\ The track configuration page provides controls to limit the genes and tissues displayed,\ and to select raw or log transformed expression level display.\ \\ RNA-seq was performed by the GTEx Laboratory, Data Analysis and Coordinating Center \ (LDACC) at the Broad Institute.\ The Illumina TruSeq protocol was used to create an unstranded polyA+ library sequenced\ on the Illumina HiSeq 2000 platform to produce 76-bp paired end reads at a depth \ averaging 50M aligned reads per sample.\ Sequence reads were aligned to the hg19/GRCh37 human genome using Tophat v1.4.1 \ assisted by the GENCODE v19 transcriptome definition. \ Gene annotations were produced by taking the union of the GENCODE exons for each gene.\ Gene expression levels in RPKM were called via the RNA-SeQC tool, after filtering for \ unique mapping, proper pairing, and exon overlap.\ For further method details, see the \ \ GTEx Portal Documentation page.\
\ UCSC obtained the gene-level expression files, gene annotations and sample metadata from the \ GTEx Portal Download page.\ Median expression level in RPKM was computed per gene/per tissue.
\ \\ The scientific goal of the GTEx project required that the donors and their biospecimen \ present with no evidence of disease. \ The tissue types collected were chosen based on their clinical significance, logistical \ feasibility and their relevance to the scientific goal of the project and the \ research community. \ Postmortem samples were collected from non-diseased donors with ages ranging from 20 to 79. 34.4% of donors were female and 65.6% male. \


\ Additional summary plots of GTEx sample characteristics are available at the \ \ GTEx Portal Tissue Summary page.
\ \ \\ The raw data for the GTEx Gene expression track can be accessed interactively through the \ \ Table Browser or Data Integrator. Metadata can be \ found in the connected tables below.\
\
For automated analysis and downloads, the track data files can be downloaded from \
our downloads server\
or the JSON API.\
Individual regions or the whole genome annotation can be accessed as text using our utility\
bigBedToBed. Instructions for downloading the utility can be found \
here. \
That utility can also be used to obtain features within a given range, e.g. \
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg19/gtex/gtexTranscExpr.bb -chrom=chr21\
-start=0 -end=100000000 stdout
\ Data can also be obtained directly from GTEx at the following link:\ \ https://gtexportal.org/home/datasets
\ \\ Statistical analysis and data interpretation was performed by The GTEx Consortium Analysis \ Working Group. \ Data was provided by the GTEx LDACC at The Broad Institute of MIT and Harvard.
\ \\ GTEx Consortium.\ \ The Genotype-Tissue Expression (GTEx) project.\ Nat Genet. 2013 Jun;45(6):580-5.\ PMID: 23715323; \ PMC: PMC4010069\
\ \\ Carithers LJ, Ardlie K, Barcus M, Branton PA, Britton A, Buia SA, Compton CC, DeLuca DS, Peter-Demchok J, Gelfand ET et al.\ \ A Novel Approach to High-Quality Postmortem Tissue Procurement: The GTEx Project.\ Biopreserv Biobank. 2015 Oct;13(5):311-9.\ PMID: 26484571; \ PMC: PMC4675181
\ \ Melé M, Ferreira PG, Reverter F, DeLuca DS, Monlong J, Sammeth M, Young TR, Goldmann JM,\ Pervouchine DD, Sullivan TJ et al.\ \ Human genomics. The human transcriptome across tissues and individuals.\ Science. 2015 May 8;348(6235):660-5.\ PMID: 25954002; PMC: PMC4547472\ \\ DeLuca DS, Levin JZ, Sivachenko A, Fennell T, Nazaire MD, Williams C, Reich M, Winckler W, Getz G.\ \ RNA-SeQC: RNA-seq metrics for quality control and process optimization.\ Bioinformatics. 2012 Jun 1;28(11):1530-2.\ PMID: 22539670; PMC: PMC3356847
\ \ expression 1 group expression\ html gtexGeneExpr\ longLabel Gene Expression in 53 tissues from GTEx RNA-seq of 8555 samples (570 donors)\ maxItems 200\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel GTEx Gene\ spectrum on\ track gtexGene\ type bed 6 +\ visibility hide\ gtexTranscExpr GTEx Transcript bigBarChart Transcript Expression in 53 tissues from GTEx RNA-seq of 8555 samples/570 donors 0 100 0 0 0 127 127 127 0 0 0\ The\ NIH Genotype-Tissue Expression (GTEx)\ project was created to establish a sample and data resource for studies on the relationship\ between genetic variation and gene expression in multiple human tissues. \ This track displays median transcript expression levels in 53 tissues, based on\ RNA-seq data from the GTEx midpoint milestone data release (V6, October 2015).\ To view the GTEx tissues in anatomical context, see the \ GTEx Body Map.\
\\ Data for this track were computed at UCSC from GTEx RNA-seq sequence data using the\ Toil\ pipeline running the kallisto transcript-level quantification tool.
\ \\ In Full and Pack display modes, expression for each transcript is represented by a colored \ bar chart, where the height of each bar represents the median expression level across all \ samples for a tissue, and the bar color indicates the tissue.\

\ The bar chart display has the same width and tissue order for all transcripts.\ Mouse hover over a bar will show the tissue and median expression level.\ The Squish display mode draws a rectangle for each gene, colored to indicate the tissue\ with highest expression level if it contributes more than 10% to the overall expression\ (and colored black if no tissue predominates).\ In Dense mode, the darkness of the grayscale rectangle displayed for the transcript reflects \ the total median expression level across all tissues.
\\ Click-through on a graph displays a boxplot of expression level quartiles with outliers, \ per tissue.
\ \\ Tissue samples were obtained using the GTEx standard operating procedures for informed consent\ and tissue collection, in conjunction with the \ \ National Cancer Institute Biorepositories and Biospecimen.\ All tissue specimens were reviewed by pathologists to characterize and\ verify organ source.\ Images from stained tissue samples can be viewed via the \ \ NCI histopathology viewer.\ The Qiagen PAXgene non-formalin tissue preservation product was used to stabilize \ tissue specimens without cross-linking biomolecules.
\\ RNA-seq was performed by the GTEx Laboratory, Data Analysis and Coordinating Center \ (LDACC) at the Broad Institute.\ The Illumina TruSeq protocol was used to create an unstranded polyA+ library sequenced\ on the Illumina HiSeq 2000 platform to produce 76-bp paired end reads at a depth \ averaging 50M aligned reads per sample.
\\ Sequence reads for this track were quantified to the hg38/GRCh38 human genome using kallisto\ assisted by the GENCODE v23 transcriptome definition. Read quantification was performed at UCSC\ by the Computational Genomics lab, using the Toil pipeline. The resulting kallisto files were\ combined to generate a transcript per million (TPM) expression matrix using the UCSC tool,\ kallistoToMatrix. Average TPM expression values for each tissue were calculated and \ used to generate a bed6+5 file that is the base of the track. This was done using the UCSC\ tool, expMatrixToBarchartBed. The bed track was then converted to a bigBed file using the \ UCSC tool, bedToBigBed.
\\ The data in the hg19/GRCh37 version of this track was generated by converting the\ coordinates from the hg38/GRCh38 track data.\ Of the 189,615 BED entries from the original hg38 track, 176,220 were mapped over by transcript\ name to hg19 using wgEncodeGencodeCompV24lift37 (~93% coverage).
\ \\ The scientific goal of the GTEx project required that the donors and their biospecimen \ present with no evidence of disease. The tissue types collected were chosen based on their \ clinical significance, logistical feasibility and their relevance to the scientific goal \ of the project and the research community. Postmortem samples were collected from \ non-diseased donors with ages ranging from 20 to 79. 34.4% of donors were female and\ 65.6% male. \

\

\ Additional summary plots of GTEx sample characteristics are available at the \ \ GTEx Portal Tissue Summary page.
\ \\ Samples were collected by the GTEx Consortium.\ RNA-seq was performed by the GTEx Laboratory, Data Analysis and Coordinating Center \ (LDACC) at the Broad Institute.\ John Vivian, Melissa Cline, and Benedict Paten of the UCSC Computational Genomics lab were\ responsible for the sequence read quantification used to produce this track. Kate Rosenbloom \ and Chris Eisenhart of the UCSC Genome Browser group were responsible for data file\ post-processing and track configuration.
\ \\ J. Vivian et al., \ \ Rapid and efficient analysis of 20,000 RNA-seq samples with Toil\ bioRxiv bioRxiv, vol. 2, p. 62497, 2016.
\\ GTEx Consortium.\ \ The Genotype-Tissue Expression (GTEx) project.\ Nat Genet. 2013 Jun;45(6):580-5.\ PMID: 23715323; \ PMC: PMC4010069
\ \\ Carithers LJ, Ardlie K, Barcus M, Branton PA, Britton A, Buia SA, Compton CC, DeLuca DS, Peter-Demchok J, Gelfand ET et al.\ \ A Novel Approach to High-Quality Postmortem Tissue Procurement: The GTEx Project.\ Biopreserv Biobank. 2015 Oct;13(5):311-9.\ PMID: 26484571; \ PMC: PMC4675181
\ \\ Melé M, Ferreira PG, Reverter F, DeLuca DS, Monlong J, Sammeth M, Young TR, Goldmann JM,\ Pervouchine DD, Sullivan TJ et al.\ \ Human genomics. The human transcriptome across tissues and individuals.\ Science. 2015 May 8;348(6235):660-5.\ PMID: 25954002; PMC: PMC4547472
\ \\ DeLuca DS, Levin JZ, Sivachenko A, Fennell T, Nazaire MD, Williams C, Reich M, Winckler W, Getz G.\ \ RNA-SeQC: RNA-seq metrics for quality control and process optimization.\ Bioinformatics. 2012 Jun 1;28(11):1530-2.\ PMID: 22539670; PMC: PMC3356847
\ \ expression 1 barChartBars Adipose-Subcutaneous Adipose-Visceral_(Omentum) Adrenal_Gland Artery-Aorta Artery-Coronary Artery-Tibial Bladder Brain-Amygdala Brain-Anterior_cingulate_cortex_(BA24) Brain-Caudate_(basal_ganglia) Brain-Cerebellar_Hemisphere Brain-Cerebellum Brain-Cortex Brain-Frontal_Cortex_(BA9) Brain-Hippocampus Brain-Hypothalamus Brain-Nucleus_accumbens_(basal_ganglia) Brain-Putamen_(basal_ganglia) Brain-Spinal_cord_(cervical_c-1) Brain-Substantia_nigra Breast-Mammary_Tissue Cells-EBV-transformed_lymphocytes Cells-Transformed_fibroblasts Cervix-Ectocervix Cervix-Endocervix Colon-Sigmoid Colon-Transverse Esophagus-Gastroesophageal_Junction Esophagus-Mucosa Esophagus-Muscularis Fallopian_Tube Heart-Atrial_Appendage Heart-Left_Ventricle Kidney-Cortex Liver Lung Minor_Salivary_Gland Muscle-Skeletal Nerve-Tibial Ovary Pancreas Pituitary Prostate Skin-Not_Sun_Exposed_(Suprapubic) Skin-Sun_Exposed_(Lower_leg) Small_Intestine-Terminal_Ileum Spleen Stomach Testis Thyroid Uterus Vagina Whole_Blood\ barChartColors \\#FFA54F #EE9A00 #8FBC8F #8B1C62 #EE6A50 #FF0000 #CDB79E #EEEE00 \\#EEEE00 #EEEE00 #EEEE00 #EEEE00 #EEEE00 #EEEE00 #EEEE00 #EEEE00 \\#EEEE00 #EEEE00 #EEEE00 #EEEE00 #00CDCD #EE82EE #9AC0CD #EED5D2 \\#EED5D2 #CDB79E #EEC591 #8B7355 #8B7355 #CDAA7D #EED5D2 #B452CD \\#7A378B #CDB79E #CDB79E #9ACD32 #CDB79E #7A67EE #FFD700 #FFB6C1 \\#CD9B1D #B4EEB4 #D9D9D9 #3A5FCD #1E90FF #CDB79E #CDB79E #FFD39B \\#A6A6A6 #008B45 #EED5D2 #EED5D2 #FF00FF\ barChartLabel Tissue types\ barChartMatrixUrl /gbdb/hgFixed/human/expMatrix/cleanGtexMatrix.tab\ barChartMetric median\ barChartSampleUrl /gbdb/hgFixed/human/expMatrix/cleanGtexSamples.tab\ barChartUnit TPM\ bigDataUrl /gbdb/hg38/gtex/gtexTranscExpr.bb\ defaultLabelFields name2, name\ group expression\ labelFields name2, name\ longLabel Transcript Expression in 53 tissues from GTEx RNA-seq of 8555 samples/570 donors\ maxItems 300\ maxLimit 8000\ shortLabel GTEx Transcript\ track gtexTranscExpr\ type bigBarChart\ gwasCatalog GWAS Catalog bed 4 + NHGRI-EBI Catalog of Published Genome-Wide Association Studies 0 100 0 90 0 127 172 127 0 0 0 https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ This track displays single nucleotide polymorphisms (SNPs) identified by published \ Genome-Wide Association Studies (GWAS), collected in the \ NHGRI-EBI GWAS Catalog\ published jointly by the National\ Human Genome Research Institute (NHGRI) and the European Bioinformatics Institute (EMBL-EBI).\ Some abbreviations\ are used above.\
\\ From http://www.ebi.ac.uk/gwas/docs/about:\
\ The Catalog is a quality controlled, manually curated, literature-derived\ collection of all published genome-wide association studies assaying at least\ 100,000 SNPs and all SNP-trait associations with p-values < 1.0 x\ 10-5 (Hindorff et al., 2009). For more details about the Catalog\ curation process and data extraction procedures, please refer to the\ Methods page.\\ \ \
\ From http://www.ebi.ac.uk/gwas/docs/methods:\
\ The GWAS Catalog data is extracted from the literature. Extracted information\ includes publication information, study cohort information such as cohort size,\ country of recruitment and subject ethnicity, and SNP-disease association\ information including SNP identifier (i.e. RSID), p-value, gene and risk\ allele. Each study is also assigned a trait that best represents the phenotype\ under investigation. When multiple traits are analysed in the same study either\ multiple entries are created, or individual SNPs are annotated with their\ specific traits. Traits are used both to query and visualise the data in the\ Catalog's web form and diagram-based query interfaces.\\ \ \
\ Data extraction and curation for the GWAS Catalog is an expert activity; each\ step is performed by scientists supported by a web-based tracking and data\ entry system which allows multiple curators to search, annotate, verify and\ publish the Catalog data. Papers that qualify for inclusion in the Catalog are\ identified through weekly PubMed searches. They then undergo two levels of\ curation. First all data, including association information for SNPs, traits\ and general information about the study, are extracted by one curator. A second\ curator then performs an additional round of curation to double-check the\ accuracy and consistency of all the information. Finally, an automated pipeline\ performs validation of the extracted data, see the\ Quality control and SNP mapping section below for more\ details. This information is then used for queries and in the production of the\ diagram.\
\ Previous versions of this track can be found on our archive download server.\
\ \\ Hindorff LA, Sethupathy P, Junkins HA, Ramos EM, Mehta JP, Collins FS, Manolio TA.\ \ Potential etiologic and functional implications of genome-wide association loci for human diseases\ and traits.\ Proc Natl Acad Sci U S A. 2009 Jun 9;106(23):9362-7.\ PMID: 19474294; PMC: PMC2687147\
\ phenDis 1 color 0,90,0\ group phenDis\ longLabel NHGRI-EBI Catalog of Published Genome-Wide Association Studies\ shortLabel GWAS Catalog\ snpTable snp144\ snpVersion 144\ track gwasCatalog\ type bed 4 +\ url https://www.ncbi.nlm.nih.gov/SNP/snp_ref.cgi?type=rs&rs=$$\ urlLabel dbSNP:\ visibility hide\ gwipsvizRiboseq GWIPS-viz Riboseq bigWig 0 3589344 Ribosome Profiling from GWIPS-viz 0 100 0 0 0 127 127 127 0 0 0\ Ribosome profiling (ribo-seq) is a technique that takes advantage of NGS\ technology to sequence ribosome-protected mRNA fragments and consequently\ allows the locations of translating ribosomes to be determined at the entire\ transcriptome level (Ingolia et al., 2009).\
\ \\ For a more detailed description of the protocol, see Ingolia et al.\ (2012). For reviews on this technique and its applications, please refer to\ Ingolia (2014) and Michel et al. (2013).\
\ \\ This track displays cumulative ribo-seq data obtained from human cells under\ different conditions and can be used for the exploration of human genomic loci\ that are being translated. The values on the y-axis represent the number of\ ribosome footprint sequence reads at a given position. As of February\ 2016, the track contains data from 9 studies (see References section for\ details). Further details about the aggregated track and additional ribo-seq\ data from these and other studies including data obtained from other organisms\ can be found at the specialized ribo-seq browser\ GWIPS-viz.\
\ \\ For each study used to generate this track, raw fastq files were downloaded from\ a repository (e.g., NCBI GEO datasets).\ Cutadapt\ was used to trim the relevant adapter sequence from the reads, after which reads\ below 25 nt in length were discarded. The trimmed reads were aligned to\ ribosomal RNA using\ Bowtie\ and aligning reads were discarded. The remaining reads were then aligned to the\ hg38 (GRCh38) genome assembly using Bowtie. An offset of 15 nt (to infer the\ position of the A-site) was added to the most 5' nucleotide coordinate of each\ uniquely-mapped read.\
\ \\ The alignment files from each of the included studies were merged to generate\ this aggregate track.\
\ \\ See individual studies at\ GWIPS-viz for a full\ description of the methods of data acquisition and processing.\
\ \\ Thanks to Audrey Michel, Stephen Kiniry and GWIPS-viz for providing the data for\ this track. If you wish to cite this track, please reference:\
\ \\ Michel AM, Fox G, M Kiran A, De Bo C, O'Connor PB, Heaphy SM, Mullan JP, Donohue CA, Higgins DG,\ Baranov PV.\ GWIPS-viz: development of a ribo-seq genome browser.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D859-64.\ PMID: 24185699; PMC: PMC3965066\
\ \\ Battle A, Khan Z, Wang SH, Mitrano A, Ford MJ, Pritchard JK, Gilad Y.\ \ Impact of regulatory variation from RNA to protein.\ Science. 2015 Feb 6;347(6222):664-7.\ PMID: 25657249;\ PMC: PMC4507520\
\ \\ Cenik C, Cenik ES, Byeon GW, Grubert F, Candille SI, Spacek D, Alsallakh B, Tilgner H, Araya CL, Tang H et al.\ \ Integrative analysis of RNA, translation and protein levels reveals distinct regulatory variation across humans.\ Genome Res. 2015 Nov;25(11):1610-21.\ PMID: 26297486;\ PMC: PMC4617958\
\ \ \\ Elkon R, Loayza-Puch F, Korkmaz G, Lopes R, van Breugel PC, Bleijerveld OB, Altelaar AM, Wolf E, Lorenzin F, Eilers M et al.\ \ Myc coordinates transcription and translation to enhance transformation and suppress invasiveness.\ EMBO Rep. 2015 Dec;16(12):1723-36.\ PMID: 26538417;\ PMC: PMC4687422\
\ \\ Jang C, Lahens NF, Hogenesch JB, Sehgal A.\ \ Ribosome profiling reveals an important role for translational control in circadian gene expression.\ Genome Res 2015 Dec;25(12):1836-47.\ PMID: 26338483;\ PMC: PMC4665005\
\ \\ Ji Z, Song R, Regev A, Struhl K.\ \ Many lncRNAs, 5'UTRs, and pseudogenes are translated and some are likely to express functional proteins.\ Elife. 2015 Dec 19;4.\ PMID: 26687005;\ PMC: PMC4739776\
\ \\ Sidrauski C, McGeachy AM, Ingolia NT, Walter P.\ \ The small molecule ISRIB reverses the effects of eIF2α phosphorylation on translation and stress granule assembly.\ Elife. 2015 Feb 26;4.\ PMID: 25719440;\ PMC: PMC4341466\
\ \\ Tanenbaum ME, Stern-Ginossar N, Weissman JS, Vale RD.\ \ Regulation of mRNA translation during mitosis.\ Elife. 2015 Aug 25;4.\ PMID: 26305499;\ PMC: PMC4548207\
\ \\ Tirosh O, Cohen Y, Shitrit A, Shani O, Le-Trilling VT, Trilling M, Friedlander G, Tanenbaum M, Stern-Ginossar N.\ \ The transcription and translation landscapes during human cytomegalovirus infection reveal novel host-pathogen interactions.\ PLoS Pathog. 2015 Nov 24;11(11):e1005288.\ PMID: 26599541;\ PMC: PMC4658056\
\ \\ Werner A, Iwasaki S, McGourty CA, Medina-Ruiz S, Teerikorpi N, Fedrigo I, Ingolia NT, Rape M.\ \ Cell fate determination by ubiquitin-dependent regulation of translation.\ Nature. 2015 Sep 24;525(7570):523-7.\ PMID: 26399832;\ PMC: PMC4602398\
\ \\ Ingolia NT.\ \ Ribosome profiling: new views of translation, from single codons to genome scale.\ Nat Rev Genet. 2014 Mar;15(3):205-13.\ PMID: 24468696\
\ \\ Ingolia NT, Brar GA, Rouskin S, McGeachy AM, Weissman JS.\ \ The ribosome profiling strategy for monitoring translation in vivo by deep sequencing of ribosome-\ protected mRNA fragments.\ Nat Protoc. 2012 Jul 26;7(8):1534-50.\ PMID: 22836135; PMC: PMC3535016\
\ \\ Ingolia NT, Ghaemmaghami S, Newman JR, Weissman JS.\ \ Genome-wide analysis in vivo of translation with nucleotide resolution using ribosome profiling.\ Science. 2009 Apr 10;324(5924):218-23.\ PMID: 19213877; PMC: PMC2746483\
\ \\ Michel AM, Baranov PV.\ \ Ribosome profiling: a Hi-Def monitor for protein synthesis at the genome-wide scale.\ Wiley Interdiscip Rev RNA. 2013 Sep-Oct;4(5):473-90.\ PMID: 23696005; PMC: PMC3823065\
\ expression 0 autoScale off\ group expression\ html gwipsvizRiboseq\ longLabel Ribosome Profiling from GWIPS-viz\ maxHeightPixels 100:32:8\ shortLabel GWIPS-viz Riboseq\ track gwipsvizRiboseq\ type bigWig 0 3589344\ viewLimits 0:2000\ visibility hide\ han945Sv Han 945 SVs bigBed 9 + Structural Variants from 945 Han Chinese (Long-read Sequencing) 0 100 0 0 0 127 127 127 0 0 0\ This track shows structural variants (SVs) identified by long-read sequencing\ of 945 Han Chinese individuals. The dataset contains 111,288 SVs merged across\ samples using SURVIVOR, including 49,518 deletions, 42,300 insertions,\ 13,503 duplications, 5,595 inversions, and 372 translocations.\
\ \\ Items are colored by SV type:\
\ Filters are available for SV type, SV length, allele frequency, and number of\ supporting samples. For insertions, the item is placed at the insertion site\ with a width of 1 bp. For translocations, only the first breakpoint is shown;\ the second breakpoint chromosome and position are listed in the item details.\
\ \\ Gong et al. 2025 performed Oxford Nanopore long-read sequencing of 945\ Han Chinese individuals on PromethION instruments with R9.4 flow cells.\ Reads were aligned to GRCh38.p13 with NGMLR v0.2.7 using ONT-tuned\ parameters, and a joint-calling strategy was used to call SVs at moderate\ coverage: per-sample discovery with\ cuteSV\ v1.0.13, merging of breakpoints within 500 bp across individuals with\ SURVIVOR\ v1.0.6, per-sample re-genotyping of the merged set with LRcaller v1.0, and\ a final BCFtools merge. SVs in centromeric, pericentromeric and gap regions\ were filtered out, yielding 111,288 high-quality SVs: 49,518 deletions,\ 42,300 insertions, 13,503 duplications, 5,595 inversions and 372\ translocations.\
\\ The site-only VCF released at\ \ OMIX accession OED00945268 (OED00945268_Han_945samples_SV.vcf.gz)\ was converted to BED for this track.\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ The raw VCF data was obtained from the\ OMIX\ repository (accession OED00945268) at the National Genomics Data Center (NGDC),\ China National Center for Bioinformation.\
\\ The source VCF also encodes phased per-sample genotypes: the sampleList\ field on the detail page is derived from the SURVIVOR SUPP_VEC bitmask\ and is an ordered list of the 1-based indices of the 945 samples carrying\ each SV. The full per-sample phased VCF can be browsed as a separate track in\ the SVs from 945 Han Chinese entry of\ the Phased Variants track collection.\
\ \\ Thanks to Gong et al. for making their structural variant calls publicly available.\
\ \\ Gong J, Sun H, Wang K, Zhao Y, Huang Y, Chen Q, Qiao H, Gao Y, Zhao J, Ling Y et al.\ \ Long-read sequencing of 945 Han individuals identifies structural variants associated with\ phenotypic diversity and disease susceptibility.\ Nat Commun. 2025 Feb 10;16(1):1494.\ PMID: 39929826; PMC: PMC11811171\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/han945.bb\ filter.AC 0:1890\ filter.alleleFreq 0:1\ filter.insLen 0:27242\ filter.sampleCount 1:945\ filter.svLen 0:99743\ filterByRange.AC on\ filterByRange.alleleFreq on\ filterByRange.insLen on\ filterByRange.sampleCount on\ filterByRange.svLen on\ filterLabel.AC Allele Count (approx 2*SUPP)\ filterLabel.alleleFreq Allele Frequency\ filterLabel.insLen Insertion Length\ filterLabel.sampleCount Number of Supporting Samples\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.alleleFreq 0:1\ filterType.svType multipleListOr\ filterValues.svType DEL,INS,DUP,INV,TRA\ itemRgb on\ longLabel Structural Variants from 945 Han Chinese (Long-read Sequencing)\ mouseOver Var: $name ($svType)\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 0 group singleCell\ longLabel Heart single cell RNA data from https://heartcellatlas.com\ shortLabel Heart Cell Atlas\ superTrack on\ track heartCellAtlas\ visibility hide\ heartAtlasAgeGroup Heart HCA Age bigBarChart Heart cell RNA binned by age group of donor from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars 40-45 45-50 50-55 55-60 60-65 65-70 70-75\ barChartColors #c22694 #c22794 #c22498 #c32c8d #bd5269 #b6615d #c63c79\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/age_group.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/age_group.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by age group of donor from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Age\ track heartAtlasAgeGroup\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ heartAtlasCellTypes Heart HCA Cells bigBarChart Heart cell RNA binned by cell type from https://heartcellatlas.org 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars adipocyte atrial_cardiomyocyte endothelial fibroblast lymphoid mesothelial myeloid neuronal not_assigned pericyte smooth_muscle_cell ventricular_cardiomyocyte doublet\ barChartColors #f1803d #c1229a #07bc02 #b5562a #eb1613 #1494b3 #de2b02 #e6af0e #c12792 #c15f4e #b06a5a #c1229b #d69f85\ barChartLimit 3\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/cell_type.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by cell type from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Cells\ track heartAtlasCellTypes\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ heartAtlasDonor Heart HCA Donor bigBarChart Heart cell RNA binned by organ donor from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars D1 D11 D2 D3 D4 D5 D6 D7 H2 H3 H4 H5 H6 H7\ barChartColors #c43483 #469615 #c53483 #c54868 #c63c79 #c3377e #9e7358 #b65e62 #c53186 #c12b90 #c22596 #c12498 #c22694 #c22794\ barChartLimit 4\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/donor.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by organ donor from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Donor\ track heartAtlasDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ heartAtlasRegion Heart HCA Region bigBarChart Heart cell RNA binned by region of collection from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars AX LA LV RA RV SP\ barChartColors #c13782 #c14d68 #c12596 #c14472 #c12696 #c02f8d\ barChartLimit 1.5\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/region.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/region.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by region of collection from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Region\ track heartAtlasRegion\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ heartAtlasSample Heart HCA Sample bigBarChart Heart cell RNA binned by biosample from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars H0015_LA_new H0015_LV H0015_RA H0015_RV H0015_apex H0015_septum H0020_LA_new H0020_LV H0020_RA H0020_RV H0020_apex H0020_septum H0025_LA H0025_LV H0025_RA H0025_RV H0025_apex H0025_septum H0026_LA H0026_LV_V3 H0026_RA H0026_RV H0026_apex H0026_septum2 H0035_LA H0035_LV H0035_RA H0035_RV H0035_apex H0035_septum H0037_Apex H0037_LA_corr H0037_LV H0037_RA_corr H0037_RV H0037_septum HCAHeart7606896 HCAHeart7656534 HCAHeart7656535 HCAHeart7656536 HCAHeart7656537 HCAHeart7656538 HCAHeart7656539 HCAHeart7664652 HCAHeart7664653 HCAHeart7664654 HCAHeart7698015 HCAHeart7698016 HCAHeart7698017 HCAHeart7702873 HCAHeart7702874 HCAHeart7702875 HCAHeart7702876 HCAHeart7702877 HCAHeart7702878 HCAHeart7702879 HCAHeart7702880 HCAHeart7702881 HCAHeart7702882 HCAHeart7728604 HCAHeart7728605 HCAHeart7728606 HCAHeart7728607 HCAHeart7728608 HCAHeart7728609 HCAHeart7745966 HCAHeart7745967 HCAHeart7745968 HCAHeart7745969 HCAHeart7745970 HCAHeart7751845 HCAHeart7757636 HCAHeart7757637 HCAHeart7757638 HCAHeart7757639 HCAHeart7829976 HCAHeart7829977 HCAHeart7829978 HCAHeart7829979 HCAHeart7833852 HCAHeart7833853 HCAHeart7833854 HCAHeart7833855 HCAHeart7835148 HCAHeart7835149 HCAHeart7836681 HCAHeart7836682 HCAHeart7836683 HCAHeart7836684 HCAHeart7843999 HCAHeart7844000 HCAHeart7844001 HCAHeart7844002 HCAHeart7844003 HCAHeart7844004 HCAHeart7850539 HCAHeart7850540 HCAHeart7850541 HCAHeart7850542 HCAHeart7850543 HCAHeart7850544 HCAHeart7850545 HCAHeart7850546 HCAHeart7850547 HCAHeart7850548 HCAHeart7850549 HCAHeart7850551 HCAHeart7880860 HCAHeart7880861 HCAHeart7880862 HCAHeart7880863 HCAHeart7888922 HCAHeart7888923 HCAHeart7888924 HCAHeart7888925 HCAHeart7888926 HCAHeart7888927 HCAHeart7888928 HCAHeart7888929 HCAHeart7905327 HCAHeart7905328 HCAHeart7905329 HCAHeart7905330 HCAHeart7905331 HCAHeart7905332 HCAHeart7964513 HCAHeart7985086 HCAHeart7985087 HCAHeart7985088 HCAHeart7985089 HCAHeart8102858 HCAHeart8102859 HCAHeart8102860 HCAHeart8102861 HCAHeart8102862 HCAHeart8102863 HCAHeart8102864 HCAHeart8102865 HCAHeart8102866 HCAHeart8102867 HCAHeart8102868 HCAHeart8287123 HCAHeart8287124 HCAHeart8287125 HCAHeart8287126 HCAHeart8287127 HCAHeart8287128\ barChartColors #c63682 #c12399 #c65169 #c12794 #c12498 #c12497 #cf5d48 #c22793 #c33f7b #c22695 #c12597 #c32b8f #c83e78 #c12992 #c43a7f #c12d8f #c12d8f #c03786 #cf5b4e #c32a90 #cb5b56 #c32d8b #c63581 #c42d8c #c22992 #c12597 #cd555c #c12993 #c32e8a #c13982 #c12498 #c9466b #c22694 #c8476c #c12497 #c12993 #85b660 #59850c #489210 #6d7611 #a4a063 #846816 #c22894 #c22793 #c42e8b #d15956 #c63e75 #c74172 #c63d77 #c73d77 #c53188 #c53188 #ca4f5a #ca4e5b #c63b78 #c63a7a #c32a91 #c63780 #c63a7c #e1cec2 #e0d2c5 #bf8d6b #c38f77 #7bbd5e #ebded6 #3e980c #83b75f #826716 #ae4719 #79be5d #45940f #e3948e #e8acc0 #e18e93 #ce4f60 #c8456c #cb476c #c53582 #c73f75 #c73f74 #c7466a #c73d77 #c7436f #c63a7b #c53187 #c73e76 #c6446d #c8466a #c73f74 #3e960a #ae4a1e #905c12 #d02f19 #af4c22 #896113 #389c0d #49900e #57860e #82650f #3c990d #40960d #a14f10 #bb4309 #cb3809 #c73a0a #ac480f #a84c0d #c43582 #c44467 #c4397d #c04b5b #c53188 #c43484 #c32a90 #c53681 #c22d8c #c22d8c #c22795 #c33385 #23ab08 #22ab09 #23aa08 #3a9b0b #16b306 #19b106 #c63c77 #cb594f #c42f8a #cf535d #c53583 #4a900c #4b910e #4f8e10 #3f970a #43950b #3e9a0f #29a70a #2ba60b #399d0d #1caf07 #3b9b0d #c43089 #c42d8d #d877ae #c12f8c #c13686 #c33685\ barChartLimit 4\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/sample.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/sample.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by biosample from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Sample\ track heartAtlasSample\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ heartAtlasSex Heart HCA Sex bigBarChart Heart cell RNA binned by sex of donor from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars Female Male\ barChartColors #c12794 #c13682\ barChartLimit 1\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/sex.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/sex.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by sex of donor from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Sex\ track heartAtlasSex\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ heartAtlasSource Heart HCA Source bigBarChart Heart cell RNA binned by source (nucleus vs whole cell) from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars CD45+ Cells Nuclei\ barChartColors #2da207 #1ab006 #c22695\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/source.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/source.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by source (nucleus vs whole cell) from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Source\ track heartAtlasSource\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ heartAtlasCellStates Heart HCA State bigBarChart Heart cell RNA binned by cell state from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars Adip1 Adip2 Adip3 Adip4 B_cells CD14+Mo CD16+Mo CD4+T_cytox CD4+T_tem CD8+T_cytox CD8+T_tem DC DOCK4+MØ1 DOCK4+MØ2 EC10_CMC-like EC1_cap EC2_cap EC3_cap EC4_immune EC5_art EC6_ven EC7_atria EC8_ln EC9_FB-like FB1 FB2 FB3 FB4 FB5 FB6 FB7 IL17RA+Mo LYVE1+MØ1 LYVE1+MØ2 LYVE1+MØ3 Mast Meso Mo_pi MØ_AgP MØ_mod NC1 NC2 NC3 NC4 NC5 NC6 NK NKT NØ PC1_vent PC2_atria PC3_str PC4_CMC-like SMC1_basic SMC2_art aCM1 aCM2 aCM3 aCM4 aCM5 doublets nan vCM1 vCM2 vCM3 vCM4 vCM5\ barChartColors #ef7f3e #ea7b3d #e87c40 #eb9d88 #c13e20 #cc3a0c #d63105 #e21e17 #d02d17 #e71a14 #a87052 #d62915 #d06946 #cf7046 #439918 #0db804 #0eb804 #10b604 #14b405 #11b604 #24a907 #b5734a #9c734c #c16036 #b7582f #b95a2e #b95e31 #b85c33 #b46239 #bb5b30 #c23876 #f3d3c8 #d93006 #b0754e #cb3b11 #c96848 #1494b3 #d73005 #d43408 #d23508 #e2aa14 #c8722b #b9ab74 #d66eb9 #e4b670 #edd9c6 #dc2315 #e31d15 #e0b09b #c45a4d #c35d48 #8d7848 #c22694 #b86756 #9d6b56 #c1229a #c12499 #c02c91 #c22a90 #d66dbb #d69f85 #c12792 #c1229a #c1229a #c22695 #c22696 #c1219b\ barChartLimit 4\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/cell_states.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/cell_states.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by cell state from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA State\ track heartAtlasCellStates\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ heartAtlasVersion Heart HCA Version bigBarChart Heart cell RNA binned by 10x chemistry version from https://heartcellatlas.org 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ This track displays data from \ Cells of the adult human heart. Single-cell and single-nucleus RNA\ sequencing (RNA-seq) was used to profile transcriptomes from six regions of the heart:\ the interventricular septum (SP), apex (AX), left ventricle (LV), right\ ventricle (RV), left atrium (LA), and right atrium (RA). A total of 11 cardiac\ cell types were identified along with their marker genes after uniform manifold\ approximation and projection (UMAP) embedding of 487,106 cells. Note that the RNA-seq\ data is generated using Tag-sequencing (Tag-seq) and does not cover all exons.
\ \
\ This track collection contains nine bar chart tracks of RNA expression in the\ human heart where cells are grouped by cell type \ (Heart HCA Cells), age \ (Heart HCA Age), donor \ (Heart HCA Donor), region of the heart \ (Heart HCA Region),\ sample (Heart HCA Sample), sex \ (Heart HCA Sex), source \ (Heart HCA Source), cell\ state (Heart HCA State), \ and 10x chemistry version \ (Heart HCA Version). \ The default track displayed is \ Heart HCA Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| lymphoid | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Heart HCA Cells subtrack, where the \ bars represent relatively pure cell types. They can give an overview of the cell composition \ within other categories in other subtracks as well.
\ \\ Healthy heart tissues were obtained from 14 UK and North American transplant\ organ donors ages 40-75. Tissues were taken from deceased donors after\ circulatory death (DCD) and after brain death (DBD). To minimize\ transcriptional degradation, heart tissues were stored and transported on ice\ until freezing or tissue dissociation. Single nuclei were isolated from\ flash-frozen tissue using mechanical homogenization with a glass Dounce tissue\ grinder. Fresh heart tissues were enzymatically dissociated and automatically\ digested using gentleMACS Octo Dissociator. Next, Hoechst-positive single\ nuclei were FACS sorted prior to library preparation. In parallel, Cell\ suspensions from fresh heart tissue were enriched for CD45+ cells using MACS LS\ columns. Libraries of single cell and single nuclei were prepared using 10x\ Genomics 3' v2 or v3. 3' gene expression libraries were sequenced on an\ Illumina HiSeq4000 and NextSeq500. In total 45,870 cells, 78,023 CD45+ enriched\ cells, and 363,213 nuclei were profiled for 11 major cell types of the heart.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Monika Litviňuková, Carlos\ Talavera-Ló, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. \ The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ singleCell 1 barChartBars V2 V3\ barChartColors #c23a7b #c12e8d\ barChartLimit 1\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/heartCellAtlas/version.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/heartCellAtlas/version.bb\ defaultLabelFields name\ html heartCellAtlas\ labelFields name,name2\ longLabel Heart cell RNA binned by 10x chemistry version from https://heartcellatlas.org\ parent heartCellAtlas\ shortLabel Heart HCA Version\ track heartAtlasVersion\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=heart-cell-atlas+global&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ adult_heart_models Heart models bigBed 12 + Adult Heart transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-Heart.bb\ longLabel Adult Heart transcript models\ parent sample_models_view on\ shortLabel Heart models\ subGroups view=sample_models_view sample=adult_heart type=models\ track adult_heart_models\ type bigBed 12 +\ visibility squish\ adult_heart_ont_post_models Heart ONT post models bigBed 12 + Adult Heart ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Heart02Rep1.bb\ itemRgb on\ longLabel Adult Heart ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Heart ONT post models\ subGroups view=per_expr_models_view sample=adult_heart type=post_capture_ont_models\ track adult_heart_ont_post_models\ type bigBed 12 +\ visibility hide\ adult_heart_ont_post_reads Heart ONT post reads bam Adult Heart ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Heart02Rep1.bam\ longLabel Adult Heart ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Heart ONT post reads\ subGroups view=per_expr_reads_view sample=adult_heart type=post_capture_ont_reads\ track adult_heart_ont_post_reads\ type bam\ visibility hide\ adult_heart_ont_pre_models Heart ONT pre models bigBed 12 + Adult Heart ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Heart02Rep1.bb\ itemRgb on\ longLabel Adult Heart ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Heart ONT pre models\ subGroups view=per_expr_models_view sample=adult_heart type=pre_capture_ont_models\ track adult_heart_ont_pre_models\ type bigBed 12 +\ visibility hide\ adult_heart_ont_pre_reads Heart ONT pre reads bam Adult Heart ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Heart02Rep1.bam\ longLabel Adult Heart ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Heart ONT pre reads\ subGroups view=per_expr_reads_view sample=adult_heart type=pre_capture_ont_reads\ track adult_heart_ont_pre_reads\ type bam\ visibility hide\ adult_heart_pacbio_post_models Heart PB post models bigBed 12 + Adult Heart PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Heart02Rep1.bb\ itemRgb on\ longLabel Adult Heart PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Heart PB post models\ subGroups view=per_expr_models_view sample=adult_heart type=post_capture_pacbio_models\ track adult_heart_pacbio_post_models\ type bigBed 12 +\ visibility hide\ adult_heart_pacbio_post_reads Heart PB post reads bam Adult Heart PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Heart02Rep1.bam\ longLabel Adult Heart PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Heart PB post reads\ subGroups view=per_expr_reads_view sample=adult_heart type=post_capture_pacbio_reads\ track adult_heart_pacbio_post_reads\ type bam\ visibility hide\ adult_heart_pacbio_pre_models Heart PB pre models bigBed 12 + Adult Heart PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Heart02Rep1.bb\ itemRgb on\ longLabel Adult Heart PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Heart PB pre models\ subGroups view=per_expr_models_view sample=adult_heart type=pre_capture_pacbio_models\ track adult_heart_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ adult_heart_pacbio_pre_reads Heart PB pre reads bam Adult Heart PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Heart02Rep1.bam\ longLabel Adult Heart PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Heart PB pre reads\ subGroups view=per_expr_reads_view sample=adult_heart type=pre_capture_pacbio_reads\ track adult_heart_pacbio_pre_reads\ type bam\ visibility hide\ gnomADPextHeart_AtrialAppendage Heart-Atrial Appendage bigWig 0 1 gnomAD pext Heart-Atrial Appendage 0 100 153 0 255 204 127 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Heart_AtrialAppendage.bw\ color 153,0,255\ longLabel gnomAD pext Heart-Atrial Appendage\ parent gnomadPext off\ shortLabel Heart-Atrial Appendage\ track gnomADPextHeart_AtrialAppendage\ visibility hide\ gnomADPextHeart_LeftVentricle Heart-Left Ventricle bigWig 0 1 gnomAD pext Heart-Left Ventricle 0 100 102 0 153 178 127 204 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Heart_LeftVentricle.bw\ color 102,0,153\ longLabel gnomAD pext Heart-Left Ventricle\ parent gnomadPext off\ shortLabel Heart-Left Ventricle\ track gnomADPextHeart_LeftVentricle\ visibility hide\ netHprcGCA_018504085v1 HG02080.mat netAlign GCA_018504085.1 chainHprcGCA_018504085v1 HG02080.mat HG02080.pri.mat.f1_v2 (May 2021 GCA_018504085.1_HG02080.pri.mat.f1_v2) HPRC project computed Chain Nets 1 100 0 0 0 255 255 0 0 0 0 hprc 0 longLabel HG02080.mat HG02080.pri.mat.f1_v2 (May 2021 GCA_018504085.1_HG02080.pri.mat.f1_v2) HPRC project computed Chain Nets\ otherDb GCA_018504085.1\ parent hprcChainNetViewnet off\ priority 84\ shortLabel HG02080.mat\ subGroups view=net sample=s084 population=eas subpop=khv hap=mat\ track netHprcGCA_018504085v1\ type netAlign GCA_018504085.1 chainHprcGCA_018504085v1\ hg38ContigDiff Hg19 Diff bed 9 . Contigs New to GRCh38/(hg38), Not Carried Forward from GRCh37/(hg19) 0 100 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/nuccore/$$\ This track shows the differences between the GRCh38 (hg38) and previous GRCh37 (hg19)\ human genome assemblies, indicating contigs (or portions of contigs) that are new\ to the hg38 assembly.\
\ \\
The following color/score key is used:\
\
\
| color | score | change from hg19 to hg38 |
|---|---|---|
| 0 | New contig added to\ hg38 to update sequence or fill gaps present in hg19 | |
| 500 | Different portions\ of this same contig used in the construction of hg38 and hg19 assemblies | |
| 1000 | Updated version of\ an hg19 contig in which sequence errors have been corrected |
\ Use the score filter to select which categories to show in the display.\
\ \\ The contig coordinates were extracted from the AGP files for both assemblies.\ Contigs that matched the same name, same version, and the same specific\ portion of sequence in both assemblies were considered identical between the two\ assemblies and were excluded from this data set. The remaining contigs are shown\ in this track.\
\ \\ The data and presentation of this track were prepared by\ Hiram Clawson, UCSC Genome\ Browser engineering.\
\ map 1 group map\ longLabel Contigs New to GRCh38/(hg38), Not Carried Forward from GRCh37/(hg19)\ scoreFilterByRange on\ shortLabel Hg19 Diff\ track hg38ContigDiff\ type bed 9 .\ url https://www.ncbi.nlm.nih.gov/nuccore/$$\ urlLabel Genbank accession:\ visibility hide\ hgmd HGMD Public 2025 bigBed 9 . Human Gene Mutation Database - Public Version 2025 0 100 0 0 0 127 127 127 0 0 0 http://www.hgmd.cf.ac.uk/ac/gene.php?gene=$P&accession=$pNOTE:
\
HGMD public is intended for use primarily by physicians and other\
professionals concerned with genetic disorders, by genetics researchers, and\
by advanced students in science and medicine. While the HGMD public database is\
open to all academic users, users seeking information about a personal medical\
or genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions.
DOWNLOADS:
\
As requested by Qiagen, this track is not available for download or mirroring but only for limited API queries, see below.\
\ This track shows the genomic positions of variants in the public version of the\ Human Gene Mutation Database (HGMD). \ UCSC does not host any further information and provides only the coordinates of\ mutations.\
\ \\ To get details on a mutation (bibliographic reference, phenotype,\ disease, nucleotide change, etc.), follow the "Link to HGMD" at the top\ of the details page. Mouse over to show the type of variant (substitution, insertion,\ deletion, regulatory or splice variant). For deletions, only start coordinates are shown\ as the end coordinates have not been provided by HGMD. Insertions are located between the two\ annotated nucleic acids.\
\ \\ The HGMD public database is produced at Cardiff University, but is free only\ for academic use. Academic users can register for a free account at the\ HGMD\ User Registration page. Download and commercial use requires a license for the HGMD Professional\ database, which also contains many mutations not yet added to the public version of HGMD public.\ The public version is usually 1-2 years behind the professional version.\
\ \The HGMD database itself does not come with a mapping to genome coordinates,\ but there is a related product called "GenomeTrax" which includes HGMD in the\ UCSC Custom Track format. Contact Qiagen for more information.
\ \Due to license restrictions, the HGMD data is not available for download or for batch queries in the Table Browser. \ However, it is available for programmatic access via the Global\ Alliance Beacon API, a web service that accepts queries in the form\ (genome, chromosome, position, allele) and returns "true" or "false" depending on whether there\ is information about this allele in the database. For more details see our \ Beacon Server.
\Subscribers of the HGMD database can also download the full database or use the HGMD API to retrieve full details, please contact Qiagen support\ for further information. Academic or non-profit users may be able to obtain a\ limited version of HGMD public from Qiagen.
\ \\ Genomic locations of HGMD variants are labeled with the gene symbol\ and the accession of the mutation, separated by a colon. All other information\ is shown on the respective HGMD variation page, accessible via the\ "Link to HGMD" at the top of the details page.\
\ \HGMD variants are originally annotated on RefSeq transcripts. You can show\ all and only those transcripts annotated by HGMD by activating the HGMD\ subtrack of the track "NCBI RefSeq".
\ \\ The mappings displayed on this track were obtained from Qiagen\ and reformatted at UCSC as a bigBed file.\
\ \\ Thanks to HGMD, Frank Schacherer and Rupert Yip from Qiagen for making these data available.\
\ \\ Stenson PD, Mort M, Ball EV, Shaw K, Phillips A, Cooper DN.\ \ The Human Gene Mutation Database: building a comprehensive mutation repository for clinical and\ molecular genetics, diagnostic testing and personalized genomic medicine.\ Hum Genet. 2014 Jan;133(1):1-9.\ PMID: 24077912; PMC: PMC3898141\
\ phenDis 1 bigDataUrl /gbdb/hg38/bbi/hgmd.bb\ group phenDis\ itemRgb on\ longLabel Human Gene Mutation Database - Public Version 2025\ maxItems 1000\ maxWindowCoverage 10000000\ mouseOverField variantType\ noScoreFilter on\ shortLabel HGMD Public 2025\ tableBrowser off hgmd\ track hgmd\ type bigBed 9 .\ url http://www.hgmd.cf.ac.uk/ac/gene.php?gene=$P&accession=$p\ urlLabel Link to HGMD\ visibility hide\ hgnc HGNC bigBed 9 + HUGO Gene Nomenclature 0 100 0 0 0 127 127 127 0 0 0 https://www.genenames.org/data/gene-symbol-report/#!/hgnc_id/$$\ The HGNC is \ responsible for approving unique symbols and names for human loci, including protein \ coding genes, ncRNA genes and pseudogenes, to allow unambiguous scientific communication.\
\ For each known human gene, the HGNC approves a gene name and symbol (short-form abbreviation).\ All approved symbols are stored in the HGNC database, www.genenames.org, a curated online repository of HGNC-approved gene \ nomenclature, gene groups and associated resources including links to genomic, proteomic, \ and phenotypic information. Each symbol is unique and we ensure that each gene is only \ given one approved gene symbol. It is necessary to provide a unique symbol for each gene \ so that we and others can talk about them, and this also facilitates electronic data \ retrieval from publications and databases. In preference, each symbol maintains \ parallel construction in different members of a gene family and can also be \ used in other species, especially other vertebrates including mouse.\ \
\ The raw data can be explored interactively with the Table Browser, or the Data Integrator. For computational analysis, genome annotations are stored in\ a bigBigFile file that can be downloaded from the\ download\ server. Regional or genome-wide annotations can be converted from binary data to human readable\ text using our command line utility bigBedToBed which can be compiled from source code or\ downloaded as a precompiled binary for your system. Files and instructions can be found in the\ utilities directory.\ \ The utility can be used to obtain features within a given range, for example:
\bigBedToBed -chrom=chr6 -start=0 -end=1000000 http://hgdownload.soe.ucsc.edu/gbdb/hg38/hgnc/hgnc.bb stdout\
\
\ \
\ Please refer to our Data Access FAQ\ for more information or our mailing list for archived user questions.
\ \\ HGNC Database, HUGO Gene Nomenclature Committee (HGNC), European Molecular Biology Laboratory, European Bioinformatics Institute (EMBL-EBI), Wellcome Genome Campus, Hinxton, Cambridge CB10 1SD, United Kingdom www.genenames.org.\ \
\ Tweedie S, Braschi B, Gray KA, Jones TEM, Seal RL, Yates B, Bruford EA. Genenames.org: the HGNC and VGNC resources in 2021. Nucleic Acids Res. PMID: 33152070 PMCID: PMC7779007 DOI: 10.1093/nar/gkaa980\
\ genes 1 bigDataUrl /gbdb/hg38/hgnc/hgnc.bb\ defaultLabelFields symbol\ filterValues.locus_type Y RNA,long non-coding RNA,micro RNA,misc RNA,ribosomal RNA,small nuclear RNA,small nucleolar RNA,transfer RNA,vault RNA,T cell receptor gene,T cell receptor pseudogene,complex locus constituent,endogenous retrovirus,gene with protein product,immunoglobulin gene,immunoglobulin pseudogene,pseudogene,readthrough,unknown\ group genes\ itemRgb on\ labelFields symbol, geneName, name, uniprot_ids, ensembl_gene_id, ucsc_id, refseq_accession\ longLabel HUGO Gene Nomenclature\ mouseOver Symbol: $symbol\ This track shows structural variants (SVs) from the second phase of the\ Human Genome Structural Variation Consortium (HGSVC2). The callset is\ derived from 32 haplotype-resolved diploid genomes (64 phased haplotypes)\ spanning five 1000 Genomes superpopulations (African, Admixed American,\ East Asian, European, South Asian). Each genome was sequenced with\ PacBio long reads (continuous long-read and HiFi) and phased with\ Strand-seq.\
\\ The track merges the two SV annotation tables from the HGSVC2 v2.0\ integrated callset freeze 4: 111,330 insertions/deletions and 416\ inversions, for a total of 111,746 SVs. Each row is a site-level variant\ with per-site allele count, carrier haplotypes, population-scale allele\ frequencies (imputed from the phased callset back into 1000 Genomes,\ insertions and deletions only) and structural annotations.\
\ \\ Items are colored by SV type:\
\ Insertions are placed at the insertion site with a width of 1 bp; deletions\ and inversions span the affected reference interval. Filters are available\ for SV type, SV length, carrier-haplotype count, distinct sample count,\ whether the site falls in a Tandem Repeat Finder region and the fraction\ of the variant overlapping segmental duplications.\
\\ The detail page shows, where available:\
\ Ebert et al. 2021 produced phased haplotype-resolved de novo assemblies for\ 32 diploid samples (64 unrelated haplotypes) across five 1000 Genomes\ superpopulations on the PacBio Sequel II platform, using continuous\ long-read sequencing (CLR, >40x) and high-fidelity sequencing (HiFi,\ >20x). Single-cell Strand-seq data from the same samples were used to\ phase the assemblies without parental trios, yielding N50 contigs >25 Mbp\ at QV > 40. SVs were discovered from the two haplotype assemblies of\ each sample with the Phased Assembly Variant (PAV) caller against GRCh38,\ and candidate SVs were orthogonally supported by at least one of seven\ other sources (read-based callers MELT, PBSV and PALMER; Bionano optical\ mapping; breakpoint k-mer analysis; PAV replication with LRA). This\ yielded the integrated nonredundant callset of 107,590 insertion/deletion\ SVs and 316 inversions. Population-scale allele frequencies (POP_*_AF) were\ obtained by graph-based re-genotyping of the HGSVC2 SVs into the\ 3,202-sample 1000 Genomes short-read cohort with PanGenie (insertions and\ deletions only).\
\\ For display, the HGSVC2 v2.0 freeze-4 annotation tables\ variants_freeze4_sv_insdel.tsv.gz (111,330 DEL+INS) and\ variants_freeze4_sv_inv.tsv.gz (416 INV) were downloaded from the\ \ IGSR HGSVC2 v2.0 integrated-callset directory and merged into a single\ bigBed; type-specific columns (POP_*_AF for insdel, RGN_REF_INNER for\ inversions) are empty on the detail page when they do not apply.\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator, and accessed\ programmatically through our API,\ track=hgsvc2Sv.\
\\ The bigBed is available from\ our\ download server as hgsvc2.bb. Example:\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/hgsvc2.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\\ The original annotation tables and VCFs are available from the\ \ HGSVC2 v2.0 integrated callset on the IGSR FTP site.\
\ \\ Thanks to the Human Genome Structural Variation Consortium (HGSVC) and\ the 1000 Genomes Project for releasing this dataset. Later HGSVC releases\ are also available as UCSC tracks:\ HGSVC3 65 SVs.\
\ \\ Ebert P, Audano PA, Zhu Q, Rodriguez-Martin B, Porubsky D, Bonder MJ, Sulovari A, Ebler J, Zhou W,\ Serra Mari R et al.\ \ Haplotype-resolved diverse human genomes and integrated analysis of structural variation.\ Science. 2021 Apr 2;372(6537).\ PMID: 33632895; PMC: PMC8026704\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/hgsvc2.bb\ filter.AC 1:35\ filter.insLen 0:108546\ filter.refSd 0:1\ filter.sampleCount 1:35\ filter.svLen 0:57207414\ filterByRange.AC on\ filterByRange.insLen on\ filterByRange.refSd on\ filterByRange.sampleCount on\ filterByRange.svLen on\ filterLabel.AC Allele Count (carrier haplotypes)\ filterLabel.insLen Insertion Length\ filterLabel.refSd Segmental Duplication Overlap\ filterLabel.refTrf In Tandem Repeat\ filterLabel.sampleCount Sample Count\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.refSd 0:1\ filterType.refTrf multipleListOr\ filterType.svType multipleListOr\ filterValues.refTrf True,False\ filterValues.svType DEL,INS,INV\ itemRgb on\ longLabel Structural Variants from 32 Haplotype-Resolved Genomes (HGSVC2 freeze 4, Ebert 2021)\ mouseOver Var: $name ($svType)\ This track shows structural variants (SVs) from the third phase of the\ Human Genome Structural Variation Consortium (HGSVC3). The callset comes\ from 65 diverse individuals across five continental groups, each sequenced\ with PacBio HiFi (~47x), Oxford Nanopore ultra-long reads (~56x) and\ complemented with Strand-seq, optical mapping, Hi-C and Iso-Seq for\ haplotype-resolved assembly. SVs were discovered from the de novo assemblies\ with PAV v2.4.0.1 and cross-validated by ten additional orthogonal callers.\
\\ The track merges the two final SV annotation tables from the HGSVC3 v1.0\ release on GRCh38: 176,231 insertions/deletions and 300 inversions, for a\ total of 176,531 SVs. Each row is a site-level variant with the list of\ carrier haplotypes and additional structural annotations.\
\\ The same track is also available natively on the T2T-CHM13 (hs1)\ assembly: HGSVC3 independently aligned all haplotype-resolved assemblies\ to both GRCh38 and T2T-CHM13 and released a separate set of annotation\ tables per reference. The hs1 track is built directly from the\ \ HGSVC3 T2T-CHM13 annotation tables (188,224 DEL+INS and 276 INV;\ 188,500 SVs total); no liftOver is involved.\
\ \\ Items are colored by SV type:\
\ Insertions are placed at the insertion site with a width of 1 bp; deletions\ and inversions span the affected reference interval. Filters are available\ for SV type, SV length, carrier-haplotype count, distinct sample count,\ whether the site falls in a Tandem Repeat Finder region and the fraction\ of the variant overlapping segmental duplications.\
\\ The detail page shows, where available:\
\ Logsdon et al. 2025 produced fully phased hybrid de novo assemblies for 65\ diverse individuals (63 from 1kGP, NA21487 from HapMap, and HG002 from\ GIAB), using PacBio HiFi (Sequel II/Revio, 30-h movies), Oxford Nanopore\ ultra-long sequencing (R9.4.1 PromethION, 96-h runs), Bionano optical\ mapping (DLE-1 on Saphyr 2nd-gen), Strand-seq, Hi-C (Proximo) and Iso-Seq.\ Assemblies were generated with Verkko v1.4.1 (primary) and hifiasm-UL\ v0.19.6 (complementary, especially for centromeres and Yq12), phased with\ the Graphasing pipeline v0.3.1-alpha, and produced 130 haplotype\ assemblies with median N50 of 130 Mbp that close 92% of previous assembly\ gaps (39% of chromosomes at telomere-to-telomere status). SVs were called\ against GRCh38 and T2T-CHM13 with PAV v2.4.1 (plus DipCall and SVIM-asm\ from the same alignments) and cross-validated with an additional ten\ callers (PBSV, Sniffles, Delly, cuteSV, DeBreak, SVIM, DeepVariant,\ Clair3, PEPPER-Margin-DeepVariant for ONT and MELT-LRA/PALMER2 for MEIs).\ Calls were merged with SV-Pop and centromere-satellite / telomere hits\ were filtered. The final GRCh38 release contains 176,231 DEL+INS plus 300\ INV (176,531 SVs total); the T2T-CHM13 release contains 188,224 DEL+INS\ plus 276 INV (188,500 SVs total).\
\\ For display, the two final HGSVC3 v1.0 annotation tables\ variants_GRCh38_sv_insdel_HGSVC2024v1.0.tsv.gz and\ variants_GRCh38_sv_inv_HGSVC2024v1.0.tsv.gz were downloaded from\ the \ IGSR HGSVC3 GRCh38 release directory and merged into a single bigBed.\ The hs1 version uses the parallel\ variants_T2T-CHM13_sv_insdel_HGSVC2024v1.0.tsv.gz and\ variants_T2T-CHM13_sv_inv_HGSVC2024v1.0.tsv.gz tables from the\ \ HGSVC3 T2T-CHM13 release directory; no liftOver is involved on hs1.\ Type-specific columns (HOM_REF/HOM_TIG/TE for insdel; RGN_REF_INNER for\ inversions) are empty on the detail page when they do not apply.\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt and\ \ doc/hs1/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator, and accessed\ programmatically through our API,\ track=hgsvc3Sv.\
\\ The bigBed is available from our download server for both assemblies:\
\ The original annotation tables are available from the\ \ HGSVC3 release on the IGSR FTP site.\
\ \\ Thanks to the Human Genome Structural Variation Consortium (HGSVC) and all\ participating sequencing and analysis centers for making the HGSVC3\ annotation tables publicly available.\
\ \\ Logsdon GA, Ebert P, Audano PA, Loftus M, Porubsky D, Ebler J, Yilmaz F, Hallast P, Prodanov T, Yoo\ D et al.\ \ Complex genetic variation in nearly complete human genomes.\ Nature. 2025 Aug;644(8076):430-441.\ PMID: 40702183; PMC: PMC12350169\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/hgsvc3.bb\ filter.AC 1:136\ filter.insLen 0:30176500\ filter.refSd 0:1\ filter.sampleCount 1:65\ filter.svLen 0:30176500\ filterByRange.AC on\ filterByRange.insLen on\ filterByRange.refSd on\ filterByRange.sampleCount on\ filterByRange.svLen on\ filterLabel.AC Allele Count (carrier haplotypes)\ filterLabel.insLen Insertion Length\ filterLabel.refSd Segmental Duplication Overlap\ filterLabel.refTrf In Tandem Repeat\ filterLabel.sampleCount Sample Count\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.refSd 0:1\ filterType.refTrf multipleListOr\ filterType.svType multipleListOr\ filterValues.refTrf True,False\ filterValues.svType DEL,INS,INV\ itemRgb on\ longLabel Structural Variants from 65 Diverse Samples (HGSVC3 ONT+HIFI)\ mouseOver Var: $name ($svType)\
\
\
\
\ Square mode provides a traditional Hi-C display in which chromosome positions are mapped along the\ top-left-to-bottom-right diagonal, and interaction values are plotted on both sides of that diagonal\ to form a square. The upper-left corner of the square corresponds to the left-most position of the\ window in view, while the bottom-right corner corresponds to the right-most position of the window.\
\ The color shade at any point within the square shows the proximity score for two genomic regions:\ the region where a vertical line drawn from that point intersects with the diagonal, and the region\ where a horizontal line from that point intersects with the diagonal. A point directly on the\ diagonal shows the score for how proximal a region is to itself (scores on the diagonal are usually\ quite high unless no data are available). A point at the extreme bottom left of the square shows the\ score for how proximal the left-most position within the window is to the right-most position within\ the window.\
\ In triangle mode, the display is quite similar to square except that only the top half of the square\ is drawn (eliminating the redundancy), and the image is rotated so that the diagonal of the square\ now lies on the horizontal axis. This display consumes less vertical space in the image, although it\ may be more difficult to ascertain exactly which positions correspond to a point within the\ triangle.\
\ In arc mode, simple arcs are drawn between the centers of interacting regions. The color of each arc\ corresponds to the proximity score. Self-interactions are not displayed.\
\
\ There are four score values available in this display: NONE, VC, VC_SQRT, and KR. NONE provides raw,\ un-normalized counts for the number of interactions between regions. VC, or Vanilla Coverage,\ normalization (Lieberman-Aiden et al., 2009) and the VC_SQRT variant normalize these count\ values based on the overall count values for each of the two interacting regions. Knight-Ruiz, or\ KR, matrix balancing (Knight and Ruiz, 2013) provides an alternative normalization method where the\ row and column sums of the contact matrix equal 1.\
\ Color intensity in the heatmap goes up to indicate higher scores, but eventually saturates at a\ maximum beyond which all scores share the same color intensity. The value of this maximum score for\ saturation can be set manually by un-checking the "Auto-scale" box. When the\ "Auto-scale" box is checked, it automatically sets the saturation maximum to be double\ (2x) the median score in the current display window.\
\
\
\ The first protocol, in situ Hi-C, was published in 2014 as a technique for obtaining full-genome\ proximity data while keeping the cell nucleus intact (Rao et al., 2014). This method uses a\ restriction enzyme to cleave DNA before linking. The second protocol, Micro-C XL, is an update to\ the Micro-C method of obtaining chromatin conformation data (Hsieh et al., 2016, Hsieh\ et al., 2015), and has largely supplanted the original. Both the original Micro-C and the\ updated version are variants of Hi-C chromatin conformation capture that use micrococcal nuclease to\ segment the genome before linking. This results in data sets with resolution down to the nucleosome\ level. The original Micro-C method had difficulty recovering higher order interactions, and the\ updated protocol makes use of additional cross-linking chemicals to address that issue.\
\ We downloaded the .hic contact matrix files with the following accessions from the 4D Nucleome\ Data Portal:\ 4DNFI18Q799K,\ 4DNFI2TK7L2F,\ 4DNFIFLJLIS5, and\ 4DNFIQYQWPF5.\ The files are parsed for display using the Straw library from the Aiden lab at Baylor College\ of Medicine.\
\
\ Knight P, Ruiz D.\ \ A fast algorithm for matrix balancing.\ IMA J Numer Anal. 2013 Jul;33(3):1029-1047.\
\ Krietenstein N, Abraham S, Venev SV, Abdennur N, Gibcus J, Hsieh TS, Parsi KM, Yang L, Maehr R,\ Mirny LA et al.\ \ Ultrastructural Details of Mammalian Chromosome Architecture.\ Mol Cell. 2020 May 7;78(3):554-565.e7.\ PMID: 32213324\
\ Lieberman-Aiden E, van Berkum NL, Williams L, Imakaev M, Ragoczy T, Telling A, Amit I, Lajoie BR,\ Sabo PJ, Dorschner MO et al.\ \ Comprehensive mapping of long-range interactions reveals folding principles of the human genome.\ Science. 2009 Oct 9;326(5950):289-93.\ PMID: 19815776; PMC: PMC2858594\
\ Rao SS, Huntley MH, Durand NC, Stamenova EK, Bochkov ID, Robinson JT, Sanborn AL, Machol I, Omer AD,\ Lander ES et al.\ \ A 3D map of the human genome at kilobase resolution reveals principles of chromatin looping.\ Cell. 2014 Dec 18;159(7):1665-80.\ PMID: 25497547; PMC: PMC5635824\
\ regulation 1 compositeTrack on\ group regulation\ longLabel Comparison of Micro-C and In situ Hi-C protocols in H1-hESC and HFFc6\ shortLabel Hi-C and Micro-C\ track hicAndMicroC\ type hic\ cons470way Hiller Lab 470 Mammals bed 4 Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group, 0 100 0 0 0 127 127 127 0 0 0\ This track shows multiple alignments of 470 mammal\ assemblies and measurements of evolutionary conservation\ from the Michael Hiller Lab. There is some duplication of different assemblies for the\ same species, hence there are 431 distinct species in this collection.\
\ \\ The multiple alignments were generated using multiz and\ other tools in the UCSC/Penn State Bioinformatics\ comparative genomics alignment pipeline.\ Conserved elements identified by phastCons are also displayed in\ this track.\
\ \
\ The base-wise conservation scores are computed using two methods\ phastCons and phyloP from the\ PHAST package,\ for all species.\
\ \\ PhastCons (which has been used in previous Conservation tracks) is a hidden\ Markov model-based method that estimates the probability that each\ nucleotide belongs to a conserved element, based on the multiple alignment.\ It considers not just each individual alignment column, but also its\ flanking columns. By contrast, phyloP separately measures conservation at\ individual columns, ignoring the effects of their neighbors. As a\ consequence, the phyloP plots have a less smooth appearance than the\ phastCons plots, with more "texture" at individual sites. The two methods\ have different strengths and weaknesses. PhastCons is sensitive to "runs"\ of conserved sites, and is therefore effective for picking out conserved\ elements. PhyloP, on the other hand, is more appropriate for evaluating\ signatures of selection at particular nucleotides or classes of nucleotides\ (e.g., third codon positions, or first positions of miRNA target sites).\
\ \\ The genome assemblies are from a variety of sources. Some are equivalent\ to UCSC genome browser assemblies, some are from NCBI Genbank assemblies,\ and some are from the DNA Zoo.\ When available in the UCSC browser system, links are provided in the table\ below. Otherwise, links are provided to source locations for the assemblies.\
\\
\\ \\
\ count \common \
nameclade \scientific \
nameassembly \taxon id \\ 1 \human \primates \Homo sapiens \Dec. 2013 (GRCh38/hg38) \9606 \\ 2 \chimpanzee \Primates \Pan troglodytes \Jan. 2018 (Clint_PTRv2/panTro6) \9598 \\ 3 \pygmy chimpanzee \Primates \Pan paniscus \May 2020 (Mhudiblu_PPA_v0/panPan3) \9597 \\ 4 \western lowland gorilla \Primates \Gorilla gorilla gorilla \Aug. 2019 (Kamilah_GGO_v0/gorGor6) \9595 \\ 5 \Sumatran orangutan \Primates \Pongo abelii \Jan. 2018 (Susie_PABv2/ponAbe3) \9601 \\ 6 \northern white-cheeked gibbon \Primates \Nomascus leucogenys \HLnomLeu4 GCA_006542625.1 \61853 \\ 7 \silvery gibbon \Primates \Hylobates moloch \HLhylMol2 GCA_009828535.2 \81572 \\ 8 \pig-tailed macaque \Primates \Macaca nemestrina \Mar. 2015 (Mnem_1.0/macNem1) \9545 \\ 9 \gelada \Primates \Theropithecus gelada \HLtheGel1 GCA_003255815.1 \9565 \\ 10 \crab-eating macaque \Primates \Macaca fascicularis \HLmacFas6 GCA_012559485.1 \9541 \\ 11 \Mona monkey \Primates \Cercopithecus mona \HLcerMon1 GCA_014849445.1 \36226 \\ 12 \Ugandan red Colobus \Primates \Piliocolobus tephrosceles \HLpilTep2 GCA_002776525.3 \591936 \\ 13 \Angolan colobus \Primates \Colobus angolensis palliatus \Mar. 2015 (Cang.pa_1.0/colAng1) \336983 \\ 14 \drill \Primates \Mandrillus leucophaeus \Mar. 2015 (Mleu.le_1.0/manLeu1) \9568 \\ 15 \sooty mangabey \Primates \Cercocebus atys \Mar. 2015 (Caty_1.0/cerAty1) \9531 \\ 16 \olive baboon \Primates \Papio anubis \HLpapAnu5 GCA_008728515.1 \9555 \\ 17 \mandrill \Primates \Mandrillus sphinx \HLmanSph1 GCA_004802615.1 \9561 \\ 18 \Hanuman langur \Primates \Semnopithecus entellus \HLsemEnt1 GCA_004025065.1_SemEnt_v1_BIUU \88029 \\ 19 \Rhesus monkey \Primates \Macaca mulatta \Feb. 2019 (Mmul_10/rheMac10) \9544 \\ 20 \Japanese macaque \Primates \Macaca fuscata \DNA zoo Macaca fuscata \9542 \\ 21 \Francois's langur \Primates \Trachypithecus francoisi \HLtraFra1 GCA_009764315.1 \54180 \\ 22 \black snub-nosed monkey \Primates \Rhinopithecus bieti \Aug. 2016 (ASM169854v1/rhiBie1) \61621 \\ 23 \golden snub-nosed monkey \Primates \Rhinopithecus roxellana \HLrhiRox2 GCA_007565055.1 \61622 \\ 24 \Red shanked douc langur \Primates \Pygathrix nemaeus \HLpygNem1 GCA_004024825.1_PygNem_v1_BIUU \54133 \\ 25 \De Brazza's monkey \Primates \Cercopithecus neglectus \HLcerNeg1 GCA_004027615.1_CertNeg_v1_BIUU \36227 \\ 26 \proboscis monkey \Primates \Nasalis larvatus \Nov. 2014 (Charlie1.0/nasLar1) \43780 \\ 27 \Allen's swamp monkey \Primates \Allenopithecus nigroviridis \DNA zoo Allenopithecus nigroviridis \54135 \\ 28 \green monkey \Primates \Chlorocebus sabaeus \Mar. 2014 (Chlorocebus_sabeus 1.1/chlSab2) \60711 \\ 29 \red guenon \Primates \Erythrocebus patas \HLeryPat1 GCA_004027335.1_EryPat_v1_BIUU \9538 \\ 30 \white-faced saki \Primates \Pithecia pithecia \HLpitPit1 GCA_004026645.1_PitPit_v1_BIUU \43777 \\ 31 \black-handed spider monkey \Primates \Ateles geoffroyi \HLateGeo1 GCA_004024785.1_AteGeo_v1_BIUU \9509 \\ 32 \Ma's night monkey \Primates \Aotus nancymaae \Jun. 2017 (Anan_2.0/aotNan1) \37293 \\ 33 \Bolivian titi \Primates \Plecturocebus donacophilus \HLpleDon1 GCA_004027715.1_CalDon_v1_BIUU \230833 \\ 34 \mantled howler monkey \Primates \Alouatta palliata \HLaloPal1 GCA_004027835.1_AloPal_v1_BIUU \30589 \\ 35 \Bolivian squirrel monkey \Primates \Saimiri boliviensis \DNA zoo Saimiri boliviensis \27679 \\ 36 \tamarin \Primates \Saguinus imperator \HLsagImp1 GCA_004024885.1_SagImp_v1_BIUU \9491 \\ 37 \Bolivian squirrel monkey \Primates \Saimiri boliviensis boliviensis \Oct. 2011 (Broad/saiBol1) \39432 \\ 38 \white-tufted-ear marmoset \Primates \Callithrix jacchus \HLcalJac4 GCA_011100555.1_mCalJac1.pat.X \9483 \\ 39 \pygmy marmoset \Primates \Callithrix pygmaea \DNA zoo Callithrix pygmaea \9493 \\ 40 \tufted capuchin \Primates \Sapajus apella \HLsapApe1 GCA_009761245.1 \9515 \\ 41 \Panamanian white-faced capuchin \Primates \Cebus capucinus imitator \Apr. 2016 (Cebus_imitator-1.0/cebCap1) \2715852 \\ 42 \white-fronted capuchin \Primates \Cebus albifrons \HLcebAlb1 GCA_004027755.1_CebAlb_v1_BIUU \9514 \\ 43 \aye-aye \Primates \Daubentonia madagascariensis \HLdauMad1 GCA_004027145.1_DauMad_v1_BIUU \31869 \\ 44 \Coquerel's sifaka \Primates \Propithecus coquereli \Mar. 2015 (Pcoq_1.0/proCoq1) \379532 \\ 45 \babakoto \Primates \Indri indri \HLindInd1 GCA_004363605.1_IndInd_v1_BIUU \34827 \\ 46 \brown lemur \Primates \Eulemur fulvus \HLeulFul1 GCA_004027275.1_EulFul_v1_BIUU \13515 \\ 47 \Sclater's lemur \Primates \Eulemur flavifrons \Aug. 2015 (Eflavifronsk33QCA/eulFla1) \87288 \\ 48 \Ring-tailed lemur \Primates \Lemur catta \HLlemCat1 GCA_004024665.1_LemCat_v1_BIUU \9447 \\ 49 \greater bamboo lemur \Primates \Prolemur simus \HLproSim1 GCA_003258685.1 \1328070 \\ 50 \mongoose lemur \Primates \Eulemur mongoz \DNA zoo Eulemur mongoz \34828 \\ 51 \Sclater's lemur \Primates \Eulemur flavifrons \DNA zoo Eulemur flavifrons \87288 \\ 52 \Lesser dwarf lemur \Primates \Cheirogaleus medius \HLcheMed1 GCA_008086735.1 \9460 \\ 53 \black lemur \Primates \Eulemur macaco \Aug. 2015 (Emacaco_refEf_BWA_oneround/eulMac1) \30602 \\ 54 \Philippine tarsier \Primates \Carlito syrichta \Sep. 2013 (Tarsius_syrichta-2.0.1/tarSyr2) \1868482 \\ 55 \gray mouse lemur \Primates \Microcebus murinus \Feb. 2017 (Mmur_3.0/micMur3) \30608 \\ 56 \Northern giant mouse lemur \Primates \Mirza zaza \HLmirZaz1 GCA_008750895.1 \339999 \\ 57 \Coquerel's mouse lemur \Primates \Mirza coquereli \HLmirCoq1 GCA_004024645.1_MizCoq_v1_BIUU \47180 \\ 58 \mouse lemur \Primates \Microcebus sp. 3 GT-2019 \HLmicSpe31 GCA_008750915.1 \2508170 \\ 59 \Northern rufous mouse lemur \Primates \Microcebus tavaratra \HLmicTav1 GCA_008750935.1 \143351 \\ 60 \slow loris \Primates \Nycticebus coucang \HLnycCou1 GCA_004027815.1_NycCou_v1_BIUU \9470 \\ 61 \small-eared galago \Primates \Otolemur garnettii \Mar. 2011 (Broad/otoGar3) \30611 \\ 62 \Sunda flying lemur \Euarchontoglires \Galeopterus variegatus \HLgalVar2 GCA_004027255.2 \482537 \\ 63 \Chinese tree shrew \Euarchontoglires \Tupaia chinensis \Jan 2013 (TupChi_1.0/tupChi1) \246437 \\ 64 \northern tree shrew \Euarchontoglires \Tupaia belangeri \Dec. 2006 (Broad/tupBel1) \37347 \\ 65 \puma \Carnivora \Puma concolor \HLpumCon1 GCA_003327715.1_PumCon1.0 \9696 \\ 66 \Amur tiger \Carnivora \Panthera tigris altaica \06 Sep 2013 (PanTig1.0/panTig1) \74533 \\ 67 \Clouded leopard \Carnivora \Neofelis nebulosa \DNA zoo Neofelis nebulosa \61452 \\ 68 \leopard \Carnivora \Panthera pardus \HLpanPar1 GCA_001857705.1_PanPar1.0 \9691 \\ 69 \bearded seal \Carnivora \Erignathus barbatus \DNA zoo Erignathus barbatus \39304 \\ 70 \jaguar \Carnivora \Panthera onca \HLpanOnc1 GCA_004023805.1_PanOnc_v1_BIUU \9690 \\ 71 \harbor seal \Carnivora \Phoca vitulina \HLphoVit1 GCA_004348235.1 \9720 \\ 72 \cheetah \Carnivora \Acinonyx jubatus \HLaciJub2 GCF_003709585.1_Aci_jub_2 \32536 \\ 73 \gray seal \Carnivora \Halichoerus grypus \HLhalGry1 GCA_012393455.1 \9711 \\ 74 \Hawaiian monk seal \Carnivora \Neomonachus schauinslandi \Jun. 2017 (ASM220157v1/neoSch1) \29088 \\ 75 \Weddell seal \Carnivora \Leptonychotes weddellii \Mar 2013 (LepWed1.0/lepWed1) \9713 \\ 76 \jaguar \Carnivora \Panthera onca \DNA zoo Panthera onca \9690 \\ 77 \Amur leopard cat \Carnivora \Prionailurus bengalensis euptilurus \HLpriBen1 GCA_005406085.1 \300877 \\ 78 \Asian black bear \Carnivora \Ursus thibetanus thibetanus \HLursThi1 GCA_009660055.1 \441215 \\ 79 \Spanish lynx \Carnivora \Lynx pardinus \HLlynPar1 GCA_900661375.1 \191816 \\ 80 \Southern elephant seal \Carnivora \Mirounga leonina \HLmirLeo1 GCA_011800145.1 \9715 \\ 81 \Canada lynx \Carnivora \Lynx canadensis \HLlynCan1 GCA_007474595.1 \61383 \\ 82 \Northern elephant seal \Carnivora \Mirounga angustirostris \DNA zoo Mirounga angustirostris \9716 \\ 83 \lion \Carnivora \Panthera leo \HLpanLeo1 GCA_008795835.1 \9689 \\ 84 \walrus \Carnivora \Odobenus rosmarus \DNA zoo Odobenus rosmarus \9707 \\ 85 \northern fur seal \Carnivora \Callorhinus ursinus \HLcalUrs1 GCA_003265705.1 \34884 \\ 86 \Pacific walrus \Carnivora \Odobenus rosmarus divergens \Jan 2013 (Oros_1.0/odoRosDiv1) \9708 \\ 87 \giant panda \Carnivora \Ailuropoda melanoleuca \HLailMel2 GCA_002007445.2 \9646 \\ 88 \California sea lion \Carnivora \Zalophus californianus \HLzalCal1 GCA_009762305.1_mZalCal1.pri \9704 \\ 89 \Steller sea lion \Carnivora \Eumetopias jubatus \HLeumJub1 GCA_004028035.1 \34886 \\ 90 \domestic cat \Carnivora \Felis catus \Nov. 2017 (Felis_catus_9.0/felCat9) \9685 \\ 91 \jaguarundi \Carnivora \Puma yagouaroundi \HLpumYag1 GCA_014898765.1 \1608482 \\ 92 \grizzly bear \Carnivora \Ursus arctos horribilis \HLursArc1 GCA_003584765.1 \116960 \\ 93 \polar bear \Carnivora \Ursus maritimus \09 May-2014 (UrsMar_1.0/ursMar1) \29073 \\ 94 \antarctic fur seal \Carnivora \Arctocephalus gazella \HLarcGaz2 GCA_900642305.1 \37190 \\ 95 \American black bear \Carnivora \Ursus americanus \HLursAme1 GCA_003344425.1 \9643 \\ 96 \American black bear \Carnivora \Ursus americanus \DNA zoo Ursus americanus \9643 \\ 97 \black-footed cat \Carnivora \Felis nigripes \HLfelNig1 GCA_004023925.1_FelNig_v1_BIUU \61379 \\ 98 \fossa \Carnivora \Cryptoprocta ferox \DNA zoo Cryptoprocta ferox \94188 \\ 99 \red fox \Carnivora \Vulpes vulpes \HLvulVul1 GCA_003160815.1 \9627 \\ 100 \dog \Carnivora \Canis lupus familiaris \Mar. 2020 (UU_Cfam_GSD_1.0/canFam4) \9615 \\ 101 \Arctic fox \Carnivora \Vulpes lagopus \HLvulLag1 GCA_004023825.1_VulLag_v1_BIUU \494514 \\ 102 \African hunting dog \Carnivora \Lycaon pictus \DNA zoo Lycaon pictus \9622 \\ 103 \dingo \Carnivora \Canis lupus dingo \HLcanLupDin1 GCA_003254725.1 \286419 \\ 104 \dog \Carnivora \Canis lupus familiaris \May 2019 (UMICH_Zoey_3.1/canFam5) \9615 \\ 105 \kinkajou \Carnivora \Potos flavus \DNA zoo Potos flavus \29067 \\ 106 \African hunting dog \Carnivora \Lycaon pictus \HLlycPic2 GCA_004216515.1 \9622 \\ 107 \lesser panda \Carnivora \Ailurus fulgens \DNA zoo Ailurus fulgens \9649 \\ 108 \spotted hyena \Carnivora \Crocuta crocuta \HLcroCro1 GCA_008692635.1 \9678 \\ 109 \striped hyena \Carnivora \Hyaena hyaena \HLhyaHya1 GCA_003009895.1 \95912 \\ 110 \Asian palm civet \Carnivora \Paradoxurus hermaphroditus \HLparHer1 GCA_004024585.1_ParHer_v1_BIUU \71117 \\ 111 \White-nosed coati \Carnivora \Nasua narica \DNA zoo Nasua narica \352831 \\ 112 \sable \Carnivora \Martes zibellina \HLmarZib1 GCA_012583365.1 \36722 \\ 113 \wolverine \Carnivora \Gulo gulo \HLgulGul1 GCA_900006375.2 \48420 \\ 114 \raccoon \Carnivora \Procyon lotor \DNA zoo Procyon lotor \9654 \\ 115 \Cacomistle \Carnivora \Bassariscus sumichrasti \DNA zoo Bassariscus sumichrasti \392507 \\ 116 \western spotted skunk \Carnivora \Spilogale gracilis \HLspiGra1 GCA_004023965.1_SpiGra_v1_BIUU \30551 \\ 117 \North American badger \Carnivora \Taxidea taxus jeffersonii \HLtaxTax1 GCA_003697995.1 \2282171 \\ 118 \ratel \Carnivora \Mellivora capensis \HLmelCap1 GCA_004024625.1_MelCap_v1_BIUU \9664 \\ 119 \meerkat \Carnivora \Suricata suricatta \HLsurSur2 GCA_004023905.1_SurSur_v1_BIUU \37032 \\ 120 \meerkat \Carnivora \Suricata suricatta \HLsurSur1 GCA_006229205.1 \37032 \\ 121 \banded mongoose \Carnivora \Mungos mungo \HLmunMug1 GCA_004023785.1_MunMun_v1_BIUU \210652 \\ 122 \dwarf mongoose \Carnivora \Helogale parvula \HLhelPar1 GCA_004023845.1_HelPar_v1_BIUU \210647 \\ 123 \Northern American river otter \Carnivora \Lontra canadensis \HLlonCan1 GCA_010015895.1 \76717 \\ 124 \giant otter \Carnivora \Pteronura brasiliensis \DNA zoo Pteronura brasiliensis \9672 \\ 125 \giant otter \Carnivora \Pteronura brasiliensis \HLpteBra1 GCA_004024605.1_PteBra_v1_BIUU \9672 \\ 126 \Southern sea otter \Carnivora \Enhydra lutris nereis \Jun. 2019 (ASM641071v1/enhLutNer1) \1049777 \\ 127 \Northern sea otter \Carnivora \Enhydra lutris kenyoni \Sep. 2017 (ASM228890v2/enhLutKen1) \391180 \\ 128 \Eurasian river otter \Carnivora \Lutra lutra \HLlutLut1 GCA_902655055.1 \9657 \\ 129 \ermine \Carnivora \Mustela erminea \HLmusErm1 GCA_009829155.1 \36723 \\ 130 \American mink \Carnivora \Neovison vison \HLneoVis1 GCA_900108605.1_NNQGG.v01 \452646 \\ 131 \European polecat \Carnivora \Mustela putorius \HLmusPut1 GCA_902460205.1 \9668 \\ 132 \domestic ferret \Carnivora \Mustela putorius furo \HLmusFur2 GCA_011764305.1 \9669 \\ 133 \Brazilian tapir \Laurasiatheria \Tapirus terrestris \HLtapTer1 GCA_004025025.1_TapTer_v1_BIUU \9801 \\ 134 \greater Indian rhinoceros \Laurasiatheria \Rhinoceros unicornis \DNA zoo Rhinoceros unicornis \9809 \\ 135 \Asiatic tapir \Laurasiatheria \Tapirus indicus \HLtapInd1 GCA_004024905.1_TapInd_v1_BIUU \9802 \\ 136 \Asiatic tapir \Laurasiatheria \Tapirus indicus \DNA zoo Tapirus indicus \9802 \\ 137 \black rhinoceros \Laurasiatheria \Diceros bicornis \HLdicBic1 GCA_004027315.2 \9805 \\ 138 \Sumatran rhinoceros \Laurasiatheria \Dicerorhinus sumatrensis sumatrensis \HLdicSum1 GCA_002844835.1_ASM284483v1 \310712 \\ 139 \northern white rhinoceros \Laurasiatheria \Ceratotherium simum cottoni \HLcerSimCot1 GCA_004027795.1_CerCot_v1_BIUU \310713 \\ 140 \southern white rhinoceros \Laurasiatheria \Ceratotherium simum simum \May 2012 (CerSimSim1.0/cerSim1) \73337 \\ 141 \Equus burchelli boehmi \Laurasiatheria \Equus burchellii boehmi \DNA zoo Equus burchellii boehmi \89250 \\ 142 \horse \Laurasiatheria \Equus caballus \Jan. 2018 (EquCab3.0/equCab3) \9796 \\ 143 \Przewalski's horse \Laurasiatheria \Equus przewalskii \Jun 2014 (Burgud/equPrz1) \9798 \\ 144 \ass \Laurasiatheria \Equus asinus \HLequAsi1 GCA_001305755.1_ASM130575v1 \9793 \\ 145 \donkey \Laurasiatheria \Equus asinus asinus \HLequAsiAsi2 GCA_003033725.1 \83772 \\ 146 \Tree pangolin \Laurasiatheria \Manis tricuspis \HLmanTri1 GCA_004765945.1 \358128 \\ 147 \Tree pangolin \Laurasiatheria \Manis tricuspis \DNA zoo Manis tricuspis \358128 \\ 148 \Chinese pangolin \Laurasiatheria \Manis pentadactyla \HLmanPen2 GCA_014570555.1 \143292 \\ 149 \Chinese pangolin \Laurasiatheria \Manis pentadactyla \Aug 2014 (M_pentadactyla-1.1.1/manPen1) \143292 \\ 150 \Malayan pangolin \Laurasiatheria \Manis javanica \HLmanJav1 GCA_001685135.1_ManJav1.0 \9974 \\ 151 \Malayan pangolin \Laurasiatheria \Manis javanica \HLmanJav2 GCA_014570535.1 \9974 \\ 152 \Hispaniolan solenodon \Laurasiatheria \Solenodon paradoxus \HLsolPar1 GCA_004363575.1_SolPar_v1_BIUU \79805 \\ 153 \eastern mole \Laurasiatheria \Scalopus aquaticus \HLscaAqu1 GCA_004024925.1_ScaAqu_v1_BIUU \71119 \\ 154 \Iberian mole \Laurasiatheria \Talpa occidentalis \HLtalOcc1 GCA_014898055.1 \50954 \\ 155 \gracile shrew mole \Laurasiatheria \Uropsilus gracilis \HLuroGra1 GCA_004024945.1_UroGra_v1_BIUU \182669 \\ 156 \star-nosed mole \Laurasiatheria \Condylura cristata \Mar 2012 (ConCri1.0/conCri1) \143302 \\ 157 \western European hedgehog \Laurasiatheria \Erinaceus europaeus \May 2012 (EriEur2.0/eriEur2) \9365 \\ 158 \European shrew \Laurasiatheria \Sorex araneus \Aug. 2008 (Broad/sorAra2) \42254 \\ 159 \Antarctic minke whale \Cetartiodactyla \Balaenoptera bonaerensis \HLbalBon1 GCA_000978805.1_ASM97880v1 \33556 \\ 160 \grey whale \Cetartiodactyla \Eschrichtius robustus \HLescRob1 GCA_004363415.1_EscRob_v1_BIUU \9764 \\ 161 \sperm whale \Cetartiodactyla \Physeter catodon \Sep. 2013 (Physeter_macrocephalus-2.0.2/phyCat1) \9755 \\ 162 \sperm whale \Cetartiodactyla \Physeter catodon \HLphyCat2 GCA_002837175.2 \9755 \\ 163 \Yangtze River dolphin \Cetartiodactyla \Lipotes vexillifer \31 Jul 2013 (Lipotes_vexillifer_v1/lipVex1) \118797 \\ 164 \beluga whale \Cetartiodactyla \Delphinapterus leucas \HLdelLeu2 GCA_002288925.3 \9749 \\ 165 \hippopotamus \Cetartiodactyla \Hippopotamus amphibius \HLhipAmp3 GCA_004027065.2 \9833 \\ 166 \hippopotamus \Cetartiodactyla \Hippopotamus amphibius \HLhipAmp1 GCA_002995585.1_ASM299558v1 \9833 \\ 167 \harbor porpoise \Cetartiodactyla \Phocoena phocoena \DNA zoo Phocoena phocoena \9742 \\ 168 \harbor porpoise \Cetartiodactyla \Phocoena phocoena \HLphoPho1 GCA_004363495.1_PhoPho_v1_BIUU \9742 \\ 169 \Wild Bactrian camel \Cetartiodactyla \Camelus ferus \HLcamFer3 GCA_009834535.1 \419612 \\ 170 \killer whale \Cetartiodactyla \Orcinus orca \Jan. 2013 (Oorc_1.1/orcOrc1) \9733 \\ 171 \Bactrian camel \Cetartiodactyla \Camelus bactrianus \HLcamBac1 GCA_000767855.1_Ca_bactrianus_MBC_1.0 \9837 \\ 172 \Indo-pacific humpbacked dolphin \Cetartiodactyla \Sousa chinensis \HLsouChi1 GCA_007760645.1 \103600 \\ 173 \Arabian camel \Cetartiodactyla \Camelus dromedarius \HLcamDro2 GCA_000803125.3 \9838 \\ 174 \alpaca \Cetartiodactyla \Vicugna pacos \Mar. 2013 (Vicugna_pacos-2.0.1/vicPac2) \30538 \\ 175 \common bottlenose dolphin \Cetartiodactyla \Tursiops truncatus \HLturTru4 GCA_011762595.1_mTurTru1.mat.Y \9739 \\ 176 \Indo-pacific bottlenose dolphin \Cetartiodactyla \Tursiops aduncus \HLturAdu1 GCA_003227395.1 \79784 \\ 177 \Indo-pacific bottlenose dolphin \Cetartiodactyla \Tursiops aduncus \DNA zoo Tursiops aduncus \79784 \\ 178 \common bottlenose dolphin \Cetartiodactyla \Tursiops truncatus \Oct. 2011 (Baylor Ttru_1.4/turTru2) \9739 \\ 179 \common bottlenose dolphin \Cetartiodactyla \Tursiops truncatus \HLturTru3 GCA_001922835.1_NIST_Tur_tru_v1 \9739 \\ 180 \pig \Cetartiodactyla \Sus scrofa \Feb. 2017 (Sscrofa11.1/susScr11) \9823 \\ 181 \okapi \Cetartiodactyla \Okapia johnstoni \DNA zoo Okapia johnstoni \86973 \\ 182 \Masai giraffe \Cetartiodactyla \Giraffa tippelskirchi \HLgirTip1 GCA_001651235.1_ASM165123v1 \439328 \\ 183 \water buffalo \Cetartiodactyla \Bubalus bubalis \HLbubBub2 GCA_003121395.1 \89462 \\ 184 \zebu cattle \Cetartiodactyla \Bos indicus \HLbosInd2 GCA_002933975.1 \9915 \\ 185 \cattle \Cetartiodactyla \Bos taurus \Apr. 2018 (ARS-UCD1.2/bosTau9) \9913 \\ 186 \wild yak \Cetartiodactyla \Bos mutus \HLbosMut2 GCA_007646595.3 \72004 \\ 187 \greater kudu \Cetartiodactyla \Tragelaphus strepsiceros \HLtraStr1 GCA_006410795.1 \9946 \\ 188 \aoudad \Cetartiodactyla \Ammotragus lervia \HLammLer1 GCA_002201775.1_ALER1.0 \9899 \\ 189 \goat \Cetartiodactyla \Capra hircus \HLcapHir2 GCA_001704415.1_ARS1 \9925 \\ 190 \wild goat \Cetartiodactyla \Capra aegagrus \HLcapAeg1 GCA_000765075.1 \9923 \\ 191 \chiru \Cetartiodactyla \Pantholops hodgsonii \May 2013 (PHO1.0/panHod1) \59538 \\ 192 \white-tailed deer \Cetartiodactyla \Odocoileus virginianus \HLodoVir3 GCA_014726795.1 \9874 \\ 193 \bighorn sheep \Cetartiodactyla \Ovis canadensis \HLoviCan2 GCA_004026945.1_OviCan_v1_BIUU \37174 \\ 194 \white-tailed deer \Cetartiodactyla \Odocoileus virginianus \DNA zoo Odocoileus virginianus \9874 \\ 195 \sheep \Cetartiodactyla \Ovis aries \HLoviAri5 GCA_011170295.1 \9940 \\ 196 \Pere David's deer \Cetartiodactyla \Elaphurus davidianus \HLelaDav1 GCA_002443075.1_Milu1.0 \43332 \\ 197 \argali \Cetartiodactyla \Ovis ammon \HLoviAmm1 GCA_003121645.1 \30527 \\ 198 \North Atlantic right whale \Artiodactyla \Eubalaena glacialis \DNA zoo Eubalaena glacialis \27606 \\ 199 \North Pacific right whale \Artiodactyla \Eubalaena japonica \HLeubJap1 GCA_004363455.1_EubJap_v1_BIUU \302098 \\ 200 \minke whale \Artiodactyla \Balaenoptera acutorostrata scammoni \Oct. 2013 (BalAcu1.0/balAcu1) \310752 \\ 201 \humpback whale \Artiodactyla \Megaptera novaeangliae \HLmegNov1 GCA_004329385.1 \9773 \\ 202 \Fin whale \Artiodactyla \Balaenoptera physalus \HLbalPhy1 GCA_008795845.1 \9770 \\ 203 \bowhead whale \Artiodactyla \Balaena mysticetus \HLbalMys1/http://alfred.liv.ac.uk/downloads/bowhead_whale/bowhead_whale_scaffolds.zip/none \27602 \\ 204 \Blue whale \Artiodactyla \Balaenoptera musculus \HLbalMus1 GCA_009873245.1 \9771 \\ 205 \pygmy Bryde's whale \Artiodactyla \Balaenoptera edeni \DNA zoo Balaenoptera edeni \9769 \\ 206 \Sowerby's beaked whale \Artiodactyla \Mesoplodon bidens \HLmesBid1 GCA_004027085.1_MesBid_v1_BIUU \48745 \\ 207 \Indus River dolphin \Artiodactyla \Platanista minor \HLplaMin1 GCA_004363435.1_PlaMin_v1_BIUU \48752 \\ 208 \Cuvier's beaked whale \Artiodactyla \Ziphius cavirostris \HLzipCav1 GCA_004364475.1_ZipCav_v1_BIUU \9760 \\ 209 \boutu \Artiodactyla \Inia geoffrensis \HLlniGeo1 GCA_004363515.1_IniGeo_v1_BIUU \9725 \\ 210 \narwhal \Artiodactyla \Monodon monoceros \HLmonMon1 GCA_005190385.2 \40151 \\ 211 \Yangtze finless porpoise \Artiodactyla \Neophocaena asiaeorientalis asiaeorientalis \HLneoAsi1 GCA_003031525.1_Neophocaena_asiaeorientalis_V1 \1706337 \\ 212 \pygmy sperm whale \Artiodactyla \Kogia breviceps \HLkogBre1 GCA_004363705.1_KogBre_v1_BIUU \27615 \\ 213 \vaquita \Artiodactyla \Phocoena sinus \HLphoSin1 GCA_008692025.1 \42100 \\ 214 \franciscana \Artiodactyla \Pontoporia blainvillei \HLponBla1 GCA_011754075.1 \48723 \\ 215 \Lama pacos huacaya \Artiodactyla \Vicugna pacos huacaya \HLvicPacHua3 GCA_000767525.1_Vi_pacos_V1.0 \273913 \\ 216 \llama \Artiodactyla \Lama glama \DNA zoo Lama glama \9844 \\ 217 \melon-headed whale \Artiodactyla \Peponocephala electra \DNA zoo Peponocephala electra \103596 \\ 218 \long-finned pilot whale \Artiodactyla \Globicephala melas \HLgloMel1 GCA_006547405.1 \9731 \\ 219 \Pacific white-sided dolphin \Artiodactyla \Lagenorhynchus obliquidens \HLlagObl1 GCA_003676395.1 \90247 \\ 220 \Vicugna mensalis \Artiodactyla \Vicugna vicugna mensalis \HLvicVicMen1 GCA_013265495.1 \273917 \\ 221 \guanaco \Artiodactyla \Lama guanicoe cacsilensis \HLlamGuaCac1 GCA_013239625.1 \273908 \\ 222 \llama \Artiodactyla \Lama glama chaku \HLlamGlaCha1 GCA_013239585.1 \273914 \\ 223 \Chacoan peccary \Artiodactyla \Catagonus wagneri \HLcatWag1 GCA_004024745.2_CatWag_v2_BIUU_UCD \51154 \\ 224 \giraffe \Artiodactyla \Giraffa camelopardalis \HLgirCam1 GCA_006408565.1 \9894 \\ 225 \giraffe \Artiodactyla \Giraffa camelopardalis \DNA zoo Giraffa camelopardalis \9894 \\ 226 \African buffalo \Artiodactyla \Syncerus caffer \HLsynCaf1 GCA_902500845.1 \9970 \\ 227 \Bos bison bison \Artiodactyla \Bison bison bison \Oct. 2014 (Bison_UMD1.0/bisBis1) \43346 \\ 228 \Chinese forest musk deer \Artiodactyla \Moschus berezovskii \HLmosBer1 GCA_006459085.1 \68408 \\ 229 \Siberian musk deer \Artiodactyla \Moschus moschiferus \HLmosMos1 GCA_004024705.2 \68415 \\ 230 \alpine musk deer \Artiodactyla \Moschus chrysogaster \HLmosChr1 GCA_006461725.1 \68412 \\ 231 \Yarkand deer \Artiodactyla \Cervus hanglu yarkandensis \HLcerHanYar1 GCA_010411085.1 \84702 \\ 232 \gaur \Artiodactyla \Bos gaurus \HLbosGau1 GCA_014182915.1 \9904 \\ 233 \gayal \Artiodactyla \Bos frontalis \HLbosFro1 GCA_007844835.1_NRC_Mithun_1 \30520 \\ 234 \white-lipped deer \Artiodactyla \Przewalskium albirostris \HLprzAlb1 GCA_006408465.1 \1088058 \\ 235 \roan antelope \Artiodactyla \Hippotragus equinus \HLhipEqu1 GCA_016433095.1 \37186 \\ 236 \Harvey's duiker \Artiodactyla \Cephalophus harveyi \HLcepHar1 GCA_006410635.1 \129224 \\ 237 \sable antelope \Artiodactyla \Hippotragus niger niger \HLhipNig1 GCA_006942125.1 \82127 \\ 238 \domestic yak \Artiodactyla \Bos grunniens \HLbosGru1 GCA_005887515.2 \30521 \\ 239 \scimitar-horned oryx \Artiodactyla \Oryx dammah \DNA zoo Oryx dammah \59534 \\ 240 \bush duiker \Artiodactyla \Sylvicapra grimmia \HLsylGri1 GCA_006408735.1 \119562 \\ 241 \Maxwell's duiker \Artiodactyla \Philantomba maxwellii \HLphiMax1 GCA_006410695.1 \907741 \\ 242 \gemsbok \Artiodactyla \Oryx gazella \HLoryGaz1 GCA_003945745.1 \9958 \\ 243 \pronghorn \Artiodactyla \Antilocapra americana \HLantAme1 GCA_007570785.1 \9891 \\ 244 \Reeves' muntjac \Artiodactyla \Muntiacus reevesi \HLmunRee1 GCA_008787405.1 \9886 \\ 245 \black muntjac \Artiodactyla \Muntiacus crinifrons \HLmunCri1 GCA_006408485.1 \71854 \\ 246 \Central European red deer \Artiodactyla \Cervus elaphus hippelaphus \HLcerEla1 GCA_002197005.1 \46360 \\ 247 \lesser kudu \Artiodactyla \Tragelaphus imberbis \HLtraImb1 GCA_006410775.1 \9947 \\ 248 \brindled gnu \Artiodactyla \Connochaetes taurinus \DNA zoo Connochaetes taurinus \9927 \\ 249 \bushbuck \Artiodactyla \Tragelaphus scriptus \HLtraScr1 GCA_006410495.1 \66440 \\ 250 \waterbuck \Artiodactyla \Kobus ellipsiprymnus \HLkobEll1 GCA_006410655.1 \9962 \\ 251 \muntjak \Artiodactyla \Muntiacus muntjak \HLmunMun1 GCA_008782695.1 \9888 \\ 252 \topi \Artiodactyla \Damaliscus lunatus \HLdamLun1 GCA_006408505.1 \9929 \\ 253 \bighorn sheep \Artiodactyla \Ovis canadensis canadensis \HLoviCan1 GCA_001039535.1 \112262 \\ 254 \lechwe \Artiodactyla \Kobus leche leche \HLkobLecLec1 GCA_014926565.1 \91880 \\ 255 \Eastern roe deer \Artiodactyla \Capreolus pygargus \HLcapPyg1 GCA_012922965.1 \48560 \\ 256 \Eurasian elk \Artiodactyla \Alces alces \HLalcAlc1 GCA_007570765.1 \9852 \\ 257 \Cobus hunteri \Artiodactyla \Beatragus hunteri \HLbeaHun1 GCA_004027495.1_BeaHun_v1_BIUU \59527 \\ 258 \impala \Artiodactyla \Aepyceros melampus \HLaepMel1 GCA_006408695.1 \9897 \\ 259 \mule deer \Artiodactyla \Odocoileus hemionus hemionus \HLodoHem1 GCA_004115125.1 \9877 \\ 260 \Bohar reedbuck \Artiodactyla \Redunca redunca \HLredRed1 GCA_006410935.1 \59556 \\ 261 \Siberian ibex \Artiodactyla \Capra sibirica \HLcapSib1 GCA_003182615.2 \72544 \\ 262 \porcupine caribou \Artiodactyla \Rangifer tarandus granti \HLranTarGra2 GCA_014898785.1 \191431 \\ 263 \reindeer \Artiodactyla \Rangifer tarandus \HLranTar1 GCA_004026565.1_RanTarSib_v1_BIUU \9870 \\ 264 \klipspringer \Artiodactyla \Oreotragus oreotragus \HLoreOre1 GCA_006410675.1 \66444 \\ 265 \Chinese water deer \Artiodactyla \Hydropotes inermis \HLhydIne1 GCA_006459105.1 \9883 \\ 266 \snow sheep \Artiodactyla \Ovis nivicola lydekkeri \HLoviNivLyd1 GCA_903231385.1 \1867112 \\ 267 \suni \Artiodactyla \Neotragus moschatus \HLneoMos1 GCA_006410615.1 \66442 \\ 268 \white-tailed deer \Artiodactyla \Odocoileus virginianus texanus \HLodoVir1 GCA_002102435.1_Ovir.te_1.0 \9880 \\ 269 \Nilgiri tahr \Artiodactyla \Hemitragus hylocrius \HLhemHyl1 GCA_004026825.1_HemHyl_v1_BIUU \330464 \\ 270 \Asiatic mouflon \Artiodactyla \Ovis orientalis \HLoviOri1 GCA_014523465.1 \469796 \\ 271 \royal antelope \Artiodactyla \Neotragus pygmaeus \HLneoPyg1 GCA_006410875.1 \1027985 \\ 272 \Grant's gazelle \Artiodactyla \Nanger granti \HLnanGra1 GCA_006408635.1 \27591 \\ 273 \Przewalski's gazelle \Artiodactyla \Procapra przewalskii \HLproPrz1 GCA_006410515.1 \157668 \\ 274 \steenbok \Artiodactyla \Raphicerus campestris \HLrapCam1 GCA_006410735.1 \59544 \\ 275 \Thomson's gazelle \Artiodactyla \Eudorcas thomsonii \HLeudTho1 GCA_006408755.1 \69308 \\ 276 \springbok \Artiodactyla \Antidorcas marsupialis \HLantMar1 GCA_006408585.1 \59523 \\ 277 \gerenuk \Artiodactyla \Litocranius walleri \HLlitWal1 GCA_006410535.1 \69311 \\ 278 \Kirk's dik-dik \Artiodactyla \Madoqua kirkii \HLmadKir1 GCA_006408675.1 \66434 \\ 279 \Hog deer \Artiodactyla \Axis porcinus \HLaxiPor1 GCA_003798545.1 \57737 \\ 280 \Java mouse-deer \Artiodactyla \Tragulus javanicus \HLtraJav1 GCA_004024965.2 \9849 \\ 281 \lesser mouse-deer \Artiodactyla \Tragulus kanchil \HLtraKan1 GCA_006408655.1 \1088131 \\ 282 \mountain goat \Artiodactyla \Oreamnos americanus \HLoreAme1 GCA_009758055.1 \34873 \\ 283 \saiga antelope \Artiodactyla \Saiga tatarica \HLsaiTat1 GCA_004024985.1_SaiTat_v1_BIUU \34875 \\ 284 \Alpine ibex \Artiodactyla \Capra ibex \HLcapIbe1 GCA_006410555.1 \72542 \\ 285 \Hoffmann's two-fingered sloth \Xenarthra \Choloepus hoffmanni \DNA zoo Choloepus hoffmanni \9358 \\ 286 \southern two-toed sloth \Xenarthra \Choloepus didactylus \HLchoDid2 GCF_015220235.1_mChoDid1.pri \27675 \\ 287 \southern two-toed sloth \Xenarthra \Choloepus didactylus \HLchoDid1 GCA_004027855.1_ChoDid_v1_BIUU \27675 \\ 288 \nine-banded armadillo \Xenarthra \Dasypus novemcinctus \Dec. 2011 (Baylor/dasNov3) \9361 \\ 289 \giant anteater \Xenarthra \Myrmecophaga tridactyla \HLmyrTri1 GCA_004026745.1_MyrTri_v1_BIUU \71006 \\ 290 \southern tamandua \Xenarthra \Tamandua tetradactyla \HLtamTet1 GCA_004025105.1_TamTet_v1_BIUU \48850 \\ 291 \Southern three-banded armadillo \Xenarthra \Tolypeutes matacus \HLtolMat1 GCA_004025125.1_TolMat_v1_BIUU \183749 \\ 292 \Chinese rufous horseshoe bat \Chiroptera \Rhinolophus sinicus \HLrhiSin1 GCA_001888835.1_ASM188883v1 \89399 \\ 293 \great roundleaf bat \Chiroptera \Hipposideros armiger \HLhipArm1 GCA_001890085.1_ASM189008v1 \186990 \\ 294 \black flying fox \Chiroptera \Pteropus alecto \Aug 2012 (ASM32557v1/pteAle1) \9402 \\ 295 \greater horseshoe bat \Chiroptera \Rhinolophus ferrumequinum \HLrhiFer5/Bat1K published/none \59479 \\ 296 \Bonin flying fox \Chiroptera \Pteropus pselaphon \HLptePse1 GCA_014363405.1 \1496133 \\ 297 \Brazilian free-tailed bat \Chiroptera \Tadarida brasiliensis \HLtadBra1 GCA_004025005.1_TadBra_v1_BIUU \9438 \\ 298 \large flying fox \Chiroptera \Pteropus vampyrus \HLpteVam2 GCA_000151845.2 \132908 \\ 299 \Malagasy flying fox \Chiroptera \Pteropus rufus \DNA zoo Pteropus rufus \196297 \\ 300 \Indian flying fox \Chiroptera \Pteropus giganteus \HLpteGig1 GCA_902729225.1 \143291 \\ 301 \Malagasy straw-colored fruit bat \Chiroptera \Eidolon dupreanum \DNA zoo Eidolon dupreanum \58063 \\ 302 \straw-colored fruit bat \Chiroptera \Eidolon helvum \HLeidHel2/DNAZoo/none \77214 \\ 303 \Cantor's roundleaf bat \Chiroptera \Hipposideros galeritus \HLhipGal1 GCA_004027415.1_HipGal_v1_BIUU \58069 \\ 304 \lesser short-nosed fruit bat \Chiroptera \Cynopterus brachyotis \HLcynBra1 GCA_009793145.1 \58060 \\ 305 \lesser dawn bat \Chiroptera \Eonycteris spelaea \HLeonSpe1 GCA_003508835.1 \58065 \\ 306 \Leschenault's rousette \Chiroptera \Rousettus leschenaultii \HLrouLes1 GCA_015472975.1 \9408 \\ 307 \Egyptian rousette \Chiroptera \Rousettus aegyptiacus \HLrouAeg4/Bat1K published/none \9407 \\ 308 \Madagascan rousette \Chiroptera \Rousettus madagascariensis \DNA zoo Rousettus madagascariensis \77223 \\ 309 \Indian false vampire \Chiroptera \Megaderma lyra \HLmegLyr2 GCA_004026885.1_MegLyr_v1_BIUU \9413 \\ 310 \Pallas's mastiff bat \Chiroptera \Molossus molossus \HLmolMol2/Bat1K published/none \27622 \\ 311 \long-tongued fruit bat \Chiroptera \Macroglossus sobrinus \HLmacSob1 GCA_004027375.1_MacSob_v1_BIUU \326083 \\ 312 \Schreibers' long-fingered bat \Chiroptera \Miniopterus schreibersii \HLminSch1 GCA_004026525.1_MinSch_v1_BIUU \9433 \\ 313 \Miniopterus schreibersii natalensis \Chiroptera \Miniopterus natalensis \HLminNat1 GCA_001595765.1 \291302 \\ 314 \hog-nosed bat \Chiroptera \Craseonycteris thonglongyai \HLcraTho1 GCA_004027555.1_CraTho_v1_BIUU \208972 \\ 315 \Antillean ghost-faced bat \Chiroptera \Mormoops blainvillei \HLmorBla1 GCA_004026545.1_MorMeg_v1_BIUU \118852 \\ 316 \Parnell's mustached bat \Chiroptera \Pteronotus parnellii \Sep. 2013 (ASM46540v1/ptePar1) \59476 \\ 317 \big brown bat \Chiroptera \Eptesicus fuscus \Jul 2012 (EptFus1.0/eptFus1) \29078 \\ 318 \greater mouse-eared bat \Chiroptera \Myotis myotis \HLmyoMyo6/Bat1K published/none \51298 \\ 319 \Brandt's bat \Chiroptera \Myotis brandtii \28 Jun 2013 (ASM41265v1/myoBra1) \109478 \\ 320 \common vampire bat \Chiroptera \Desmodus rotundus \HLdesRot2 \9430 \\ 321 \California big-eared bat \Chiroptera \Macrotus californicus \HLmacCal1 GCA_007922815.1 \9419 \\ 322 \Northern long-eared myotis \Chiroptera \Myotis septentrionalis \DNA zoo Myotis septentrionalis \258941 \\ 323 \little brown bat \Chiroptera \Myotis lucifugus \DNA zoo Myotis lucifugus \59463 \\ 324 \little brown bat \Chiroptera \Myotis lucifugus \Jul. 2010 (Broad Institute Myoluc2.0/myoLuc2) \59463 \\ 325 \Lesser long-nosed bat \Chiroptera \Leptonycteris yerbabuenae \HLlepYer1/GIGADB/none \700936 \\ 326 \Vespertilio Davidii \Chiroptera \Myotis davidii \Aug 2012 (ASM32734v1/myoDav1) \225400 \\ 327 \Schizostoma hirsutum \Chiroptera \Micronycteris hirsuta \HLmicHir1 GCA_004026765.1_MicHir_v1_BIUU \148065 \\ 328 \tailed tailless bat \Chiroptera \Anoura caudifer \HLanoCau1 GCA_004027475.1_AnoCau_v1_BIUU \27642 \\ 329 \Murina feae \Chiroptera \Murina aurata feae \HLmurAurFea1 GCA_004026665.1_MurFea_v1_BIUU \1453894 \\ 330 \greater bulldog bat \Chiroptera \Noctilio leporinus \HLnocLep1 GCA_004026585.1_NocLep_v1_BIUU \94963 \\ 331 \Seba's short-tailed bat \Chiroptera \Carollia perspicillata \HLcarPer3 GCA_004027735.1_CarPer_v1_BIUU \40233 \\ 332 \pale spear-nosed bat \Chiroptera \Phyllostomus discolor \HLphyDis3/Bat1K published/none \89673 \\ 333 \stripe-headed round-eared bat \Chiroptera \Tonatia saurophila \HLtonSau1 GCA_004024845.1_TonSau_v1_BIUU \171122 \\ 334 \Jamaican fruit-eating bat \Chiroptera \Artibeus jamaicensis \HLartJam1 GCA_004027435.1_ArtJam_v1_BIUU \9417 \\ 335 \Jamaican fruit-eating bat \Chiroptera \Artibeus jamaicensis \HLartJam2 GCA_014825515.1 \9417 \\ 336 \Honduran yellow-shouldered bat \Chiroptera \Sturnira hondurensis \HLstuHon1 GCA_014824575.1 \192404 \\ 337 \hoary bat \Chiroptera \Aeorestes cinereus \HLaeoCin1 GCA_011751065.1 \257879 \\ 338 \pallid bat \Chiroptera \Antrozous pallidus \HLantPal1 GCA_007922775.1 \9440 \\ 339 \evening bat \Chiroptera \Nycticeius humeralis \HLnycHum2 GCA_007922795.1 \27670 \\ 340 \red bat \Chiroptera \Lasiurus borealis \HLlasBor1 GCA_004026805.1_LasBor_v1_BIUU \258930 \\ 341 \Kuhl's pipistrelle \Chiroptera \Pipistrellus kuhlii \HLpipKuh2/Bat1K published/none \59472 \\ 342 \common pipistrelle \Chiroptera \Pipistrellus pipistrellus \HLpipPip1 GCA_004026625.1_PipPip_v1_BIUU \59474 \\ 343 \common pipistrelle \Chiroptera \Pipistrellus pipistrellus \HLpipPip2 GCA_903992545.1 \59474 \\ 344 \gray squirrel \Glires \Sciurus carolinensis \HLsciCar1 GCA_902686445.1 \30640 \\ 345 \Eurasian red squirrel \Glires \Sciurus vulgaris \HLsciVul1 GCA_902686455.1_mSciVul1.1 \55149 \\ 346 \South African ground squirrel \Glires \Xerus inauris \HLxerIna1 GCA_004024805.1_XerIna_v1_BIUU \234690 \\ 347 \mountain beaver \Glires \Aplodontia rufa \HLaplRuf1 GCA_004027875.1_AplRuf_v1_BIUU \51342 \\ 348 \yellow-bellied marmot \Glires \Marmota flaviventris \HLmarFla1 GCA_003676075.2 \93162 \\ 349 \Alpine marmot \Glires \Marmota marmota marmota \HLmarMar1 GCF_001458135.1_marMar2.1 \9994 \\ 350 \Vancouver Island marmot \Glires \Marmota vancouverensis \HLmarVan1 GCA_005458795.1 \93167 \\ 351 \Himalayan marmot \Glires \Marmota himalayana \HLmarHim1 GCA_005280165.1 \93163 \\ 352 \Daurian ground squirrel \Glires \Spermophilus dauricus \HLspeDau1 GCA_002406435.1_ASM240643v1 \99837 \\ 353 \woodchuck \Glires \Marmota monax \HLmarMon1 GCA_901343595.1_MONAX5 \9995 \\ 354 \woodchuck \Glires \Marmota monax \HLmarMon2 GCA_014533835.1 \9995 \\ 355 \Arctic ground squirrel \Glires \Urocitellus parryii \HLuroPar1 GCA_003426925.1 \9999 \\ 356 \Gunnison's prairie dog \Glires \Cynomys gunnisoni \HLcynGun1 GCA_011316645.1 \45479 \\ 357 \thirteen-lined ground squirrel \Glires \Ictidomys tridecemlineatus \Nov. 2011 (Broad/speTri2) \43179 \\ 358 \Fat dormouse \Glires \Glis glis \HLgliGli1 GCA_004027185.1_GliGli_v1_BIUU \41261 \\ 359 \springhare \Glires \Pedetes capensis \HLpedCap1 GCA_007922755.1 \10023 \\ 360 \American beaver \Glires \Castor canadensis \DNA zoo Castor canadensis \51338 \\ 361 \woodland dormouse \Glires \Graphiurus murinus \HLgraMur1 GCA_004027655.1_GraMur_v1_BIUU \51346 \\ 362 \Mountain hare \Glires \Lepus timidus \HLlepTim1 GCA_009760805.1 \62621 \\ 363 \snowshoe hare \Glires \Lepus americanus \HLlepAme1 GCA_004026855.1_LepAme_v1_BIUU \48086 \\ 364 \European rabbit \Glires \Oryctolagus cuniculus cuniculus \HLoryCunCun4 GCA_013371645.1 \568996 \\ 365 \rabbit \Glires \Oryctolagus cuniculus \Apr. 2009 (Broad/oryCun2) \9986 \\ 366 \rabbit \Glires \Oryctolagus cuniculus \HLoryCun3 GCA_009806435.1 \9986 \\ 367 \brush rabbit \Glires \Sylvilagus bachmani \DNA zoo Sylvilagus bachmani \365149 \\ 368 \crested porcupine \Glires \Hystrix cristata \HLhysCri1 GCA_004026905.1_HysCri_v1_BIUU \10137 \\ 369 \North American porcupine \Glires \Erethizon dorsatum \HLereDor1 GCA_006547115.1 \34844 \\ 370 \Brazilian porcupine \Glires \Coendou prehensilis \DNA zoo Coendou prehensilis \187985 \\ 371 \hazel dormouse \Glires \Muscardinus avellanarius \HLmusAve1 GCA_004027005.1_MusAve_v1_BIUU \39082 \\ 372 \naked mole-rat \Glires \Heterocephalus glaber \Jan. 2012 (Broad HetGla_female_1.0/hetGla2) \10181 \\ 373 \Damara mole-rat \Glires \Fukomys damarensis \HLfukDam2 GCA_012274545.1 \885580 \\ 374 \Upper Galilee mountains blind mole rat \Glires \Nannospalax galili \Jun 2014 (S.galili_v1.0/nanGal1) \1026970 \\ 375 \long-tailed chinchilla \Glires \Chinchilla lanigera \May 2012 (ChiLan1.0/chiLan1) \34839 \\ 376 \punctate agouti \Glires \Dasyprocta punctata \HLdasPun1 GCA_004363535.1_DasPun_v1_BIUU \34846 \\ 377 \northern gundi \Glires \Ctenodactylus gundi \HLcteGun1 GCA_004027205.1_CteGun_v1_BIUU \10166 \\ 378 \Gobi jerboa \Glires \Allactaga bullata \HLallBul1 GCA_004027895.1_AllBul_v1_BIUU \1041416 \\ 379 \Stephens's kangaroo rat \Glires \Dipodomys stephensi \HLdipSte1 GCA_004024685.1_DipSte_v1_BIUU \323379 \\ 380 \Ord's kangaroo rat \Glires \Dipodomys ordii \Dec. 2014 (Dord_2.0/dipOrd2) \10020 \\ 381 \hoary bamboo rat \Glires \Rhizomys pruinosus \HLrhiPru1 GCA_009823505.1 \53275 \\ 382 \pacarana \Glires \Dinomys branickii \HLdinBra1 GCA_004027595.1_DinBra_v1_BIUU \108858 \\ 383 \lesser Egyptian jerboa \Glires \Jaculus jaculus \May 2012 (JacJac1.0/jacJac1) \51337 \\ 384 \meadow jumping mouse \Glires \Zapus hudsonius \HLzapHud1 GCA_004024765.1_ZapHud_v1_BIUU \160400 \\ 385 \Patagonian cavy \Glires \Dolichotis patagonum \HLdolPat1 GCA_004027295.1_DolPat_v1_BIUU \29091 \\ 386 \Pacific pocket mouse \Glires \Perognathus longimembris pacificus \HLperLonPac1 GCA_004363475.1_PerLonPac_v1_BIUU \214514 \\ 387 \capybara \Glires \Hydrochoerus hydrochaeris \HLhydHyd1 GCA_004027455.1_HydHyd_v1_BIUU \10149 \\ 388 \American pika \Glires \Ochotona princeps \May 2012 (OchPri3.0/ochPri3) \9978 \\ 389 \Brazilian guinea pig \Glires \Cavia aperea \Jan. 2014 (CavAp1.0/cavApe1) \37548 \\ 390 \dassie-rat \Glires \Petromus typicus \HLpetTyp1 GCA_004026965.1_PetTyp_v1_BIUU \10183 \\ 391 \Montane guinea pig \Glires \Cavia tschudii \HLcavTsc1 GCA_004027695.1_CavTsc_v1_BIUU \143287 \\ 392 \domestic guinea pig \Glires \Cavia porcellus \Feb. 2008 (Broad/cavPor3) \10141 \\ 393 \Greater cane rat \Glires \Thryonomys swinderianus \HLthrSwi1 GCA_004025085.1_ThrSwi_v1_BIUU \10169 \\ 394 \degu \Glires \Octodon degus \Apr 2012 (OctDeg1.0/octDeg1) \10160 \\ 395 \Gambian giant pouched rat \Glires \Cricetomys gambianus \HLcriGam1 GCA_004027575.1_CriGam_v1_BIUU \10085 \\ 396 \desert woodrat \Glires \Neotoma lepida \HLneoLep1 GCA_001675575.1 \56216 \\ 397 \social tuco-tuco \Glires \Ctenomys sociabilis \HLcteSoc1 GCA_004027165.1_CteSoc_v1_BIUU \43321 \\ 398 \nutria \Glires \Myocastor coypus \HLmyoCoy1 GCA_004027025.1_MyoCoy_v1_BIUU \10157 \\ 399 \northern rock mouse \Glires \Peromyscus nasutus \DNA zoo Peromyscus nasutus \97212 \\ 400 \Chinese hamster \Glires \Cricetulus griseus \HLcriGri3 GCA_003668045.1 \10029 \\ 401 \Hesperomys crinitus \Glires \Peromyscus crinitus \DNA zoo Peromyscus crinitus \144753 \\ 402 \muskrat \Glires \Ondatra zibethicus \HLondZib1 GCA_004026605.1_OndZib_v1_BIUU \10060 \\ 403 \Peromyscus californicus subsp. insignis \Glires \Peromyscus californicus insignis \HLperCal2 GCA_007827085.2 \564181 \\ 404 \cactus mouse \Glires \Peromyscus eremicus \HLperEre1 GCA_902702925.1 \42410 \\ 405 \southern grasshopper mouse \Glires \Onychomys torridus \HLonyTor1 GCA_903995425.1 \38674 \\ 406 \golden hamster \Glires \Mesocricetus auratus \Mar 2013 (MesAur1.0/mesAur1) \10036 \\ 407 \white-footed mouse \Glires \Peromyscus leucopus \HLperLeu1 GCA_004664715.1 \10041 \\ 408 \Northern mole vole \Glires \Ellobius talpinus \HLellTal1 GCA_001685095.1_ETalpinus_0.1 \329620 \\ 409 \oldfield mouse \Glires \Peromyscus polionotus subgriseus \HLperPol1 GCA_003704135.2 \369710 \\ 410 \prairie deer mouse \Glires \Peromyscus maniculatus bairdii \HLperManBai2 GCA_003704035.1 \230844 \\ 411 \hispid cotton rat \Glires \Sigmodon hispidus \HLsigHis1 GCA_004025045.1_SigHis_v1_BIUU \42415 \\ 412 \Transcaucasian mole vole \Glires \Ellobius lutescens \HLellLut1 GCA_001685075.1_ASM168507v1 \39086 \\ 413 \Bank vole \Glires \Myodes glareolus \HLmyoGla2 GCA_902806735.1 \447135 \\ 414 \Eurasian water vole \Glires \Arvicola amphibius \HLarvAmp1 GCA_903992535.1 \1047088 \\ 415 \fat sand rat \Glires \Psammomys obesus \HLpsaObe1 GCA_002215935.2 \48139 \\ 416 \golden spiny mouse \Glires \Acomys russatus \HLacoRus1 GCA_903995435.1 \60746 \\ 417 \African woodland thicket rat \Glires \Grammomys surdaster \HLgraSur1 GCA_004785775.1 \491861 \\ 418 \African grass rat \Glires \Arvicanthis niloticus \HLarvNil1 GCA_011762505.1_mArvNil1.pat.X \61156 \\ 419 \root vole \Glires \Microtus oeconomus \HLmicOec1 GCA_007455595.1 \64717 \\ 420 \short-tailed field vole \Glires \Microtus agrestis \HLmicAgr2 GCA_902806775.1 \29092 \\ 421 \reed vole \Glires \Microtus fortis \HLmicFor1 GCA_014885135.1 \100897 \\ 422 \Egyptian spiny mouse \Glires \Acomys cahirinus \HLacoCah1 GCA_004027535.1_AcoCah_v1_BIUU \10068 \\ 423 \Common vole \Glires \Microtus arvalis \HLmicArv1 GCA_007455615.1 \47230 \\ 424 \prairie vole \Glires \Microtus ochrogaster \Oct. 2012 (MicOch1.0/micOch1) \79684 \\ 425 \great gerbil \Glires \Rhombomys opimus \HLrhoOpi1 GCA_010120015.1 \186474 \\ 426 \southern multimammate mouse \Glires \Mastomys coucha \HLmasCou1 GCA_008632895.1 \35658 \\ 427 \Mongolian gerbil \Glires \Meriones unguiculatus \HLmerUng1 GCA_002204375.1 \10047 \\ 428 \black rat \Glires \Rattus rattus \HLratRat7 GCA_011064425.1 \10117 \\ 429 \Norway rat \Glires \Rattus norvegicus \HLratNor7 GCA_015227675.1 \10116 \\ 430 \Norway rat \Glires \Rattus norvegicus \Jul. 2014 (RGSC 6.0/rn6) \10116 \\ 431 \shrew mouse \Glires \Mus pahari \HLmusPah1 GCA_900095145.2 \10093 \\ 432 \Ryukyu mouse \Glires \Mus caroli \HLmusCar1 GCA_900094665.2_CAROLI_EIJ_v1.1 \10089 \\ 433 \steppe mouse \Glires \Mus spicilegus \HLmusSpi1 GCA_003336285.1 \10103 \\ 434 \house mouse \Glires \Mus musculus \Jun. 2020 (GRCm39/mm39) \10090 \\ 435 \house mouse \Glires \Mus musculus \Dec. 2011 (GRCm38/mm10) \10090 \\ 436 \western wild mouse \Glires \Mus spretus \HLmusSpr1 GCA_001624865.1_SPRET_EiJ_v1 \10096 \\ 437 \European woodmouse \Glires \Apodemus sylvaticus \HLapoSyl1 GCA_001305905.1 \10129 \\ 438 \dugong \Afrotheria \Dugong dugon \HLdugDug1 GCA_015147995.1 \29137 \\ 439 \Florida manatee \Afrotheria \Trichechus manatus latirostris \Oct. 2011 (Broad v1.0/triMan1) \127582 \\ 440 \Asiatic elephant \Afrotheria \Elephas maximus \DNA zoo Elephas maximus \9783 \\ 441 \African savanna elephant \Afrotheria \Loxodonta africana \HLloxAfr4/ftp://ftp.broadinstitute.org/pub/assemblies/mammals/elephant/loxAfr4//none \9785 \\ 442 \aardvark \Afrotheria \Orycteropus afer afer \May 2012 (OryAfe1.0/oryAfe1) \1230840 \\ 443 \Steller's sea cow \Afrotheria \Hydrodamalis gigas \HLhydGig1 GCA_013391785.1 \63631 \\ 444 \Cape golden mole \Afrotheria \Chrysochloris asiatica \Aug 2012 (ChrAsi1.0/chrAsi1) \185453 \\ 445 \yellow-spotted hyrax \Afrotheria \Heterohyrax brucei \HLhetBru1 GCA_004026845.1_HetBruBak_v1_BIUU \77598 \\ 446 \Cape rock hyrax \Afrotheria \Procavia capensis \HLproCap3 GCA_004026925.2 \9813 \\ 447 \Cape elephant shrew \Afrotheria \Elephantulus edwardii \Aug 2012 (EleEdw1.0/eleEdw1) \28737 \\ 448 \small Madagascar hedgehog \Afrotheria \Echinops telfairi \Nov. 2012 (Broad/echTel2) \9371 \\ 449 \Talazac's shrew tenrec \Afrotheria \Microgale talazaci \HLmicTal1 GCA_004026705.1_MicTal_v1_BIUU \176115 \\ 450 \common wombat \Metatheria \Vombatus ursinus \HLvomUrs1 GCA_900497805.2 \29139 \\ 451 \koala \Metatheria \Phascolarctos cinereus \HLphaCin1 GCA_002099425.1 \38626 \\ 452 \Agile Gracile Mouse Opossum \Metatheria \Gracilinanus agilis \HLgraAgi1 GCA_016433145.1 \191870 \\ 453 \common brushtail \Metatheria \Trichosurus vulpecula \HLtriVul1 GCA_011100635.1_mTriVul1.pri \9337 \\ 454 \North American opossum \Metatheria \Didelphis virginiana \DNA zoo Didelphis virginiana \9267 \\ 455 \ground cuscus \Metatheria \Phalanger gymnotis \DNA zoo Phalanger gymnotis \65615 \\ 456 \gray short-tailed opossum \Metatheria \Monodelphis domestica \Oct. 2006 (Broad/monDom5) \13616 \\ 457 \Leadbeater's possum \Metatheria \Gymnobelideus leadbeateri \HLgymLea1 GCA_011680675.1 \38618 \\ 458 \Tasmanian wolf \Metatheria \Thylacinus cynocephalus \HLthyCyn1 GCA_007646695.1 \9275 \\ 459 \coppery ringtail possum \Metatheria \Pseudochirops cupreus \DNA zoo Pseudochirops cupreus \37702 \\ 460 \eastern gray kangaroo \Metatheria \Macropus giganteus \DNA zoo Macropus giganteus \9317 \\ 461 \golden ringtail possum \Metatheria \Pseudochirops corinnae \DNA zoo Pseudochirops corinnae \65629 \\ 462 \western gray kangaroo \Metatheria \Macropus fuliginosus \DNA zoo Macropus fuliginosus \9316 \\ 463 \tammar wallaby \Metatheria \Macropus eugenii \DNA zoo Macropus eugenii \9315 \\ 464 \red kangaroo \Metatheria \Osphranter rufus \DNA zoo Osphranter rufus \9321 \\ 465 \Western ringtail oppossum \Metatheria \Pseudocheirus occidentalis \DNA zoo Pseudocheirus occidentalis \656515 \\ 466 \tammar wallaby \Metatheria \Macropus eugenii \Sep. 2009 (TWGS Meug_1.1/macEug2) \9315 \\ 467 \yellow-footed antechinus \Metatheria \Antechinus flavipes \HLantFla1 GCA_016432865.1_AdamAnt \38775 \\ 468 \Tasmanian devil \Metatheria \Sarcophilus harrisii \HLsarHar2 GCA_902635505.1 \9305 \\ 469 \platypus \Monotremata \Ornithorhynchus anatinus \HLornAna3 GCA_004115215.1 \9258 \\ 470 \Australian echidna \Monotremata \Tachyglossus aculeatus \HLtacAcu1 GCA_015852505.1 \9261 \
\ Table 1. Genome assemblies included in the 470-way Conservation track.\
\ Downloads for data in this track are available:\
\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ size of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ The names of selected species are colored according to their clade,\ alternating between blue and green.\ Note that excluding species from the pairwise display does not alter the\ the conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. Missing sequence in any\ assembly is highlighted in the track display by regions of yellow when zoomed\ out and by Ns when displayed at base level. The following conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment. The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\\ Codon translation is available in base-level display mode if the\ displayed region is identified as a coding segment. To display this annotation,\ select the species for translation from the pull-down menu in the Codon\ Translation configuration section at the top of the page. Then, select one of\ the following modes:\
\ Codon translation uses the following gene tracks as the basis for translation:\
\ \\
\ Table 2. Gene tracks used for codon translation.\\ Gene Track Species \ RefSeq Genes aardvark, American pika, Amur tiger, Angolan colobus, big brown bat, black flying fox, black snub-nosed monkey, Bolivian squirrel monkey, Brandt's bat, Cape elephant shrew, Cape golden mole, cattle, chimpanzee, Chinese tree shrew, Coquerel's sifaka, degu, dog, domestic cat, domestic guinea pig, drill, European shrew, Florida manatee, golden hamster, gray mouse lemur, green monkey, Hawaiian monk seal, horse, house mouse, house mouse, human, killer whale, lesser Egyptian jerboa, little brown bat, long-tailed chinchilla, Ma's night monkey, minke whale, naked mole-rat, nine-banded armadillo, Northern sea otter, Norway rat, Ord's kangaroo rat, Pacific walrus, Panamanian white-faced capuchin, Philippine tarsier, pig, pig-tailed macaque, polar bear, prairie vole, Przewalski's horse, pygmy chimpanzee, rabbit, Rhesus monkey, small Madagascar hedgehog, small-eared galago, sooty mangabey, southern white rhinoceros, star-nosed mole, Sumatran orangutan, thirteen-lined ground squirrel, Upper Galilee mountains blind mole rat, Vespertilio Davidii, Weddell seal, western European hedgehog, western lowland gorilla, Yangtze River dolphin \ Ensembl Genes Bos bison bison, Brazilian guinea pig, dog, gray short-tailed opossum, northern tree shrew \ Xeno RefGene alpaca, black lemur, Chinese pangolin, common bottlenose dolphin, proboscis monkey, Sclater's lemur, Southern sea otter, tammar wallaby \ no annotation African buffalo, African grass rat, African hunting dog, African hunting dog, African savanna elephant, African woodland thicket rat, Agile Gracile Mouse Opossum, Allen's swamp monkey, Alpine ibex, Alpine marmot, alpine musk deer, American beaver, American black bear, American black bear, American mink, Amur leopard cat, antarctic fur seal, Antarctic minke whale, Antillean ghost-faced bat, aoudad, Arabian camel, Arctic fox, Arctic ground squirrel, argali, Asian black bear, Asian palm civet, Asiatic elephant, Asiatic mouflon, Asiatic tapir, Asiatic tapir, ass, Australian echidna, aye-aye, babakoto, Bactrian camel, banded mongoose, Bank vole, bearded seal, beluga whale, bighorn sheep, bighorn sheep, black muntjac, black rat, black rhinoceros, black-footed cat, black-handed spider monkey, Blue whale, Bohar reedbuck, Bolivian squirrel monkey, Bolivian titi, Bonin flying fox, boutu, bowhead whale, Brazilian free-tailed bat, Brazilian porcupine, Brazilian tapir, brindled gnu, brown lemur, brush rabbit, bush duiker, bushbuck, Cacomistle, cactus mouse, California big-eared bat, California sea lion, Canada lynx, Cantor's roundleaf bat, Cape rock hyrax, capybara, Central European red deer, Chacoan peccary, cheetah, Chinese forest musk deer, Chinese hamster, Chinese pangolin, Chinese rufous horseshoe bat, Chinese water deer, chiru, Clouded leopard, Cobus hunteri, common bottlenose dolphin, common bottlenose dolphin, common brushtail, common pipistrelle, common pipistrelle, common vampire bat, Common vole, common wombat, coppery ringtail possum, Coquerel's mouse lemur, crab-eating macaque, crested porcupine, Cuvier's beaked whale, Damara mole-rat, dassie-rat, Daurian ground squirrel, De Brazza's monkey, desert woodrat, dingo, domestic ferret, domestic yak, donkey, dugong, dwarf mongoose, eastern gray kangaroo, eastern mole, Eastern roe deer, Egyptian rousette, Egyptian spiny mouse, Equus burchelli boehmi, ermine, Eurasian elk, Eurasian red squirrel, Eurasian river otter, Eurasian water vole, European polecat, European rabbit, European woodmouse, evening bat, Fat dormouse, fat sand rat, Fin whale, fossa, franciscana, Francois's langur, Gambian giant pouched rat, gaur, gayal, gelada, gemsbok, gerenuk, giant anteater, giant otter, giant otter, giant panda, giraffe, giraffe, goat, Gobi jerboa, golden ringtail possum, golden snub-nosed monkey, golden spiny mouse, gracile shrew mole, Grant's gazelle, gray seal, gray squirrel, great gerbil, great roundleaf bat, greater bamboo lemur, greater bulldog bat, Greater cane rat, greater horseshoe bat, greater Indian rhinoceros, greater kudu, greater mouse-eared bat, grey whale, grizzly bear, ground cuscus, guanaco, Gunnison's prairie dog, Hanuman langur, harbor porpoise, harbor porpoise, harbor seal, Harvey's duiker, hazel dormouse, Hesperomys crinitus, Himalayan marmot, hippopotamus, hippopotamus, Hispaniolan solenodon, hispid cotton rat, hoary bamboo rat, hoary bat, Hoffmann's two-fingered sloth, Hog deer, hog-nosed bat, Honduran yellow-shouldered bat, humpback whale, Iberian mole, impala, Indian false vampire, Indian flying fox, Indo-pacific bottlenose dolphin, Indo-pacific bottlenose dolphin, Indo-pacific humpbacked dolphin, Indus River dolphin, jaguar, jaguar, jaguarundi, Jamaican fruit-eating bat, Jamaican fruit-eating bat, Japanese macaque, Java mouse-deer, kinkajou, Kirk's dik-dik, klipspringer, koala, Kuhl's pipistrelle, Lama pacos huacaya, large flying fox, Leadbeater's possum, lechwe, leopard, Leschenault's rousette, lesser dawn bat, Lesser dwarf lemur, lesser kudu, Lesser long-nosed bat, lesser mouse-deer, lesser panda, lesser short-nosed fruit bat, lion, little brown bat, llama, llama, long-finned pilot whale, long-tongued fruit bat, Madagascan rousette, Malagasy flying fox, Malagasy straw-colored fruit bat, Malayan pangolin, Malayan pangolin, mandrill, mantled howler monkey, Masai giraffe, Maxwell's duiker, meadow jumping mouse, meerkat, meerkat, melon-headed whale, Miniopterus schreibersii natalensis, Mona monkey, Mongolian gerbil, mongoose lemur, Montane guinea pig, mountain beaver, mountain goat, Mountain hare, mouse lemur, mule deer, muntjak, Murina feae, muskrat, narwhal, Nilgiri tahr, North American badger, North American opossum, North American porcupine, North Atlantic right whale, North Pacific right whale, Northern American river otter, Northern elephant seal, northern fur seal, Northern giant mouse lemur, northern gundi, Northern long-eared myotis, Northern mole vole, northern rock mouse, Northern rufous mouse lemur, northern white rhinoceros, northern white-cheeked gibbon, Norway rat, nutria, okapi, oldfield mouse, olive baboon, pacarana, Pacific pocket mouse, Pacific white-sided dolphin, pale spear-nosed bat, Pallas's mastiff bat, pallid bat, Parnell's mustached bat, Patagonian cavy, Pere David's deer, Peromyscus californicus subsp. insignis, platypus, porcupine caribou, prairie deer mouse, pronghorn, Przewalski's gazelle, puma, punctate agouti, pygmy Bryde's whale, pygmy marmoset, pygmy sperm whale, rabbit, raccoon, ratel, red bat, red fox, red guenon, red kangaroo, Red shanked douc langur, reed vole, Reeves' muntjac, reindeer, Ring-tailed lemur, roan antelope, root vole, royal antelope, Ryukyu mouse, sable, sable antelope, saiga antelope, Schizostoma hirsutum, Schreibers' long-fingered bat, scimitar-horned oryx, Sclater's lemur, Seba's short-tailed bat, sheep, short-tailed field vole, shrew mouse, Siberian ibex, Siberian musk deer, silvery gibbon, slow loris, snow sheep, snowshoe hare, social tuco-tuco, South African ground squirrel, Southern elephant seal, southern grasshopper mouse, southern multimammate mouse, southern tamandua, Southern three-banded armadillo, southern two-toed sloth, southern two-toed sloth, Sowerby's beaked whale, Spanish lynx, sperm whale, sperm whale, spotted hyena, springbok, springhare, steenbok, Steller sea lion, Steller's sea cow, Stephens's kangaroo rat, steppe mouse, straw-colored fruit bat, stripe-headed round-eared bat, striped hyena, Sumatran rhinoceros, Sunda flying lemur, suni, tailed tailless bat, Talazac's shrew tenrec, tamarin, tammar wallaby, Tasmanian devil, Tasmanian wolf, Thomson's gazelle, topi, Transcaucasian mole vole, Tree pangolin, Tree pangolin, tufted capuchin, Ugandan red Colobus, Vancouver Island marmot, vaquita, Vicugna mensalis, walrus, water buffalo, waterbuck, western gray kangaroo, Western ringtail oppossum, western spotted skunk, western wild mouse, white-faced saki, white-footed mouse, white-fronted capuchin, white-lipped deer, White-nosed coati, white-tailed deer, white-tailed deer, white-tailed deer, white-tufted-ear marmoset, Wild Bactrian camel, wild goat, wild yak, wolverine, woodchuck, woodchuck, woodland dormouse, Yangtze finless porpoise, Yarkand deer, yellow-bellied marmot, yellow-footed antechinus, yellow-spotted hyrax, zebu cattle,\
\ Pairwise alignments with the human genome were generated for\ each species using lastz from repeat-masked genomic sequence.\ Pairwise alignments were then linked into chains using a dynamic programming\ algorithm that finds maximally scoring chains of gapless subsections\ of the alignments organized in a kd-tree.\ The scoring matrix and parameters for pairwise alignment and chaining\ were tuned for each species based on phylogenetic distance from the reference.\ High-scoring chains were then placed along the genome, with\ gaps filled by lower-scoring chains, to produce an alignment net.\
\ \\ The phyloP are phylogenetic methods that rely\ on a tree model containing the tree topology, branch lengths representing\ evolutionary distance at neutrally evolving sites, the background distribution\ of nucleotides, and a substitution rate matrix.\ The\ all-species tree model for this track was\ generated using the phyloFit program from the PHAST package\ (REV model, EM algorithm, medium precision) using multiple alignments of\ 4-fold degenerate sites extracted from the 470-way alignment\ (msa_view). The 4d sites were derived from the RefSeq (Reviewed+Coding) gene\ set, filtered to select single-coverage long transcripts.\
\\ This same tree model was used in the phyloP calculations; however, the\ background frequencies were modified to maintain reversibility.\ The resulting tree model:\ all species.\
\\ The phyloP program supports several different methods for computing\ p-values of conservation or acceleration, for individual nucleotides or\ larger elements (\ http://compgen.cshl.edu/phast/). Here it was used\ to produce separate scores at each base (--wig-scores option), considering\ all branches of the phylogeny rather than a particular subtree or lineage\ (i.e., the --subtree option was not used). The scores were computed by\ performing a likelihood ratio test at each alignment column (--method LRT),\ and scores for both conservation and acceleration were produced (--mode\ CONACC).\
\ \This track was created using the following programs:\
\ Harris RS.\ Improved pairwise alignment of genomic DNA.\ Ph.D. Thesis. Pennsylvania State University, USA. 2007.\
\ \\ Cooper GM, Stone EA, Asimenos G, NISC Comparative Sequencing Program., Green ED, Batzoglou S, Sidow\ A.\ \ Distribution and intensity of constraint in mammalian genomic sequence.\ Genome Res. 2005 Jul;15(7):901-13.\ PMID: 15965027;\ PMC: PMC1172034;\ DOI: 10.1101/gr.3577405\
\ \\ Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A.\ \ Detection of nonneutral substitution rates on mammalian phylogenies.\ Genome Res. 2010 Jan;20(1):110-21.\ PMID: 19858363;\ PMC: PMC2798823\
\ \\ Siepel A, Haussler D.\ Phylogenetic Hidden Markov Models.\ In: Nielsen R, editor. Statistical Methods in Molecular Evolution.\ New York: Springer; 2005. pp. 325-351.\ DOI: 10.1007/0-387-27733-1_12\
\ \\ Siepel A, Pollard KS, and Haussler D. New methods for detecting\ lineage-specific selection. In Proceedings of the 10th International\ Conference on Research in Computational Molecular Biology (RECOMB 2006), pp. 190-205.\ DOI: 10.1007/11732990_17\
\ compGeno 1 compositeTrack on\ dragAndDrop subTracks\ group compGeno\ longLabel Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group,\ shortLabel Hiller Lab 470 Mammals\ subGroup1 view Views align=Multiz_Alignments phyloP=Basewise_Conservation_(phyloP) phastcons=Element_Conservation_(phastCons) elements=Conserved_Elements\ track cons470way\ type bed 4\ visibility hide\ hprcDecomposed HPRC All Variants vcfTabix HPRC variants decomposed from hprc-v1.0-mc.grch38.vcfbub.a100k.wave.vcf.gz (Liao et al 2023), no size filtering 0 100 0 0 0 127 127 127 0 0 0\ This track shows short nucleotide variants of a few base pairs when aligning\ HPRC genomes to the hg38 reference assembly. The alignment was made with the\ Minigraph-cactus approach described in the references below.\
\ \There are three subtracks in this superTrack:\
\ VCF Decomposition from\ HPRC Pangenome Resources Github:\ "The Raw VCF files contain a site for each bubble in the graph. Nested bubbles will result in\ overlapping sites. The nesting relationships are denoted with the PS (parent snarl), LV (level) and\ AT (allele traversal) tags and need to be taken into account when interpreting the VCF.\ Alternatively, you can use the 'Decomposed VCFs' which have been normalized by using\ vcfbub to 'pop'\ bubbles with alleles larger than 100k and\ vcfwave\ to realign each alt\ (script). Note that in order to reproduce the PanGenie analyses from the papers, you should instead\ use the\ PanGenie HPRC Workflow. This workflow has a\ CHM13 branch to use when working with that reference.\
\ The exact tools and commands used to produce the VCFs are given\ here."
\ \\ The Name of the items are the pair of node labels that denote the site's location\ in the graph, with the '>' and '<' denoting the forward and reverse\ orientation of the node. Mouseover on items in "squish" and "pack" modes shows the items Name and\ Genotypes. Mouseover on items in "full" mode shows Alleles.\ \
\ The Minigraph-Cactus HPRC v1.0 graph was converted to VCF using vg deconstruct.\ This result was further postprocessed using vcfbub to flatten nested sites then\ vcfwave to normalize by realigning alt alleles to the reference. All steps are\ described in Hickey et al 2023. The postprocessing command lines and data can be found on\ Github.\ Finally, the resulting VCF was filtered by length and split into two VCFs using a cutoff of 3bp.\
\ \\ Thanks to Glenn Hickey for providing the HAL file from the HPRC project and for making these VCFs from them.\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q,\ Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ PMID: 33177663;\ PMC: PMC7673649;\ DOI: 10.1038/s41586-020-2871-y\
\ \\ Glenn Hickey, Jean Monlong, Jana Ebler, Adam M Novak, Jordan M Eizenga,\ Yan Gao; Human Pangenome Reference Consortium; Tobias Marschall, Heng Li,\ Benedict Paten\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nature Biotechnology. 2023 May 10. doi: 10.1038/s41587-023-01793-w.\ PMID: 37165083;\ DOI: 10.1038/s41587-023-01793-w\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ PMID: 21665927;\ PMC: PMC3166836;\ DOI: 10.1101/gr.123356.111\
\ \\ Wen-Wei Liao, Mobin Asri, Jana Ebler, ...et al, Heng Lin,\ Benedict Paten\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ PMID: 37165242;\ PMC: PMC1017212;\ DOI: 10.1038/s41586-023-05896-x\
\ hprc 1 bigDataUrl /gbdb/hg38/hprc/decomposed.vcf.gz\ configureByPopup off\ dataVersion August 2023\ html hprcVCF\ longLabel HPRC variants decomposed from hprc-v1.0-mc.grch38.vcfbub.a100k.wave.vcf.gz (Liao et al 2023), no size filtering\ maxWindowToDraw 200000\ parent hprcVCF\ shortLabel HPRC All Variants\ showHardyWeinberg on\ track hprcDecomposed\ type vcfTabix\ visibility hide\ hprc2v21Sv HPRC v2.1 233 SVs bigBed 9 + Structural Variants from HPRC v2.1 Pangenome Graph (233 samples, minigraph-cactus) 0 100 0 0 0 127 127 127 0 0 0\ A pangenome graph holds many human genomes at once. Sequence that the\ genomes share collapses onto common paths, and the places where they\ differ show up as bubbles in the graph. This track shows the structural\ variants found in version 2.1 of the Human Pangenome Reference Consortium\ (HPRC) minigraph-cactus graph, which was built from haplotype-resolved\ PacBio HiFi assemblies of 233 samples. Only larger events are shown here:\ insertions and deletions of at least 50 bp. HPRC produces one variant file\ per reference path, so the events are measured against GRCh38 on hg38 and\ against T2T-CHM13 on hs1, and each assembly shows its own native callset.\
\\ On hg38 there are about 550,000 such alleles (roughly 422,000 insertions and\ 128,000 deletions). On hs1 there are about 541,000 (roughly 348,000\ insertions and 193,000 deletions). The two sets are not lifted between\ assemblies; the counts differ because an insertion against one reference can\ be a deletion against the other.\
\ \\ Items are colored by SV type:\
\| \ | Insertion (INS) |
|---|---|
| \ | Deletion (DEL) |
\ An insertion is drawn as a 1 bp anchor at the point where the extra\ sequence goes in. A deletion spans the stretch of reference that is\ missing. Each variant keeps its allele count, allele frequency, the\ number of samples with data, and the level it sits at in the graph's\ snarl tree. A snarl level of 0 is a top-level bubble; higher numbers are\ bubbles nested inside a parent bubble. All of these can be used as\ filters.\
\ \\ HPRC release 2 does not yet have a peer-reviewed paper. The graph was\ built with minigraph-cactus from haplotype-resolved PacBio HiFi assemblies\ of 233 samples, including T2T-CHM13 and the diverse 1000 Genomes Project\ panel, using GRCh38 as the reference path. Variants were called from the\ graph with vg deconstruct. HPRC keeps the sample list and assembly\ provenance in\ \ alignments_v2.0.csv.\
\\ We started from the per-reference files provided by the HPRC graph team,\ hprc-v2.1-mc-grch38.gref95.ro.vcf.gz for hg38 and\ hprc-v2.1-mc-chm13.gref95.ro.vcf.gz for hs1. These are the raw\ vg deconstruct output: each graph bubble is one multi-allelic\ record with its graph traversals attached, and there are no per-allele type\ or length fields. To turn a file into a track, we compared every alternate\ allele to the reference allele after trimming the sequence they share at\ each end. An allele was kept when the net length change was at least 50 bp,\ and labeled an insertion when the alternate is longer or a deletion when it\ is shorter. At this size no balanced, equal-length substitutions came up,\ and the files carry no inversion calls, so the track has only insertions and\ deletions. On hg38, 549,649 alleles were kept (40,678 at nested snarl\ levels); on hs1, 541,176 (70,200 nested), after removing byte-identical\ duplicate records. Because these files are not broken\ down into atomic indels, one bubble can appear as a single large allele\ rather than several small ones, so the counts are not comparable to a\ wave-decomposed callset. Allele counts, frequencies and sample counts come\ straight from the VCF.\
\\ The conversion script and autoSql schema are in\ \ makeDb/scripts/lrSv and the build steps are in the makeDoc at\ \ doc/hg38/lrSv.txt, and the track configuration is in\ trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator, and read programmatically\ through our API,\ track=hprc2v21Sv. For automated download and analysis the variants\ are in a bigBed file on our download server, one per assembly:\ \ hg38 and\ \ hs1. You can pull out one region or the whole set with\ bigBedToBed, for example\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/hprc2v21.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\ \\ Thanks to the Human Pangenome Reference Consortium for building and\ releasing the release-2 minigraph-cactus pangenome, and to Glenn Hickey\ for the v2.1 deconstructed VCF.\
\ \\ HPRC release 2 is not yet described in a peer-reviewed publication. The\ release announcement has background and data-access details:\ \ HPRC data release 2.\
\ varRep 1 bigDataUrl /gbdb/hg38/lrSv/hprc2v21.bb\ filter.AC 0:463\ filter.alleleFreq 0:1\ filter.insLen 0:1064897\ filter.snarlLevel 0:7\ filter.svLen 0:99835\ filterByRange.AC on\ filterByRange.alleleFreq on\ filterByRange.insLen on\ filterByRange.snarlLevel on\ filterByRange.svLen on\ filterLabel.AC Allele Count\ filterLabel.alleleFreq Allele Frequency\ filterLabel.insLen Insertion Length\ filterLabel.snarlLevel Snarl Level\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.alleleFreq 0:1\ filterType.svType multipleListOr\ filterValues.svType INS,DEL\ itemRgb on\ longLabel Structural Variants from HPRC v2.1 Pangenome Graph (233 samples, minigraph-cactus)\ mouseOver Var: $name ($svType)\ This track shows short nucleotide variants of a few base pairs when aligning\ HPRC genomes to the hg38 reference assembly. The alignment was made with the\ Minigraph-cactus approach described in the references below.\
\ \There are three subtracks in this superTrack:\
\ VCF Decomposition from\ HPRC Pangenome Resources Github:\ "The Raw VCF files contain a site for each bubble in the graph. Nested bubbles will result in\ overlapping sites. The nesting relationships are denoted with the PS (parent snarl), LV (level) and\ AT (allele traversal) tags and need to be taken into account when interpreting the VCF.\ Alternatively, you can use the 'Decomposed VCFs' which have been normalized by using\ vcfbub to 'pop'\ bubbles with alleles larger than 100k and\ vcfwave\ to realign each alt\ (script). Note that in order to reproduce the PanGenie analyses from the papers, you should instead\ use the\ PanGenie HPRC Workflow. This workflow has a\ CHM13 branch to use when working with that reference.\
\ The exact tools and commands used to produce the VCFs are given\ here."
\ \\ The Name of the items are the pair of node labels that denote the site's location\ in the graph, with the '>' and '<' denoting the forward and reverse\ orientation of the node. Mouseover on items in "squish" and "pack" modes shows the items Name and\ Genotypes. Mouseover on items in "full" mode shows Alleles.\ \
\ The Minigraph-Cactus HPRC v1.0 graph was converted to VCF using vg deconstruct.\ This result was further postprocessed using vcfbub to flatten nested sites then\ vcfwave to normalize by realigning alt alleles to the reference. All steps are\ described in Hickey et al 2023. The postprocessing command lines and data can be found on\ Github.\ Finally, the resulting VCF was filtered by length and split into two VCFs using a cutoff of 3bp.\
\ \\ Thanks to Glenn Hickey for providing the HAL file from the HPRC project and for making these VCFs from them.\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q,\ Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ PMID: 33177663;\ PMC: PMC7673649;\ DOI: 10.1038/s41586-020-2871-y\
\ \\ Glenn Hickey, Jean Monlong, Jana Ebler, Adam M Novak, Jordan M Eizenga,\ Yan Gao; Human Pangenome Reference Consortium; Tobias Marschall, Heng Li,\ Benedict Paten\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nature Biotechnology. 2023 May 10. doi: 10.1038/s41587-023-01793-w.\ PMID: 37165083;\ DOI: 10.1038/s41587-023-01793-w\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ PMID: 21665927;\ PMC: PMC3166836;\ DOI: 10.1101/gr.123356.111\
\ \\ Wen-Wei Liao, Mobin Asri, Jana Ebler, ...et al, Heng Lin,\ Benedict Paten\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ PMID: 37165242;\ PMC: PMC1017212;\ DOI: 10.1038/s41586-023-05896-x\
\ hprc 1 bigDataUrl /gbdb/hg38/hprc/decomposedUnder4.vcf.gz\ configureByPopup off\ dataVersion August 2023\ html hprcVCF\ longLabel HPRC VCF variants filtered for items size <= 3bp\ maxWindowToDraw 200000\ parent hprcVCF\ shortLabel HPRC Variants <= 3bp\ showHardyWeinberg on\ track hprcVCFDecomposedUnder4\ type vcfTabix\ visibility pack\ hprcVCFDecomposedOver3 HPRC Variants > 3bp vcfTabix HPRC VCF variants filtered for items size > 3bp 0 100 0 0 0 127 127 127 0 0 0\ This track shows short nucleotide variants of a few base pairs when aligning\ HPRC genomes to the hg38 reference assembly. The alignment was made with the\ Minigraph-cactus approach described in the references below.\
\ \There are three subtracks in this superTrack:\
\ VCF Decomposition from\ HPRC Pangenome Resources Github:\ "The Raw VCF files contain a site for each bubble in the graph. Nested bubbles will result in\ overlapping sites. The nesting relationships are denoted with the PS (parent snarl), LV (level) and\ AT (allele traversal) tags and need to be taken into account when interpreting the VCF.\ Alternatively, you can use the 'Decomposed VCFs' which have been normalized by using\ vcfbub to 'pop'\ bubbles with alleles larger than 100k and\ vcfwave\ to realign each alt\ (script). Note that in order to reproduce the PanGenie analyses from the papers, you should instead\ use the\ PanGenie HPRC Workflow. This workflow has a\ CHM13 branch to use when working with that reference.\
\ The exact tools and commands used to produce the VCFs are given\ here."
\ \\ The Name of the items are the pair of node labels that denote the site's location\ in the graph, with the '>' and '<' denoting the forward and reverse\ orientation of the node. Mouseover on items in "squish" and "pack" modes shows the items Name and\ Genotypes. Mouseover on items in "full" mode shows Alleles.\ \
\ The Minigraph-Cactus HPRC v1.0 graph was converted to VCF using vg deconstruct.\ This result was further postprocessed using vcfbub to flatten nested sites then\ vcfwave to normalize by realigning alt alleles to the reference. All steps are\ described in Hickey et al 2023. The postprocessing command lines and data can be found on\ Github.\ Finally, the resulting VCF was filtered by length and split into two VCFs using a cutoff of 3bp.\
\ \\ Thanks to Glenn Hickey for providing the HAL file from the HPRC project and for making these VCFs from them.\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q,\ Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ PMID: 33177663;\ PMC: PMC7673649;\ DOI: 10.1038/s41586-020-2871-y\
\ \\ Glenn Hickey, Jean Monlong, Jana Ebler, Adam M Novak, Jordan M Eizenga,\ Yan Gao; Human Pangenome Reference Consortium; Tobias Marschall, Heng Li,\ Benedict Paten\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nature Biotechnology. 2023 May 10. doi: 10.1038/s41587-023-01793-w.\ PMID: 37165083;\ DOI: 10.1038/s41587-023-01793-w\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ PMID: 21665927;\ PMC: PMC3166836;\ DOI: 10.1101/gr.123356.111\
\ \\ Wen-Wei Liao, Mobin Asri, Jana Ebler, ...et al, Heng Lin,\ Benedict Paten\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ PMID: 37165242;\ PMC: PMC1017212;\ DOI: 10.1038/s41586-023-05896-x\
\ hprc 1 bigDataUrl /gbdb/hg38/hprc/decomposedOver3.vcf.gz\ configureByPopup off\ dataVersion August 2023\ html hprcVCF\ longLabel HPRC VCF variants filtered for items size > 3bp\ maxWindowToDraw 200000\ parent hprcVCF\ shortLabel HPRC Variants > 3bp\ showHardyWeinberg on\ track hprcVCFDecomposedOver3\ type vcfTabix\ visibility hide\ hgIkmc IKMC Genes Mapped bed 12 International Knockout Mouse Consortium Genes Mapped to Human Genome 0 100 0 0 0 127 127 127 0 0 0 http://www.mousephenotype.org/data/genes/$$\ This track shows genes targeted by \ International Knockout Mouse Consortium (IKMC)\ mapped to the human genome. IKMC is a \ collaboration to generate a public resource of mouse embryonic stem (ES)\ cells containing a null mutation in every gene in the mouse genome.\ Gene targets are color-coded by status:\
\ The KnockOut Mouse Project Data\ Coordination Center (KOMP DCC) is the central database resource\ for coordinating mouse gene targeting within IKMC and provides\ web-based query and display tools for IKMC data. In addition, the\ KOMP DCC website provides a tool for the scientific community to\ nominate genes of interest to be knocked out by the KOMP initiative.
\ \\ IKMC members include\
\ Using complementary targeting strategies, the IKMC centers\ design and create targeting vectors, mutant ES cell lines and, to some\ extent, mutant mice, embryos or sperm. Materials are distributed to\ the research community.
\\ The KOMP Repository\ archives, maintains, and distributes IKMC products. Researchers can\ order products and get product information from the\ Repository. Researchers can also express interest in products that are\ still in the pipeline. They will then receive email notification as\ soon as KOMP generated products are available for distribution.
\\ The process for ordering EUCOMM materials can be found \ here.
\\ The process for ordering TIGM materials can be found \ here.
\\ Information on NorCOMM products and services can be found \ here.\
\ Genes were mapped to the human genome by IKMC.\
\ \\ Thanks to the International Knockout Mouse Consortium, and Carol Bult in \ particular, for providing these data.
\ \\ Austin CP, Battey JF, Bradley A, Bucan M, Capecchi M, Collins FS, Dove WF, Duyk G, Dymecki S, Eppig\ JT et al.\ \ The knockout mouse project.\ Nat Genet. 2004 Sep;36(9):921-4.\ PMID: 15340423; PMC: PMC2716027\
\ \\ Collins FS, Finnell RH, Rossant J, Wurst W.\ \ A new partner for the international knockout mouse consortium.\ Cell. 2007 Apr 20;129(2):235.\ PMID: 17448981\
\ \\ International Mouse Knockout Consortium, Collins FS, Rossant J, Wurst W.\ \ A mouse for all reasons.\ Cell. 2007 Jan 12;128(1):9-13.\ PMID: 17218247\
\ genes 1 exonNumbers off\ group genes\ itemRgb on\ longLabel International Knockout Mouse Consortium Genes Mapped to Human Genome\ mgiUrl https://www.informatics.jax.org//marker/$$\ mgiUrlLabel MGI Report:\ noScoreFilter .\ origAssembly hg19\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel IKMC Genes Mapped\ track hgIkmc\ type bed 12\ url http://www.mousephenotype.org/data/genes/$$\ urlLabel KOMP Data Coordination Center:\ visibility hide\ ileumWangCellType Ileum Cells bigBarChart Ileum cells binned by cell type from Wang et al 2020 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-intestine+ileum&gene=$$\ This track shows data from \ Single-cell transcriptome analysis reveals differential nutrient absorption\ functions in human intestine. Droplet-based single-cell RNA sequencing\ (scRNA-seq) was used to survey gene expression profiles of the epithelium in\ the human ileum, colon, and rectum. A total of 7 cell clusters were identified:\ enterocytes (EC), goblet cells (G), paneth-like cells (PLC), enteroendocrine\ cells (EEC), progenitor cells (PRO), transient-amplifying cells (TA) and stem\ cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in ileum\ cells where cells are grouped by cell type\ (Ileum Cells) or donor\ (Ileum Donor). The default track\ displayed is Ileum Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. Note that the Ileum Donor track \ is colored by donor for improved clarity.
\ \\ Using single-cell RNA sequencing, RNA profiles of intestinal epithelial cells\ were obtained for 6,167 cells from two human ileum samples. Tissue samples\ belonged to a male donor age 60 with Neuroendocrine Carcinoma (Ileum-1) and a\ female donor age 67 with Adenocarcinoma (Ileum-2). The healthy intestinal\ mucous membranes used for each sample were cut away from the tumor border in\ surgically removed ileum tissue. Additionally, the intestinal tissues were\ washed in Hank's balanced salt solution (HBSS) to remove mucus, blood cells,\ and muscle tissue. The sample was enriched for epithelial cells through \ centrifugation before being dissociated with Tryple to obtain single-cell \ suspensions. RNA-seq libraries were prepared using 10x Genomics 3' v2 kit and \ sequenced on an Illumina Hiseq X Ten PE150.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The UCSC command line utility\ matrixClusterColumns, matrixToBarChart, and bedToBigBed were used to transform\ these into a bar chart format bigBed file that can be visualized. The coloring\ was done by defining colors for the broad level cell classes and then using\ another UCSC utility, hcaColorCells, to interpolate the colors across all cell\ types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The\ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720
\ \ singleCell 1 barChartBars enteroendocrine_cell enterocyte goblet_cell paneth-like_cell progenitor_cell stem_cell transit-amplifying_cell\ barChartColors #bcd0f3 #0198c0 #568bfd #629be4 #436ca1 #9ea0a1 #919eb1\ barChartLimit 1.6\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/ileumWang/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/ileumWang/cell_type.bb\ defaultLabelFields name\ html ileumWang\ labelFields name,name2\ longLabel Ileum cells binned by cell type from Wang et al 2020\ parent ileumWang\ shortLabel Ileum Cells\ track ileumWangCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-intestine+ileum&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ ileumWangDonor Ileum Donor bigBarChart Ileum cells binned by organ donor from Wang et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-intestine+ileum&gene=$$\ This track shows data from \ Single-cell transcriptome analysis reveals differential nutrient absorption\ functions in human intestine. Droplet-based single-cell RNA sequencing\ (scRNA-seq) was used to survey gene expression profiles of the epithelium in\ the human ileum, colon, and rectum. A total of 7 cell clusters were identified:\ enterocytes (EC), goblet cells (G), paneth-like cells (PLC), enteroendocrine\ cells (EEC), progenitor cells (PRO), transient-amplifying cells (TA) and stem\ cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in ileum\ cells where cells are grouped by cell type\ (Ileum Cells) or donor\ (Ileum Donor). The default track\ displayed is Ileum Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. Note that the Ileum Donor track \ is colored by donor for improved clarity.
\ \\ Using single-cell RNA sequencing, RNA profiles of intestinal epithelial cells\ were obtained for 6,167 cells from two human ileum samples. Tissue samples\ belonged to a male donor age 60 with Neuroendocrine Carcinoma (Ileum-1) and a\ female donor age 67 with Adenocarcinoma (Ileum-2). The healthy intestinal\ mucous membranes used for each sample were cut away from the tumor border in\ surgically removed ileum tissue. Additionally, the intestinal tissues were\ washed in Hank's balanced salt solution (HBSS) to remove mucus, blood cells,\ and muscle tissue. The sample was enriched for epithelial cells through \ centrifugation before being dissociated with Tryple to obtain single-cell \ suspensions. RNA-seq libraries were prepared using 10x Genomics 3' v2 kit and \ sequenced on an Illumina Hiseq X Ten PE150.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The UCSC command line utility\ matrixClusterColumns, matrixToBarChart, and bedToBigBed were used to transform\ these into a bar chart format bigBed file that can be visualized. The coloring\ was done by defining colors for the broad level cell classes and then using\ another UCSC utility, hcaColorCells, to interpolate the colors across all cell\ types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The\ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720
\ \ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/ileumWang/donor.colors\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/ileumWang/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/ileumWang/donor.bb\ defaultLabelFields name\ html ileumWang\ labelFields name,name2\ longLabel Ileum cells binned by organ donor from Wang et al 2020\ parent ileumWang\ shortLabel Ileum Donor\ track ileumWangDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-intestine+ileum&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ ileumWang Ileum Wang Ileum single cell sequencing from Wang et al 2020 0 100 0 0 0 127 127 127 0 0 0\ This track shows data from \ Single-cell transcriptome analysis reveals differential nutrient absorption\ functions in human intestine. Droplet-based single-cell RNA sequencing\ (scRNA-seq) was used to survey gene expression profiles of the epithelium in\ the human ileum, colon, and rectum. A total of 7 cell clusters were identified:\ enterocytes (EC), goblet cells (G), paneth-like cells (PLC), enteroendocrine\ cells (EEC), progenitor cells (PRO), transient-amplifying cells (TA) and stem\ cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in ileum\ cells where cells are grouped by cell type\ (Ileum Cells) or donor\ (Ileum Donor). The default track\ displayed is Ileum Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. Note that the Ileum Donor track \ is colored by donor for improved clarity.
\ \\ Using single-cell RNA sequencing, RNA profiles of intestinal epithelial cells\ were obtained for 6,167 cells from two human ileum samples. Tissue samples\ belonged to a male donor age 60 with Neuroendocrine Carcinoma (Ileum-1) and a\ female donor age 67 with Adenocarcinoma (Ileum-2). The healthy intestinal\ mucous membranes used for each sample were cut away from the tumor border in\ surgically removed ileum tissue. Additionally, the intestinal tissues were\ washed in Hank's balanced salt solution (HBSS) to remove mucus, blood cells,\ and muscle tissue. The sample was enriched for epithelial cells through \ centrifugation before being dissociated with Tryple to obtain single-cell \ suspensions. RNA-seq libraries were prepared using 10x Genomics 3' v2 kit and \ sequenced on an Illumina Hiseq X Ten PE150.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. The UCSC command line utility\ matrixClusterColumns, matrixToBarChart, and bedToBigBed were used to transform\ these into a bar chart format bigBed file that can be visualized. The coloring\ was done by defining colors for the broad level cell classes and then using\ another UCSC utility, hcaColorCells, to interpolate the colors across all cell\ types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The\ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720
\ \ singleCell 0 group singleCell\ longLabel Ileum single cell sequencing from Wang et al 2020\ shortLabel Ileum Wang\ superTrack on\ track ileumWang\ visibility hide\ ucscToINSDC INSDC bed 4 Accession at INSDC - International Nucleotide Sequence Database Collaboration 0 100 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/nuccore/$$\ This track associates UCSC Genome Browser chromosome names to accession\ names from the International Nucleotide Sequence Database Collaboration (INSDC).\
\ \\ The data were downloaded from the NCBI assembly database.\
\ \ The data for this track was prepared by\
Hiram Clawson.\
\
map 1 group map\
longLabel Accession at INSDC - International Nucleotide Sequence Database Collaboration\
shortLabel INSDC\
track ucscToINSDC\
type bed 4\
url https://www.ncbi.nlm.nih.gov/nuccore/$$\
urlLabel INSDC link:\
visibility hide\
ghInteraction Interactions bigInteract GeneHancer Regulatory Elements and Gene Interactions 2 100 0 0 0 127 127 127 0 0 0 https://www.genecards.org/cgi-bin/carddisp.pl?gene=$ \
This track represents the genome-wide predicted binding \
sites for TF (transcription factor) binding profiles in the \
JASPAR \
database CORE collection.\
\
Shaded boxes represent predicted binding sites for each of the TF profiles\
in the JASPAR CORE collection. The shading of the boxes indicates \
the p-value of the profile's match to that position (scaled between \
0-1000 scores, where 0 corresponds to a p-value of 1 and 1000 to a \
p-value ≤ 10-10). Thus, the darker the shade, the \
lower (better) the p-value. \
The default view shows only predicted binding sites with scores of 400 or greater but\
can be adjusted in the track settings. Multi-select filters allow viewing of\
particular transcription factors. At window sizes of greater than\
10,000 base pairs, this track turns to density graph mode. \
Zoom to a smaller region and click into an item to see more detail. \
From BED format documentation:\
\
Conversion table: \
For each TF binding profile in the JASPAR database CORE collection, genomes were scanned for matches.\
\
For the computation of relative scores and p-values, we used PWMScan (Ambrosini et al. 2018). \
We selected TFBS predictions with a PWM relative score ≥ 0.8 and a p-value < 0.05.\
P-values were scaled between 0 (corresponding to a p-value of 1) and 1000 (p-value ≤ 10-10)\
for colouring of the genome tracks and to allow for comparison of prediction confidence between \
different profiles.\
\
Please refer to the supplementary information of the JASPAR 2020 manuscript for more details.\
\
The JASPAR 2026 update expanded the JASPAR CORE collection by 12% (306\
added or upgraded profiles), culminating to a set of 2633 non-redundant\
TF binding profiles. Genome sequences were scanned with JASPAR 2026\
CORE TF binding profiles for each taxon independently using PWMScan.\
TFBS predictions were selected with a PWM relative score ≥ 0.8 and a\
p-value < 0.05. P-values were scaled between 0 (corresponding to a\
p-value of 1) and 1000 (p-value ≤ 10-10) for coloring of the genome\
tracks and to allow for comparison of prediction confidence between\
different profiles. More information on the methods can be found in the\
JASPAR 2026 \
publication or on the\
JASPAR website. \
The JASPAR 2024 update expanded the JASPAR CORE collection by 20% (329 added and 72 upgraded\
profiles). The new profiles were introduced after manual curation, in which 26 629 TF binding\
motifs were curated and obtained as PFMs or discovered from ChIP-seq/-exo or DAP-seq data. 2500\
profiles from JASPAR 2022 were revised to either promote them to the CORE collection, update the\
associated metadata, or remove them because of validation inconsistencies or poor quality. The\
JASPAR database stores and focuses mostly on PFMs as the model of choice for TF-DNA interactions.\
More information on the methods can be found in the\
\
JASPAR 2024 publication or on the\
JASPAR website. \
JASPAR 2022 contains updated transcription factor binding sites\
with additional transcription factor profiles. More information on the methods can be found in the\
\
JASPAR 2022 publication\
JASPAR 2022 publication or on the\
JASPAR website. \
JASPAR 2020 scanned DNA sequences with JASPAR CORE TF-binding profiles \
for each taxa independently using PWMScan. TFBS predictions were selected with \
a PWM relative score ≥ 0.8 and a p-value < 0.05. P-values were scaled \
between 0 (corresponding to a p-value of 1) and 1000 (p-value ≤ 10-10) for \
coloring of the genome tracks and to allow for comparison of prediction \
confidence between different profiles. \
JASPAR 2018 used the TFBS Perl module (Lenhard and Wasserman 2002) \
and FIMO (Grant, Bailey, and Noble 2011), as distributed within the MEME suite \
(version 4.11.2) (Bailey et al. 2009). For scanning genomes with the \
BioPerl TFBS module, profiles were converted to PWMs and matches were kept with a \
relative score ≥ 0.8. For the FIMO scan, profiles were reformatted to MEME motifs \
and matches with a p-value < 0.05 were kept. TFBS predictions that were not \
consistent between the two methods (TFBS Perl module and FIMO) were removed. The \
remaining TFBS predictions were colored according \
to their FIMO p-value to allow for comparison of prediction confidence between \
different profiles. \
JASPAR Transcription Factor Binding data includes billions of items.\
Because of the data size, the Table Browser does not allow "Genome" as a query region for this\
track. Limited regions can be explored interactively with the\
Table Browser and cross-referenced with \
Data Integrator, although positional\
queries that are too big can lead to timing out. This results in a black page\
or truncated output. In this case, you may try reducing the chromosomal query to\
a smaller window. \
For programmatic access, \
the track can be accessed using the Genome Browser's \
REST API. \
JASPAR annotations can be downloaded from the\
Genome Browser's download server\
as a bigBed file. This compressed binary format can be remotely queried through\
command line utilities. Please note that some of the download files can be quite large. \
The utilities for working with bigBed-formatted binary files can be downloaded\
here.\
Run a utility with no arguments to see a brief description of the utility and its options.\
Description
\
Display Conventions and Configuration
\
\
\
\
\
\
\
shade \
\
\
\
\
\
\
\
\
\
\
\
\
score in range \
≤ 166 \
167-277 \
278-388 \
389-499 \
500-611 \
612-722 \
723-833 \
834-944 \
≥ 945 \
\
\
\
\
\
Item score \
0 \
100 \
131 \
200 \
300 \
400 \
500 \
600 \
700 \
800 \
900 \
1000 \
\
\
p-value \
1 \
0.1 \
0.049 \
10-2 \
10-3 \
10-4 \
10-5 \
10-6 \
10-7 \
10-8 \
10-9 \
≤ 10-10 \
Methods
\
Brief overview of each release
\
Data Access
\
\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/jaspar/JASPAR2024.bb -chrom=chr1 -start=200000 -end=200400 stdout\ \
\ All data are freely available.\ Additional resources are available directly from the JASPAR group:
\The JASPAR group provides TFBS predictions for many additional species and \ genomes. The 2026 release is available as a native track on the following genomes, and additionally \ on mm10 and araTha1 by connection to their \ \ Public Hub or by clicking the assembly links below:
\| Species | \Genome assembly versions | \
| Human - Homo sapiens | \hg38 | \
| Mouse - Mus musculus | \mm39 | \
| Zebrafish - Danio rerio | \danRer11 | \
| Fruitfly - Drosophila melanogaster | \dm6 | \
| Nematode - Caenorhabditis elegans | \ce11 | \
| Vase tunicate - Ciona intestinalis | \ci3 | \
| Thale cress - Arabidopsis thaliana | \araTha1 | \
| Yeast - Saccharomyces cerevisiae | \sacCer3 | \
| Chicken - Gallus gallus | \galGal6 | \
\ The JASPAR database is a joint effort between several labs (please see the latest JASPAR \ paper, below). Binding site predictions and UCSC tracks were computed by the CBGR team \ at NCMBM using code developed at the Wasserman Lab. For enquiries about the data, \ please contact Anthony Mathelier (\ \ anthony.\ mathelier@ncmbm.\ uio.\ no\ \ ) or Ieva Rauluseviciute (\ \ ieva.\ rauluseviciute@ncmbm.\ uio.\ no\ \ ).\
\ \\\CBGR
\
\ Computational Biology & Gene Regulation
\ Norwegian Centre for Molecular Biosciences and Medicine (NCMBM)
\ University of Oslo
\ Oslo, Norway\
\\ \ \Wasserman Lab
\
\ Centre for Molecular Medicine and Therapeutics
\ BC Children's Hospital Research Institute
\ Department of Medical Genetics
\ University of British Columbia
\ Vancouver, Canada\
\ Ovek Baydar D, Rauluseviciute I, Aronsen DR, Blanc-Mathieu R, Bonthuis I, de Beukelaer H, Ferenc K,\ Jegou A, Kumar V, Lemma RB et al.\ \ JASPAR 2026: expansion of transcription factor binding profiles and integration of deep learning models.\ Nucleic Acids Res. 2026;\ PMID: 41325984; PMC: PMC12807658\
\ \\ Sandelin A, Alkema W, Engstrom P, Wasserman WW, Lenhard B.\ \ JASPAR: an open-access database for eukaryotic transcription factor binding profiles.\ Nucleic Acids Res. 2004;.\ PMID: 14681366\
\ \ regulation 1 compositeTrack on\ exonArrows on\ filter.score 400\ filterByRange.score 0:1000\ group regulation\ longLabel JASPAR Transcription Factor Binding Site Database\ maxWindowCoverage 15000\ noGenomeReason JASPAR files contain billions of items. The Table Browser allows regional queries for this track, but those may timeout if the regions are too big. See the Data Access section in the track description page for other ways to query this data, such as command-line tools and our API.\ noParentConfig on\ shortLabel JASPAR Transcription Factors\ spectrum on\ tableBrowser tbNoGenome\ track jaspar\ type bigBed 6 .\ url http://jaspar.genereg.net/search?q=$$&collection=all&tax_group=all&tax_id=all&type=all&class=all&family=all&version=all\ urlLabel View on JASPAR:\ visibility hide\ KICH KICH bigLolly 12 + Kidney Chromophobe 0 100 0 0 0 127 127 127 0 0 0 phenDis 1 autoScale on\ bigDataUrl /gbdb/hg38/gdcCancer/KICH.bb\ configurable off\ group phenDis\ lollyField 13\ longLabel Kidney Chromophobe\ parent gdcCancer off\ priority \ shortLabel KICH\ track KICH\ type bigLolly 12 +\ urls case_id=https://portal.gdc.cancer.gov/cases/193294\ kidneyStewartBroadCellType Kidney Broad CT bigBarChart Kidney RNA binned by broad cell type from Stewart et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ This track displays data from Spatiotemporal immune zonation of the human kidney. \ Droplet-based single-cell RNA sequencing (scRNA-seq) was used to profile 40,268 \ mature human kidney cells. After principal component analysis, identified clusters \ were manually curated into four major cellular compartments using canonical markers \ as found in Stewart et al., 2019: endothelial, immune, fibroblast, and epithelium.\ \
\ This track collection contains six bar chart tracks of RNA expression in the\ human kidney where cells are grouped by merged cell type \ (Kidney Cells), broad cell type \ (Kidney Broad CT), detailed cell type \ (Kidney Details), compartment\ (Kidney Compartment), experiment \ (Kidney Experiment), and project \ (Kidney Project).\ The default track displayed is \ Kidney Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| kidney specific | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ 14 mature healthy human kidney samples were obtained from individuals (ages\ 1-72) that either underwent tumor nephrectomy (n=10) or from kidneys donated\ for transplantation (n=4) but were unsuitable for use. Kidney tissues from\ tumor nephrectomies were collected from unaffected areas estimated to be\ corticomedullary. Samples were enzymatically dissociated and enriched for live\ cells (experiment set 1) or enriched for leukocytes with a density gradient and\ then for live cells (experiment set 2). Single cell libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina HiSeq4000.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. \ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Benjamin J Stewart, John R Ferdinand, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ singleCell 1 barChartBars Ascending_vasa_recta_endothelium B_cell CD4_T_cell CD8_T_cell Connecting_tubule Descending_vasa_recta_endothelium Epithelial_progenitor_cell Fibroblast Glomerular_endothelium Intercalated_cell MNP-a/classical_monocyte_derived MNP-b/non-classical_monocyte_derived MNP-c/dendritic_cell MNP-d/Tissue_macrophage Mast_cell Myofibroblast NK_cell NKT_cell Neutrophil Pelvic_epithelium Peritubular_capillary_endothelium Plasmacytoid_dendritic_cell Podocyte Principal_cell Proximal_tubule Thick_ascending_limb_of_Loop_of_Henle Transitional_urothelium\ barChartColors #5bd05a #ec374a #f7354b #f7354b #5f66ed #5fcd5b #60afce #e0cdc4 #0ab707 #181dda #e77258 #e67259 #e2745e #e8a497 #eec7c9 #c88b6c #eb384a #f4364b #e5c8c1 #5cb6cf #05bb04 #edc6c6 #9f968b #6496d4 #0e0ceb #181cd9 #bfd7e4\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/kidneyStewart/broad_celltype.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/kidneyStewart/broad_celltype.bb\ defaultLabelFields name\ html kidneyStewart\ labelFields name,name2\ longLabel Kidney RNA binned by broad cell type from Stewart et al 2019\ parent kidneyStewart\ shortLabel Kidney Broad CT\ track kidneyStewartBroadCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ kidneyStewartCellType Kidney Cells bigBarChart Kidney RNA binned by merged cell type from Stewart et al 2019 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ This track displays data from Spatiotemporal immune zonation of the human kidney. \ Droplet-based single-cell RNA sequencing (scRNA-seq) was used to profile 40,268 \ mature human kidney cells. After principal component analysis, identified clusters \ were manually curated into four major cellular compartments using canonical markers \ as found in Stewart et al., 2019: endothelial, immune, fibroblast, and epithelium.\ \
\ This track collection contains six bar chart tracks of RNA expression in the\ human kidney where cells are grouped by merged cell type \ (Kidney Cells), broad cell type \ (Kidney Broad CT), detailed cell type \ (Kidney Details), compartment\ (Kidney Compartment), experiment \ (Kidney Experiment), and project \ (Kidney Project).\ The default track displayed is \ Kidney Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| kidney specific | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ 14 mature healthy human kidney samples were obtained from individuals (ages\ 1-72) that either underwent tumor nephrectomy (n=10) or from kidneys donated\ for transplantation (n=4) but were unsuitable for use. Kidney tissues from\ tumor nephrectomies were collected from unaffected areas estimated to be\ corticomedullary. Samples were enzymatically dissociated and enriched for live\ cells (experiment set 1) or enriched for leukocytes with a density gradient and\ then for live cells (experiment set 2). Single cell libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina HiSeq4000.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. \ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Benjamin J Stewart, John R Ferdinand, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ singleCell 1 barChartBars ascending_vasa_recta_endothelial_cell B_cell T_cell_CD4+ T_cell_CD8+ connecting_tubule_cell descending_vasa_recta_endothelial_cell epithelial_progenitor_cell fibroblast glomerular_endothelial_cell intercalated_cell mononuclear_phagocyte natural_killer_cell other_immune_cell pelvic_epithelial_cell peritubular_capillary_endothelial_cell podocyte principal_cell proximal_tubule_cell thick_ascending_loop_of_Henle transitional_urothelium_cell\ barChartColors #5bd05a #ec374a #f7354b #f7354b #5f66ed #5fcd5b #60afce #c98b6b #0ab707 #181dda #de2a02 #f1374b #e7a69c #5cb6cf #05bb04 #9f968b #6496d4 #0e0ceb #181cd9 #bfd7e4\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/kidneyStewart/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/kidneyStewart/cell_type.bb\ defaultLabelFields name\ html kidneyStewart\ labelFields name,name2\ longLabel Kidney RNA binned by merged cell type from Stewart et al 2019\ parent kidneyStewart\ shortLabel Kidney Cells\ track kidneyStewartCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ kidneyStewartCompartment Kidney Compartment bigBarChart Kidney RNA binned by compartment from Stewart et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ This track displays data from Spatiotemporal immune zonation of the human kidney. \ Droplet-based single-cell RNA sequencing (scRNA-seq) was used to profile 40,268 \ mature human kidney cells. After principal component analysis, identified clusters \ were manually curated into four major cellular compartments using canonical markers \ as found in Stewart et al., 2019: endothelial, immune, fibroblast, and epithelium.\ \
\ This track collection contains six bar chart tracks of RNA expression in the\ human kidney where cells are grouped by merged cell type \ (Kidney Cells), broad cell type \ (Kidney Broad CT), detailed cell type \ (Kidney Details), compartment\ (Kidney Compartment), experiment \ (Kidney Experiment), and project \ (Kidney Project).\ The default track displayed is \ Kidney Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| kidney specific | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ 14 mature healthy human kidney samples were obtained from individuals (ages\ 1-72) that either underwent tumor nephrectomy (n=10) or from kidneys donated\ for transplantation (n=4) but were unsuitable for use. Kidney tissues from\ tumor nephrectomies were collected from unaffected areas estimated to be\ corticomedullary. Samples were enzymatically dissociated and enriched for live\ cells (experiment set 1) or enriched for leukocytes with a density gradient and\ then for live cells (experiment set 2). Single cell libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina HiSeq4000.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. \ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Benjamin J Stewart, John R Ferdinand, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ singleCell 1 barChartBars PT lymphoid myeloid non_PT\ barChartColors #0e0dea #fb344a #dd2a02 #257684\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/kidneyStewart/compartment.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/kidneyStewart/compartment.bb\ defaultLabelFields name\ html kidneyStewart\ labelFields name,name2\ longLabel Kidney RNA binned by compartment from Stewart et al 2019\ parent kidneyStewart\ shortLabel Kidney Compartment\ track kidneyStewartCompartment\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ kidneyStewartDetailedCellType Kidney Details bigBarChart Kidney RNA binned by detailed cell type from Stewart et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ This track displays data from Spatiotemporal immune zonation of the human kidney. \ Droplet-based single-cell RNA sequencing (scRNA-seq) was used to profile 40,268 \ mature human kidney cells. After principal component analysis, identified clusters \ were manually curated into four major cellular compartments using canonical markers \ as found in Stewart et al., 2019: endothelial, immune, fibroblast, and epithelium.\ \
\ This track collection contains six bar chart tracks of RNA expression in the\ human kidney where cells are grouped by merged cell type \ (Kidney Cells), broad cell type \ (Kidney Broad CT), detailed cell type \ (Kidney Details), compartment\ (Kidney Compartment), experiment \ (Kidney Experiment), and project \ (Kidney Project).\ The default track displayed is \ Kidney Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| kidney specific | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ 14 mature healthy human kidney samples were obtained from individuals (ages\ 1-72) that either underwent tumor nephrectomy (n=10) or from kidneys donated\ for transplantation (n=4) but were unsuitable for use. Kidney tissues from\ tumor nephrectomies were collected from unaffected areas estimated to be\ corticomedullary. Samples were enzymatically dissociated and enriched for live\ cells (experiment set 1) or enriched for leukocytes with a density gradient and\ then for live cells (experiment set 2). Single cell libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina HiSeq4000.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. \ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Benjamin J Stewart, John R Ferdinand, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ singleCell 1 barChartBars Ascending_vasa_recta_endothelium B_cell CD4_T_cell CD8_T_cell Connecting_tubule Descending_vasa_recta_endothelium Distinct_proximal_tubule_1 Distinct_proximal_tubule_2 Epithelial_progenitor_cell Fibroblast Glomerular_endothelium Indistinct_intercalated_cell MNP-a/classical_monocyte_derived MNP-b/non-classical_monocyte_derived MNP-c/dendritic_cell MNP-d/Tissue_macrophage Mast_cell Myofibroblast NK_cell NKT_cell Neutrophil Pelvic_epithelium Peritubular_capillary_endothelium_1 Peritubular_capillary_endothelium_2 Plasmacytoid_dendritic_cell Podocyte Principal_cell Proliferating_Proximal_Tubule Proximal_tubule Thick_ascending_limb_of_Loop_of_Henle Transitional_urothelium Type_A_intercalated_cell Type_B_intercalated_cell\ barChartColors #5bd05a #ec374a #f7354b #f7354b #5f66ed #5fcd5b #bfd5e4 #5d5df3 #60afce #e0cdc4 #0ab707 #6b6cdf #e77258 #e67259 #e2745e #e8a497 #eec7c9 #c88b6c #eb384a #f4364b #e5c8c1 #5cb6cf #07ba05 #65c860 #edc6c6 #9f968b #6496d4 #615fef #0e0dea #181cd9 #bfd7e4 #656be5 #6873df\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/kidneyStewart/detailed_cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/kidneyStewart/detailed_cell_type.bb\ defaultLabelFields name\ html kidneyStewart\ labelFields name,name2\ longLabel Kidney RNA binned by detailed cell type from Stewart et al 2019\ parent kidneyStewart\ shortLabel Kidney Details\ track kidneyStewartDetailedCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ kidneyStewartExperiment Kidney Experiment bigBarChart Kidney RNA binned by Experiment from Stewart et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ This track displays data from Spatiotemporal immune zonation of the human kidney. \ Droplet-based single-cell RNA sequencing (scRNA-seq) was used to profile 40,268 \ mature human kidney cells. After principal component analysis, identified clusters \ were manually curated into four major cellular compartments using canonical markers \ as found in Stewart et al., 2019: endothelial, immune, fibroblast, and epithelium.\ \
\ This track collection contains six bar chart tracks of RNA expression in the\ human kidney where cells are grouped by merged cell type \ (Kidney Cells), broad cell type \ (Kidney Broad CT), detailed cell type \ (Kidney Details), compartment\ (Kidney Compartment), experiment \ (Kidney Experiment), and project \ (Kidney Project).\ The default track displayed is \ Kidney Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| kidney specific | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ 14 mature healthy human kidney samples were obtained from individuals (ages\ 1-72) that either underwent tumor nephrectomy (n=10) or from kidneys donated\ for transplantation (n=4) but were unsuitable for use. Kidney tissues from\ tumor nephrectomies were collected from unaffected areas estimated to be\ corticomedullary. Samples were enzymatically dissociated and enriched for live\ cells (experiment set 1) or enriched for leukocytes with a density gradient and\ then for live cells (experiment set 2). Single cell libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina HiSeq4000.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. \ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Benjamin J Stewart, John R Ferdinand, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ singleCell 1 barChartBars PapRCC RCC1 RCC2 RCC3 Teen_Tx TxK1 TxK2 TxK3 TxK4 VHL_RCC Wilms1 Wilms2 Wilms3\ barChartColors #cec2e1 #415c71 #1712e1 #2b1fc6 #0d0cec #1d16db #6f6ddd #928faf #e03752 #100ee8 #2118d4 #7581cf #251cce\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/kidneyStewart/Experiment.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/kidneyStewart/Experiment.bb\ defaultLabelFields name\ html kidneyStewart\ labelFields name,name2\ longLabel Kidney RNA binned by Experiment from Stewart et al 2019\ parent kidneyStewart\ shortLabel Kidney Experiment\ track kidneyStewartExperiment\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ kidneyStewartProject Kidney Project bigBarChart Kidney RNA binned by project from Stewart et al 2019 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ This track displays data from Spatiotemporal immune zonation of the human kidney. \ Droplet-based single-cell RNA sequencing (scRNA-seq) was used to profile 40,268 \ mature human kidney cells. After principal component analysis, identified clusters \ were manually curated into four major cellular compartments using canonical markers \ as found in Stewart et al., 2019: endothelial, immune, fibroblast, and epithelium.\ \
\ This track collection contains six bar chart tracks of RNA expression in the\ human kidney where cells are grouped by merged cell type \ (Kidney Cells), broad cell type \ (Kidney Broad CT), detailed cell type \ (Kidney Details), compartment\ (Kidney Compartment), experiment \ (Kidney Experiment), and project \ (Kidney Project).\ The default track displayed is \ Kidney Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| kidney specific | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ 14 mature healthy human kidney samples were obtained from individuals (ages\ 1-72) that either underwent tumor nephrectomy (n=10) or from kidneys donated\ for transplantation (n=4) but were unsuitable for use. Kidney tissues from\ tumor nephrectomies were collected from unaffected areas estimated to be\ corticomedullary. Samples were enzymatically dissociated and enriched for live\ cells (experiment set 1) or enriched for leukocytes with a density gradient and\ then for live cells (experiment set 2). Single cell libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina HiSeq4000.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. \ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Benjamin J Stewart, John R Ferdinand, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ singleCell 1 barChartBars Experiment_set_1 Experiment_set_2\ barChartColors #0d0bed #c8385f\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/kidneyStewart/Project.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/kidneyStewart/Project.bb\ defaultLabelFields name\ html kidneyStewart\ labelFields name,name2\ longLabel Kidney RNA binned by project from Stewart et al 2019\ parent kidneyStewart\ shortLabel Kidney Project\ track kidneyStewartProject\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=kidney-atlas+mature-full&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ kidneyStewart Kidney Stewart Kidney single cell data from Stewart et al 2019 0 100 0 0 0 127 127 127 0 0 0\ This track displays data from Spatiotemporal immune zonation of the human kidney. \ Droplet-based single-cell RNA sequencing (scRNA-seq) was used to profile 40,268 \ mature human kidney cells. After principal component analysis, identified clusters \ were manually curated into four major cellular compartments using canonical markers \ as found in Stewart et al., 2019: endothelial, immune, fibroblast, and epithelium.\ \
\ This track collection contains six bar chart tracks of RNA expression in the\ human kidney where cells are grouped by merged cell type \ (Kidney Cells), broad cell type \ (Kidney Broad CT), detailed cell type \ (Kidney Details), compartment\ (Kidney Compartment), experiment \ (Kidney Experiment), and project \ (Kidney Project).\ The default track displayed is \ Kidney Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| kidney specific | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ 14 mature healthy human kidney samples were obtained from individuals (ages\ 1-72) that either underwent tumor nephrectomy (n=10) or from kidneys donated\ for transplantation (n=4) but were unsuitable for use. Kidney tissues from\ tumor nephrectomies were collected from unaffected areas estimated to be\ corticomedullary. Samples were enzymatically dissociated and enriched for live\ cells (experiment set 1) or enriched for leukocytes with a density gradient and\ then for live cells (experiment set 2). Single cell libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina HiSeq4000.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the UCSC Cell Browser. \ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed\ were used to transform these into a bar chart format bigBed file that can be\ visualized. The coloring was done by defining colors for the broad level cell\ classes and then using another UCSC utility, hcaColorCells, to interpolate the\ colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Benjamin J Stewart, John R Ferdinand, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ singleCell 0 group singleCell\ longLabel Kidney single cell data from Stewart et al 2019\ shortLabel Kidney Stewart\ superTrack on\ track kidneyStewart\ visibility hide\ gnomADPextKidney_Cortex Kidney-Cortex bigWig 0 1 gnomAD pext Kidney-Cortex 0 100 34 255 221 144 255 238 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Kidney_Cortex.bw\ color 34,255,221\ longLabel gnomAD pext Kidney-Cortex\ parent gnomadPext off\ shortLabel Kidney-Cortex\ track gnomADPextKidney_Cortex\ visibility hide\ liftHg19 LiftOver & ReMap chain UCSC LiftOver and NCBI ReMap: Genome alignments to convert annotations to hg19 0 100 0 0 0 127 127 127 0 0 0\ This track shows alignments from the hg38 to the hg19 genome assembly, used by the UCSC\ liftOver tool and \ NCBI's ReMap\ service, respectively.\ \
The track has three subtracks, one for UCSC and two for NCBI alignments.
\\ The alignments are shown as "chains" of alignable regions. The display is similar to\ the other chain tracks, see our \ \ chain display documentation for more information.\
\ \ \\ UCSC liftOver chain files for hg19 to hg38 can be obtained from a dedicated directory on our\ \ Download server. The NCBI chain file can be obtained from the\ \ MySQL tables directory on our download server, the filename is 'chainHg19ReMap.txt.gz'.\
\ \\ Both tables can also be explored interactively with the\ Table Browser or the\ Data Integrator.\
\ \\ Thanks to NCBI for making the ReMap data available and to Angie Hinrichs for the file conversion.\
\ map 1 compositeTrack on\ group map\ longLabel UCSC LiftOver and NCBI ReMap: Genome alignments to convert annotations to hg19\ shortLabel LiftOver & ReMap\ track liftHg19\ type chain\ visibility hide\ lincRNAsTranscripts lincRNA TUCP genePred lincRNA and TUCP transcripts 3 100 100 50 0 175 150 128 0 0 0This track displays the Human Body Map lincRNAs (large intergenic non\ coding RNAs) and TUCPs (transcripts of uncertain coding potential), as well as their\ expression levels across 22 human tissues and cell lines. The Human Body Map catalog was generated\ by integrating previously existing annotation sources with transcripts that were de-novo assembled\ from RNA-Seq data. These transcripts were collected from ~4 billion RNA-Seq reads across 24 tissues \ and cell types.
\ \Expression abundance was estimated by Cufflinks (Trapnell et al., 2010) based on RNA-Seq. \ Expression abundances were estimated on the gene locus level, rather than for each transcript \ separately and are given as raw FPKM. The prefixes tcons_ and tcons_l2_ are used to describe \ lincRNAs and TUCP transcripts, respectively. Specific details about the catalog generation and data \ sets used for this study can be found in Cabili et al (2011). Extended \ characterization of each transcript in the human body map catalog can be found at the Human lincRNA\ Catalog website.
\ \Expression abundance scores range from 0 to 1000, and are displayed from light blue to dark blue\ respectively:
\ \ \01000
\ \The body map RNA-Seq data was kindly provided by the Gene Expression\ Applications research group at Illumina.
\ \\ Cabili MN, Trapnell C, Goff L, Koziol M, Tazon-Vega B, Regev A, Rinn JL.\ \ Integrative annotation of human large intergenic noncoding RNAs reveals global properties and\ specific subclasses.\ Genes Dev. 2011 Sep 15;25(18):1915-27.\ PMID: 21890647; PMC: PMC3185964\
\ \\ Trapnell C, Williams BA, Pertea G, Mortazavi A, Kwan G, van Baren MJ, Salzberg SL, Wold BJ, Pachter\ L.\ \ Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform\ switching during cell differentiation.\ Nat Biotechnol. 2010 May;28(5):511-5.\ PMID: 20436464; PMC: PMC3146043\
\ genes 1 altColor 175,150,128\ color 100,50,0\ html lincRNAs\ longLabel lincRNA and TUCP transcripts\ noInherit on\ origAssembly hg19\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel lincRNA TUCP\ superTrack nonCodingRNAs pack\ track lincRNAsTranscripts\ type genePred\ gnomADPextLiver Liver bigWig 0 1 gnomAD pext Liver 0 100 170 187 102 212 221 178 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Liver.bw\ color 170,187,102\ longLabel gnomAD pext Liver\ parent gnomadPext off\ shortLabel Liver\ track gnomADPextLiver\ visibility hide\ liverMacParlandBroadCellType Liver Broad bigBarChart Liver cells binned by broad cell type from MacParland et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-liver&gene=$$\ This track shows data from \ Single cell RNA sequencing of human liver reveals distinct intrahepatic\ macrophage populations. Liver tissue was analyzed using droplet-based \ single-cell RNA-sequencing (scRNA-seq) and subsequent clustering distinguished 20\ hepatic cell populations based on their identified marker genes found in\ MacParland et al., 2018.
\ \\ There are three bar chart tracks in this track collection with liver cells\ grouped by either broad cell type \ (Liver Broad), specific cell type \ (Liver Cells) and donor \ (Liver Donor). The default track displayed is \ Liver Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune | |
| endothelial | |
| fibroblast | |
| epithelial | |
| stem cell | |
| hepatocyte |
\ Cells that fall into multiple classes will be colored by blending the colors associated \ with those classes. The colors will be purest in the \ Liver Cells subtrack,\ where the bars represent relatively pure cell types. They can give an overview\ of the cell composition within other categories in other subtracks as well.
\ \ \ \ \ \\ Fresh liver samples were taken from 5 neurologically deceased donors (NDD)\ deemed acceptable for liver transplantation. The caudate lobe of the liver was\ surgically separated and flushed with HTK solution to leave only tissue\ resident cells that were used to prepare a cell suspension for scRNA-seq\ analysis. Samples were prepared using 10x Genomics 3' v2 library kit and\ sequenced on the Illumina HiSeq 2500. A total of 8,444 transcriptional profiles\ were obtained for organ specific and non-organ specific cells from healthy\ hepatic tissue.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used \ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Sonya MacParland and to the many authors who worked on producing and\ publishing this data set. The data were integrated into the UCSC Genome Browser\ by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The UCSC work \ was paid for by the Chan Zuckerberg Initiative.
\ \\ MacParland SA, Liu JC, Ma XZ, Innes BT, Bartczak AM, Gage BK, Manuel J, Khuu N, Echeverri J, Linares\ I et al.\ \ Single cell RNA sequencing of human liver reveals distinct intrahepatic macrophage populations.\ Nat Commun. 2018 Oct 22;9(1):4383.\ PMID: 30348985; PMC: PMC6197289
\ singleCell 1 barChartBars B-cell Cholangiocyte Endothelial Erythroid Hepatocyte Kupffer Stellate T/NK-cell\ barChartColors #dc7b91 #908ffd #075bdb #d3c4db #af01af #d92b07 #e7cbbe #eb364f\ barChartLimit 1.5\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/liverMacParland/BroadCellType.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/liverMacParland/BroadCellType.bb\ defaultLabelFields name\ html liverMacParland\ labelFields name,name2\ longLabel Liver cells binned by broad cell type from MacParland et al 2018\ parent liverMacParland\ shortLabel Liver Broad\ track liverMacParlandBroadCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-liver&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ liverMacParlandCellType Liver Cells bigBarChart Liver cells binned by cell type from MacParland et al 2018 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-liver&gene=$$\ This track shows data from \ Single cell RNA sequencing of human liver reveals distinct intrahepatic\ macrophage populations. Liver tissue was analyzed using droplet-based \ single-cell RNA-sequencing (scRNA-seq) and subsequent clustering distinguished 20\ hepatic cell populations based on their identified marker genes found in\ MacParland et al., 2018.
\ \\ There are three bar chart tracks in this track collection with liver cells\ grouped by either broad cell type \ (Liver Broad), specific cell type \ (Liver Cells) and donor \ (Liver Donor). The default track displayed is \ Liver Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune | |
| endothelial | |
| fibroblast | |
| epithelial | |
| stem cell | |
| hepatocyte |
\ Cells that fall into multiple classes will be colored by blending the colors associated \ with those classes. The colors will be purest in the \ Liver Cells subtrack,\ where the bars represent relatively pure cell types. They can give an overview\ of the cell composition within other categories in other subtracks as well.
\ \ \ \\ Map of the human liver and its associated cell types. The liver is constructed\ of hepatic lobules which are composed of a portal triad (hepatic artery, the\ portal vein and the bile duct), hepatocytes aligned between a capillary\ network, and a central vein.\ \
\
\
\
MacParland et al. Nat\
Commun. 2018. / CC BY 4.0\
\
\
\
\ Fresh liver samples were taken from 5 neurologically deceased donors (NDD)\ deemed acceptable for liver transplantation. The caudate lobe of the liver was\ surgically separated and flushed with HTK solution to leave only tissue\ resident cells that were used to prepare a cell suspension for scRNA-seq\ analysis. Samples were prepared using 10x Genomics 3' v2 library kit and\ sequenced on the Illumina HiSeq 2500. A total of 8,444 transcriptional profiles\ were obtained for organ specific and non-organ specific cells from healthy\ hepatic tissue.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used \ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Sonya MacParland and to the many authors who worked on producing and\ publishing this data set. The data were integrated into the UCSC Genome Browser\ by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The UCSC work \ was paid for by the Chan Zuckerberg Initiative.
\ \\ MacParland SA, Liu JC, Ma XZ, Innes BT, Bartczak AM, Gage BK, Manuel J, Khuu N, Echeverri J, Linares\ I et al.\ \ Single cell RNA sequencing of human liver reveals distinct intrahepatic macrophage populations.\ Nat Commun. 2018 Oct 22;9(1):4383.\ PMID: 30348985; PMC: PMC6197289
\ singleCell 1 barChartBars B_cell cholangiocyte erythroid_cell hepatocyte macrophage_(inflammatory) liver_sinusoidal_endothelial_1_(LSEC_1) liver_sinusoidal_endothelial_2,3_(LSEC_2,3) natural_killer_like macrophage_(non-inflammatory) plasma_B_cell portal_endothelial_cell stellate_cell T_cell_alpha/beta T_cell_gamma/delta_1 T_cell_gamma/delta_2\ barChartColors #f1798a #908ffd #d3c4db #af01af #d42c0d #5e97d5 #5d8fe8 #f0798a #e3725c #c27d9a #58d05c #e7cbbe #e93650 #e87a8c #cc7d95\ barChartLimit 1.5\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/liverMacParland/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/liverMacParland/cell_type.bb\ defaultLabelFields name\ html liverMacParland\ labelFields name,name2\ longLabel Liver cells binned by cell type from MacParland et al 2018\ parent liverMacParland\ shortLabel Liver Cells\ track liverMacParlandCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-liver&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ liverMacParlandDonor Liver Donor bigBarChart Liver cells binned by organ donor from MacParland et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-liver&gene=$$\ This track shows data from \ Single cell RNA sequencing of human liver reveals distinct intrahepatic\ macrophage populations. Liver tissue was analyzed using droplet-based \ single-cell RNA-sequencing (scRNA-seq) and subsequent clustering distinguished 20\ hepatic cell populations based on their identified marker genes found in\ MacParland et al., 2018.
\ \\ There are three bar chart tracks in this track collection with liver cells\ grouped by either broad cell type \ (Liver Broad), specific cell type \ (Liver Cells) and donor \ (Liver Donor). The default track displayed is \ Liver Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune | |
| endothelial | |
| fibroblast | |
| epithelial | |
| stem cell | |
| hepatocyte |
\ Cells that fall into multiple classes will be colored by blending the colors associated \ with those classes. The colors will be purest in the \ Liver Cells subtrack,\ where the bars represent relatively pure cell types. They can give an overview\ of the cell composition within other categories in other subtracks as well.
\ \ \ \\ Contribution of cells from each liver sample to each cell cluster. Note that\ the liver number corresponds to the donor number (e.g. Liver 1 = Donor 1).
\ \\
\
\
MacParland et al. Nat\
Commun. 2018. / CC BY 4.0
\ t-SNE plot of human liver resident cells colored by source donor (Liver 1-5)\ and labeled with cluster number.
\ \\
\
\
MacParland et al. Nat\
Commun. 2018. / CC BY 4.0
\ Fresh liver samples were taken from 5 neurologically deceased donors (NDD)\ deemed acceptable for liver transplantation. The caudate lobe of the liver was\ surgically separated and flushed with HTK solution to leave only tissue\ resident cells that were used to prepare a cell suspension for scRNA-seq\ analysis. Samples were prepared using 10x Genomics 3' v2 library kit and\ sequenced on the Illumina HiSeq 2500. A total of 8,444 transcriptional profiles\ were obtained for organ specific and non-organ specific cells from healthy\ hepatic tissue.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used \ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Sonya MacParland and to the many authors who worked on producing and\ publishing this data set. The data were integrated into the UCSC Genome Browser\ by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The UCSC work \ was paid for by the Chan Zuckerberg Initiative.
\ \\ MacParland SA, Liu JC, Ma XZ, Innes BT, Bartczak AM, Gage BK, Manuel J, Khuu N, Echeverri J, Linares\ I et al.\ \ Single cell RNA sequencing of human liver reveals distinct intrahepatic macrophage populations.\ Nat Commun. 2018 Oct 22;9(1):4383.\ PMID: 30348985; PMC: PMC6197289
\ singleCell 1 barChartBars P1TLH P2TLH P3TLH P4TLH P5TLH\ barChartColors #ae3f5a #9112a6 #ad03ae #dd3751 #d63856\ barChartLimit 1.5\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/liverMacParland/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/liverMacParland/donor.bb\ defaultLabelFields name\ html liverMacParland\ labelFields name,name2\ longLabel Liver cells binned by organ donor from MacParland et al 2018\ parent liverMacParland\ shortLabel Liver Donor\ track liverMacParlandDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-liver&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ liverMacParland Liver MacParland Liver single cell sequencing from MacParland et al 2018 0 100 0 0 0 127 127 127 0 0 0\ This track shows data from \ Single cell RNA sequencing of human liver reveals distinct intrahepatic\ macrophage populations. Liver tissue was analyzed using droplet-based \ single-cell RNA-sequencing (scRNA-seq) and subsequent clustering distinguished 20\ hepatic cell populations based on their identified marker genes found in\ MacParland et al., 2018.
\ \\ There are three bar chart tracks in this track collection with liver cells\ grouped by either broad cell type \ (Liver Broad), specific cell type \ (Liver Cells) and donor \ (Liver Donor). The default track displayed is \ Liver Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| immune | |
| endothelial | |
| fibroblast | |
| epithelial | |
| stem cell | |
| hepatocyte |
\ Cells that fall into multiple classes will be colored by blending the colors associated \ with those classes. The colors will be purest in the \ Liver Cells subtrack,\ where the bars represent relatively pure cell types. They can give an overview\ of the cell composition within other categories in other subtracks as well.
\ \ \\ The default track displayed is liver RNA grouped by cell type.
\ \ \\ Fresh liver samples were taken from 5 neurologically deceased donors (NDD)\ deemed acceptable for liver transplantation. The caudate lobe of the liver was\ surgically separated and flushed with HTK solution to leave only tissue\ resident cells that were used to prepare a cell suspension for scRNA-seq\ analysis. Samples were prepared using 10x Genomics 3' v2 library kit and\ sequenced on the Illumina HiSeq 2500. A total of 8,444 transcriptional profiles\ were obtained for organ specific and non-organ specific cells from healthy\ hepatic tissue.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used \ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on \ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \ \\ Thanks to Sonya MacParland and to the many authors who worked on producing and\ publishing this data set. The data were integrated into the UCSC Genome Browser\ by Jim Kent and Brittney Wick then reviewed by Daniel Schmelter. The UCSC work \ was paid for by the Chan Zuckerberg Initiative.
\ \\ MacParland SA, Liu JC, Ma XZ, Innes BT, Bartczak AM, Gage BK, Manuel J, Khuu N, Echeverri J, Linares\ I et al.\ \ Single cell RNA sequencing of human liver reveals distinct intrahepatic macrophage populations.\ Nat Commun. 2018 Oct 22;9(1):4383.\ PMID: 30348985; PMC: PMC6197289
\ singleCell 0 group singleCell\ longLabel Liver single cell sequencing from MacParland et al 2018\ shortLabel Liver MacParland\ superTrack on\ track liverMacParland\ visibility hide\ adult_liver_models Liver models bigBed 12 + Adult Liver transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-Liver.bb\ longLabel Adult Liver transcript models\ parent sample_models_view on\ shortLabel Liver models\ subGroups view=sample_models_view sample=adult_liver type=models\ track adult_liver_models\ type bigBed 12 +\ visibility squish\ adult_liver_ont_post_models Liver ONT post models bigBed 12 + Adult Liver ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Liver01Rep1.bb\ itemRgb on\ longLabel Adult Liver ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Liver ONT post models\ subGroups view=per_expr_models_view sample=adult_liver type=post_capture_ont_models\ track adult_liver_ont_post_models\ type bigBed 12 +\ visibility hide\ adult_liver_ont_post_reads Liver ONT post reads bam Adult Liver ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Liver01Rep1.bam\ longLabel Adult Liver ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Liver ONT post reads\ subGroups view=per_expr_reads_view sample=adult_liver type=post_capture_ont_reads\ track adult_liver_ont_post_reads\ type bam\ visibility hide\ adult_liver_ont_pre_models Liver ONT pre models bigBed 12 + Adult Liver ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Liver01Rep1.bb\ itemRgb on\ longLabel Adult Liver ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Liver ONT pre models\ subGroups view=per_expr_models_view sample=adult_liver type=pre_capture_ont_models\ track adult_liver_ont_pre_models\ type bigBed 12 +\ visibility hide\ adult_liver_ont_pre_reads Liver ONT pre reads bam Adult Liver ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Liver01Rep1.bam\ longLabel Adult Liver ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Liver ONT pre reads\ subGroups view=per_expr_reads_view sample=adult_liver type=pre_capture_ont_reads\ track adult_liver_ont_pre_reads\ type bam\ visibility hide\ adult_liver_pacbio_post_models Liver PB post models bigBed 12 + Adult Liver PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Liver01Rep1.bb\ itemRgb on\ longLabel Adult Liver PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Liver PB post models\ subGroups view=per_expr_models_view sample=adult_liver type=post_capture_pacbio_models\ track adult_liver_pacbio_post_models\ type bigBed 12 +\ visibility hide\ adult_liver_pacbio_post_reads Liver PB post reads bam Adult Liver PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Liver01Rep1.bam\ longLabel Adult Liver PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Liver PB post reads\ subGroups view=per_expr_reads_view sample=adult_liver type=post_capture_pacbio_reads\ track adult_liver_pacbio_post_reads\ type bam\ visibility hide\ adult_liver_pacbio_pre_models Liver PB pre models bigBed 12 + Adult Liver PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Liver01Rep1.bb\ itemRgb on\ longLabel Adult Liver PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Liver PB pre models\ subGroups view=per_expr_models_view sample=adult_liver type=pre_capture_pacbio_models\ track adult_liver_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ adult_liver_pacbio_pre_reads Liver PB pre reads bam Adult Liver PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Liver01Rep1.bam\ longLabel Adult Liver PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Liver PB pre reads\ subGroups view=per_expr_reads_view sample=adult_liver type=pre_capture_pacbio_reads\ track adult_liver_pacbio_pre_reads\ type bam\ visibility hide\ longReadVariants Long-read SVs Structural Variants from Long-read Sequencing 0 100 0 0 0 127 127 127 0 0 0\ This track collection contains structural variant (SV) calls derived from long-read sequencing\ studies. Structural variants are genomic rearrangements larger than ~50 bp, including\ deletions, insertions, duplications, inversions, and translocations. Long-read sequencing\ technologies can span repetitive regions and resolve complex rearrangements\ that are difficult to detect with short-read methods.\
\ \\ SV length statistics (min / median / max) are computed from the svLen\ field of each track, in base pairs. Some tracks include sites with\ svLen=0 (complex events where the reference and alternate alleles\ differ in sequence but not in length).\
\\ For short-read structural-variant comparators (CCDG 17,795, 1KG 3202,\ ToMMo 48K CNV) see the companion\ Short-read SVs supertrack.\
\\ Polymorphic Mobile Element Insertions (Alu, L1, SVA, HERVK,\ snRNA) called from HGSVC3 long-read assemblies are released as a\ separate track collection; see the\ Mobile Insertions tracks. Those MEIs are\ the insertions identified in the 65 HGSVC3 samples relative to the\ reference, available on both GRCh38/hg38 and T2T-CHM13/hs1.\
\| Dataset | \N samples | \Cohort / disease | \Disease cases | \Coverage | \SV count | \Min | \Median | \Max | \
|---|---|---|---|---|---|---|---|---|
| All merged | \— | \All long-read SV datasets merged on identical position+type+length, with per-database AC | \mixed | \mixed (PacBio HiFi, ONT) | \2,317,508 | \1 | \147 | \57,207,413 | \
| CoLoRSdb | \1,427 | \Consortium of Long-Read Sequencing, joint callset | \No | \mixed (HiFi) | \426,239 | \20 | \33 | \101,381 | \
| Han 945 | \945 | \Han Chinese, general population | \No | \~17x ONT | \111,288 | \1 | \254 | \99,744 | \
| 1KG ONT 100 | \100 | \1000 Genomes, 5 superpopulations / 19 subpopulations | \No | \~37x ONT (R9.4.1) | \113,159 | \1 | \167 | \98,290 | \
| 1KG ONT Vienna | \1,019 | \1000 Genomes, diverse | \No | \~17x ONT | \148,375 | \2 | \157 | \49,171 | \
| ToMMo Japanese | \333 (111 trios) | \Japanese, general population | \No | \~22x ONT | \74,201 | \51 | \158 | \99,985 | \
| AoU 1K | \1,027 | \All of Us, self-identified Black/African American; biobank includes a variety of conditions (diabetes, hearing loss, etc.) | \Yes (mixed) | \~8x HiFi | \540,155 | \50 | \152 | \9,998 | \
| GA4K | \502 | \Children's Mercy, pediatric rare disease probands + families | \Yes (probands) | \~27x HiFi | \115,554 | \50 | \186 | \809,712 | \
| deCODE 3,622 | \3,622 | \Icelandic general population | \No | \~17x ONT | \119,453 | \1 | \154 | \861,081 | \
| HPRC v2.1 | \233 | \HPRC release-2 pangenome (CHM13 + diverse 1KG assemblies) | \No | \~60x HiFi + ~30x ONT (pangenome graph) | \549,649 | \50 | \261 | \1,064,897 | \
| HGSVC2 | \32 | \HGSVC2 haplotype-resolved assemblies (5 superpopulations) | \No | \>40x PacBio CLR + >20x HiFi (+ Strand-seq) | \111,746 | \50 | \168 | \57,207,413 | \
| HGSVC3 | \65 | \HGSVC3 diverse reference assemblies | \No | \~47x HiFi + ~56x ONT | \176,531 | \50 | \154 | \30,176,500 | \
| Arab APR | \53 | \UAE-resident Arabs from 8 countries (Arab Pangenome Reference) | \No | \~35x HiFi + ~54x ONT (+ Hi-C, pangenome graph) | \72,656 | \1 | \121 | \584,016 | \
| CPC | \58 | \Chinese Pangenome Consortium, 36 minority ethnic groups (HPRC-specific SVs removed) | \No | \~30x HiFi (pangenome graph) | \36,030 | \50 | \134 | \8,998,096 | \
| SVatalog 101 | \101 | \Cystic fibrosis (CF) patients from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT). Long-read WGS used for GWAS LD fine-mapping | \Yes (all CF) | \~50x PacBio CLR (34, Sequel I) + ~76x HiFi (67, Sequel II) | \87,068 | \4 | \160 | \1,321,484 | \
\ Note: there is likely some overlap in sample composition across these collections.\ For example, 1000 Genomes samples are also included in HPRC and CoLoRSdb.\
\ \\ Structural variants from the Consortium of Long-Read Sequencing database\ (CoLoRSdb), from 1,427 PacBio HiFi long-read whole-genome sequences.\ ~426k SVs (insertions, deletions, inversions) called with pbsv and\ merged with Jasmine, with allele frequencies, genotype counts and\ Hardy-Weinberg statistics across the cohort.\
\ \\ Structural variants from 945 Han Chinese individuals. ~111k SVs\ (deletions, insertions, duplications, inversions, translocations) merged with SURVIVOR.\ Includes allele frequencies and per-sample support.\
\ \\ Structural variants from Oxford Nanopore long-read sequencing of 100\ 1000 Genomes samples (5 superpopulations, 19 subpopulations) released\ by the 1000 Genomes ONT Sequencing Consortium and described in\ Gustafson et al. 2024. ~114k SVs (insertions, deletions, duplications,\ inversions) called with five callers and merged with Jasmine. This is a\ separate dataset from the Vienna 1KG-ONT release below; the 100 samples\ here do not overlap with the 1,019 samples in the Vienna release.\
\ \\ Structural variants from 1,019 individuals across 26 populations (1000 Genomes ONT).\ ~161k SVs annotated with SVAN, classifying insertions and deletions by mechanism\ of origin (mobile elements, VNTRs, processed pseudogenes, etc.).\ Original coordinates are on T2T-CHM13 (hs1); the hg38 version was created via liftOver.\ This is a separate dataset from the 1KG ONT 100 (Gustafson et al.) track above;\ the 1,019 samples here do not overlap with the 100 samples in that release.\
\ \\ Structural variants from 333 Japanese individuals (111 trios) from the Tohoku Medical\ Megabank (ToMMo). ~74k SVs (deletions and insertions) with trio-based Mendelian\ error rates and allele frequencies.\
\ \\ Structural variants from 1,027 individuals from the All of Us (AoU) Research Program,\ sequenced with PacBio HiFi long reads. AoU is a deeply phenotyped biobank\ that includes participants with a range of conditions (e.g. diabetes,\ hearing loss, hypertension), so the cohort is not disease-free.\ ~541k SVs (insertions and deletions) with population-specific allele\ frequencies, gene annotations, and clinical trait associations.\
\ \\ Structural variants from 502 probands and family members enrolled in the\ Genomic Answers for Kids (GA4K) pediatric rare-disease program at Children's\ Mercy Research Institute, sequenced with PacBio HiFi long reads. ~116k\ replicated SVs (deletions, insertions, duplications, inversions) called with\ pbsv and merged with JASMINE. The matched GA4K small-variant callset (SNVs\ and short indels) lives alongside other population allele-frequency resources\ as GA4K 552 PacBio LR in the Variant\ Frequencies track collection.\
\ \\ High-confidence structural variants from 3,622 Icelanders (deCODE genetics),\ sequenced with Oxford Nanopore long reads. ~134k SVs (deletions, insertions\ and combined insertion/deletion events). Site-only callset with annotated\ surrounding tandem-repeat regions.\
\ \\ Structural variants derived from the Human Pangenome Reference Consortium\ release-2.1 minigraph-cactus pangenome graph, built from 233 PacBio HiFi\ haplotype-resolved assemblies (CHM13 + diverse 1000 Genomes samples).\ About 550k SV-sized alleles (insertions and deletions) extracted from the\ graph with vg deconstruct.\
\ \\ Structural variants from 32 haplotype-resolved diploid genomes (HGSVC2\ freeze 4, Ebert et al. 2021). ~112k SVs (deletions, insertions and\ inversions) called from phased de novo assemblies with PAV, with\ per-variant 1000 Genomes population allele frequencies (insertions and\ deletions) and rich structural/gene annotations. An earlier HGSVC release\ complementary to HGSVC3.\
\ \\ Structural variants from 65 diverse individuals sequenced and de novo\ assembled by the Human Genome Structural Variation Consortium phase 3\ (HGSVC3). ~177k haplotype-resolved SVs (deletions, insertions and\ inversions) called with PAV and cross-validated with ten additional callers,\ with per-site carrier haplotype lists and structural annotations.\
\ \\ Structural variants from the Arab Pangenome Reference (APR), a\ haplotype-resolved pangenome graph built from 53 UAE-resident Arab individuals\ drawn from eight countries (PacBio HiFi + ultralong ONT + Hi-C; Nassir et al.\ 2025). ~73k SVs on hg38 (deletions, insertions, complex and mixed snarls),\ lifted from the native T2T-CHM13 assembly; the hs1 track uses the native\ coordinates.\
\ \\ Structural variants from the Chinese Pangenome Consortium (CPC), 58 samples\ spanning 36 minority ethnic groups (PacBio HiFi pangenome graph; Gao et al.\ 2023). This track shows the CPC contribution to the joint CPC+HPRC graph with\ HPRC-specific SVs removed. ~36k SVs on hg38 (deletions, insertions and mixed\ snarls), lifted from the native T2T-CHM13 assembly; the hs1 track is native.\
\ \\ Structural variants from 101 long-read whole-genome sequences released\ alongside the GWAS SVatalog tool (Chirmade et al. 2026). The samples come\ from the CF Canada-Sick Kids Program in Individual CF Therapy (CFIT), a\ cystic-fibrosis (CF) patient cohort assembled to model patient-specific\ responses to CFTR modulator therapies (most participants are F508del\ homozygotes or F508del / minimal-function compound heterozygotes; a smaller\ number carry rare nonsense or missense CFTR mutations). ~87k SVs\ (deletions, insertions, duplications, inversions and complex events)\ annotated with gene overlaps, ClinGen / gnomAD constraint scores,\ OMIM / ClinVar / DGV / Decipher regional annotations.\
\ \ \\ Each subtrack has its own documentation page with details on how to download\ and intersect the underlying annotations. The build process for all subtracks\ is recorded in the UCSC makeDoc,\ doc/hg38/lrSv.txt\ (and doc/hs1/lrSv.txt\ for T2T-CHM13); the conversion scripts are in\ makeDb/scripts/lrSv,\ and the track configuration is in\ trackDb/human/lrSv.ra.\
\ \\ Gong J, Sun H, Wang K, Zhao Y, Huang Y, Chen Q, Qiao H, Gao Y, Zhao J, Ling Y et al.\ \ Long-read sequencing of 945 Han individuals identifies structural variants associated with\ phenotypic diversity and disease susceptibility.\ Nat Commun. 2025 Feb 10;16(1):1494.\ PMID: 39929826; PMC: PMC11811171\
\ \\ Schloissnig S, Pani S, Ebler J, Hain C, Tsapalou V, Söylev A, Hüther P, Ashraf H, Prodanov T,\ Asparuhova M et al.\ \ Structural variation in 1,019 diverse humans based on long-read sequencing.\ Nature. 2025 Aug;644(8076):442-452.\ PMID: 40702182; PMC: PMC12350158\
\ \ \\ Otsuki A, Okamura Y, Ishida N, Tadaka S, Takayama J, Kumada K, Kawashima J, Taguchi K, Minegishi N,\ Kuriyama S et al.\ \ Construction of a trio-based structural variation panel utilizing activated T lymphocytes and long-\ read sequencing technology.\ Commun Biol. 2022 Sep 20;5(1):991.\ PMID: 36127505; PMC: PMC9489684\
\ \ \ \\ Garimella KV, Li Q, Wertz J, Lee SK, Cunial F, Huang Y, Mostovoy Y, Lorig-Roach R, English A, Su H\ et al.\ \ Population-scale Long-read Sequencing in the All of Us Research Program.\ medRxiv. 2025 Oct 5;.\ PMID: 41256123; PMC: PMC12622093\
\ \ \ \\ Cohen ASA, Farrow EG, Abdelmoity AT, Alaimo JT, Amudhavalli SM, Anderson JT, Bansal L, Bartik L,\ Baybayan P, Belden B et al.\ \ Genomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes.\ Genet Med. 2022 Jun;24(6):1336-1348.\ PMID: 35305867\
\ \ \ \\ Beyter D, Ingimundardottir H, Oddsson A, Eggertsson HP, Bjornsson E, Jonsson H, Atlason BA,\ Kristmundsdottir S, Mehringer S, Hardarson MT et al.\ \ Long-read sequencing of 3,622 Icelanders provides insight into the role of structural variants in\ human diseases and other traits.\ Nat Genet. 2021 Jun;53(6):779-786.\ PMID: 33972781\
\ \ \ \\ Logsdon GA, Ebert P, Audano PA, Loftus M, Porubsky D, Ebler J, Yilmaz F, Hallast P, Prodanov T, Yoo\ D et al.\ \ Complex genetic variation in nearly complete human genomes.\ Nature. 2025 Aug;644(8076):430-441.\ PMID: 40702183; PMC: PMC12350169\
\ \ \ \ \ \ \\ Chirmade S, Wang Z, Mastromatteo S, Sanders E, Thiruvahindrapuram B, Nalpathamkalam T, Pellecchia G,\ Lin F, Keenan K, Patel RV et al.\ \ GWAS SVatalog: a visualization tool to aid fine-mapping of GWAS loci with structural variations.\ Heredity (Edinb). 2026 Mar;135(3):199-210.\ PMID: 41203876; PMC: PMC13031531\
\ \ \ \\ Gustafson JA, Gibson SB, Damaraju N, Zalusky MPG, Hoekzema K, Twesigomwe D, Yang L, Snead AA,\ Richmond PA, De Coster W et al.\ \ High-coverage nanopore sequencing of samples from the 1000 Genomes Project to build a comprehensive\ catalog of human genetic variation.\ Genome Res. 2024 Nov 20;34(11):2061-2073.\ PMID: 39358015; PMC: PMC11610458\
\ \ \ \\ Ebert P, Audano PA, Zhu Q, Rodriguez-Martin B, Porubsky D, Bonder MJ, Sulovari A, Ebler J, Zhou W,\ Serra Mari R et al.\ \ Haplotype-resolved diverse human genomes and integrated analysis of structural variation.\ Science. 2021 Apr 2;372(6537).\ PMID: 33632895; PMC: PMC8026704\
\ \ \ \\ Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, Corvelo A, Clarke WE, Musunuri R,\ Nagulapalli K et al.\ \ High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602\ trios.\ Cell. 2022 Sep 1;185(18):3426-3440.e19.\ PMID: 36055201; PMC: PMC9439720\
\ \ varRep 0 filter.AC 0:30000\ filter.insLen 0:30176500\ filter.svLen 0:250000000\ filterByRange.AC on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Allele Count\ filterLabel.insLen Insertion Length (bp)\ filterLabel.svLen SV Length (bp)\ filterLabel.svType SV Type\ filterType.svType multipleListOr\ filterValues.svType DEL|DEL (Deletion),INS|INS (Insertion),INV|INV (Inversion),CPX|CPX (Complex rearrangement),DUP|DUP (Duplication),INSDEL|INSDEL (Insertion-deletion),MIXED|MIXED (multi-allele snarl),TRA|TRA (Translocation)\ group varRep\ html lrSv\ longLabel Structural Variants from Long-read Sequencing\ noScoreFilter on\ shortLabel Long-read SVs\ superTrack on\ track longReadVariants\ visibility hide\ long_read_transcripts Long-read Transcripts Transcripts and other data generated using long-read sequencing technology (PacBio and Oxford Nanopore) 0 100 0 0 0 127 127 127 0 0 0 \\ This collection is for long-read RNA-seq transcript models and primary data\ generated from experiments using third-generation sequencing technology\ (PacBio and Oxford Nanopore). The initial set is long-read models from\ ENCODE4 PacBio Iso-Seq experiments. More data sets will be added to this\ collection in the future.\
\ rna 0 group rna\ longLabel Transcripts and other data generated using long-read sequencing technology (PacBio and Oxford Nanopore)\ shortLabel Long-read Transcripts\ superTrack on\ track long_read_transcripts\ lovdComp LOVD Variants bigBed 4 + LOVD: Leiden Open Variation Database Public Variants 0 100 0 0 0 127 127 127 0 0 0NOTE:
\
LOVD is intended for use primarily by physicians and other\
professionals concerned with genetic disorders, by genetics researchers, and\
by advanced students in science and medicine. While the LOVD database is\
open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions. Further, please be\
sure to visit the LOVD web site for the very latest, as they are continually \
updating data.
DOWNLOADS:
\
LOVD databases are owned by their respective curators\
and are not available for download or mirroring \
by any third party without their permission. Batch queries on this track are only available via the\
UCSC Beacon API (see below). See also the\
LOVD web site\
for a list of database installations and the respective curators.
\ This track shows the genomic positions of all public entries in public\ installations of the Leiden Open Variation Database system (LOVD) and the effect of the \ variant, if annotated. \ Due to the copyright restrictions of the LOVD databases, UCSC is not allowed to\ host any further information. To get details on a variant (bibliographic\ reference, phenotype, disease, patient, etc.), follow the\ "Link to LOVD" to the central server at Leiden, which will then redirect you\ to the details page on the particular LOVD server reporting this variant.\
\ \\ Since Apr 2020, similar to the ClinVar track, the data is split into two subtracks, for variants\ with a length of < 50 bp and >= 50 bp, respectively.\
\ \\ LOVD is a flexible, freely-available tool for gene-centered collection and\ display of DNA variations. It is not a database itself, but rather a platform\ where curators store and analyze data. While the LOVD team and the biggest LOVD\ sites are run at the Leiden University Medical Center, LOVD installations and their\ curators are spread over the whole world. Most LOVD databases report at least \ some of their content back to Leiden to allow global cross-database search, which\ is, among others, exported to this UCSC Genome Browser track every month.\
\\ A few LOVD databases are entirely missing from this track. Reasons include configuration issues and\ intentionally blocked data search. During the last check in November 2019, the following databases\ did not export any variants:\
The LOVD data is not available for download or for batch queries in the Table Browser. \ However, it is available for programmatic access via the Global\ Alliance Beacon API, a web service that accepts queries in the form\ (genome, chromosome, position, allele) and returns "true" or "false" depending\ on whether there is information about this allele in the database. For more details see our \ Beacon Server.
\ \\ To find all LOVD databases that contain variants of a given gene, you can get a list of databases by\ constructing a url in the format geneSymbol.lovd.nl, for example,\ tp53.lovd.nl. You can\ then use the LOVD API to retrieve more detailed information from a particular database. See the\ LOVD FAQ.
\ \\ Genomic locations of LOVD variation entries are labeled with the gene symbol\ and the description of the mutation according to Human Gene Variation Society\ standards. For instance, the label AGRN:c.172G>A means that the cDNA of AGRN is\ mutated from G to A at position 172.\
\ \\ Since October 2017, the functional effect for variants is shown on the details page, if annotated.\ The possible values are:\
\ All other information is shown on the respective LOVD variation page, accessible via the\ "Link to LOVD" above.\
\ \\ The mappings displayed in this track were provided by LOVD.\
\ \\ Thanks to the LOVD team, Ivo Fokkema, Peter Taschner, Johan den Dunnen, and all LOVD curators who\ gave permission to show their data.
\ \\ Fokkema IF, Taschner PE, Schaafsma GC, Celli J, Laros JF, den Dunnen JT.\ \ LOVD v.2.0: the next generation in gene variant databases.\ Hum Mutat. 2011 May;32(5):557-63.\ PMID: 21520333\
\ phenDis 1 compositeTrack on\ group phenDis\ html lovdComp\ longLabel LOVD: Leiden Open Variation Database Public Variants\ shortLabel LOVD Variants\ tableBrowser off lovdComp\ track lovdComp\ type bigBed 4 +\ visibility hide\ lrg LRG Regions bigBed 12 + Locus Reference Genomic (LRG) / RefSeqGene Sequences Mapped to Dec. 2013 (GRCh38/hg38) Assembly 0 100 72 167 38 163 211 146 0 0 0 http://ftp.ebi.ac.uk/pub/databases/lrgex/$$.xml\ Locus Reference Genomic (LRG)\ sequences are manually curated, stable DNA sequences that surround a\ locus (typically a gene) and provide an unchanging coordinate system\ for reporting sequence variants. They are not necessarily identical\ to the corresponding sequence in a particular reference genome\ assembly (such as Dec. 2013 (GRCh38/hg38)), but can be mapped to each version of a\ reference genome assembly in order to convert between the stable LRG\ variant coordinates and the various assembly coordinates.\
\ \\ We import the data from the LRG database at the EBI. \ The NCBI RefSeqGene database is almost identical to LRG, \ but it may contain a few more sequences. See the NCBI documentation.\
\ \\ Each LRG record also includes at least one stable transcript\ on which variants may be reported. These transcripts\ appear in the LRG Transcripts track in the Gene and Gene Predictions\ track section.
\ \\ LRG sequences are suggested by the community studying a locus (for example,\ Locus-Specific Database curators, research laboratories, mutation consortia).\ LRG curators then examine the submitted transcript as well as other known\ transcripts at the locus, in the context of alignment and public expression\ data.\ For more information on the selection and annotation process, see the \ LRG FAQ,\ (Dalgleish, et al.) and (MacArthur, et al.).\
\ \\ This track was produced at UCSC using\ LRG XML files.\ Thanks to\ LRG collaborators\ for making these data available.\
\ \\ Dalgleish R, Flicek P, Cunningham F, Astashyn A, Tully RE, Proctor G, Chen Y, McLaren WM, Larsson P,\ Vaughan BW et al.\ \ Locus Reference Genomic sequences: an improved basis for describing human DNA variants.\ Genome Med. 2010 Apr 15;2(4):24.\ PMID: 20398331; PMC: PMC2873802 \
\ \\ MacArthur JA, Morales J, Tully RE, Astashyn A, Gil L, Bruford EA, Larsson P, Flicek P, Dalgleish R,\ Maglott DR et al.\ \ Locus Reference Genomic: reference sequences for the reporting of clinically relevant sequence\ variants.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D873-8.\ PMID: 24285302; PMC: PMC3965024\
\ map 1 baseColorDefault diffBases\ baseColorUseSequence lrg\ color 72,167,38\ group map\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Locus Reference Genomic (LRG) / RefSeqGene Sequences Mapped to Dec. 2013 (GRCh38/hg38) Assembly\ noScoreFilter .\ searchIndex name,ncbiAcc\ shortLabel LRG Regions\ showDiffBasesAllScales .\ track lrg\ type bigBed 12 +\ url http://ftp.ebi.ac.uk/pub/databases/lrgex/$$.xml\ urlLabel Link to LRG report:\ urls hgncId="https://www.genenames.org/data/gene-symbol-report/#!/hgnc_id/HGNC:$$" ncbiAcc="https://www.ncbi.nlm.nih.gov/nuccore/$$"\ visibility hide\ lrgTranscriptAli LRG Transcripts bigPsl Locus Reference Genomic (LRG) / RefSeqGene Fixed Transcript Annotations 0 100 54 125 29 127 127 127 0 0 0 http://ftp.ebi.ac.uk/pub/databases/lrgex/$<_lrgParent>.xml#transcripts_anchor\ This track shows the fixed (unchanging) transcript(s) associated with\ each \ Locus Reference Genomic (LRG) sequence.\ LRG\ sequences are manually curated, stable DNA sequences that surround a\ locus (typically a gene) and provide an unchanging coordinate system\ for reporting sequence variants. They are not necessarily identical\ to the corresponding sequence in a particular reference genome\ assembly (such as Dec. 2013 (GRCh38/hg38)), but can be mapped to each version of a\ reference genome assembly in order to convert between the stable LRG\ variant coordinates and the various assembly coordinates.\
\\ We import the data from the LRG database at the EBI. \ The NCBI RefSeqGene database is almost identical to LRG, \ but it may contain a few more sequences. See the NCBI documentation.\
\ \\ The LRG Regions track, in the Mapping and Sequencing Tracks section,\ includes more information about the LRG including the HGNC gene symbol\ for the gene at that locus, source of the LRG sequence, and summary of\ differences between LRG sequence and the genome assembly.\
\ \\ LRG sequences are suggested by the community studying a locus (for example,\ Locus-Specific Database curators, research laboratories, mutation consortia).\ LRG curators then examine the submitted transcript as well as other known\ transcripts at the locus, in the context of alignment and public expression\ data.\ For more information on the selection and annotation process, see the \ LRG FAQ,\ (Dalgleish, et al.) and (MacArthur, et al.).\
\ \\ This track was produced at UCSC using\ LRG XML files.\ Thanks to\ LRG\ collaborators for making these data available.\
\ \\ Dalgleish R, Flicek P, Cunningham F, Astashyn A, Tully RE, Proctor G, Chen Y, McLaren WM, Larsson P,\ Vaughan BW et al.\ \ Locus Reference Genomic sequences: an improved basis for describing human DNA variants.\ Genome Med. 2010 Apr 15;2(4):24.\ PMID: 20398331; PMC: PMC2873802\
\ \\ MacArthur JA, Morales J, Tully RE, Astashyn A, Gil L, Bruford EA, Larsson P, Flicek P, Dalgleish R,\ Maglott DR et al.\ \ Locus Reference Genomic: reference sequences for the reporting of clinically relevant sequence\ variants.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D873-8.\ PMID: 24285302; PMC: PMC3965024\
\ genes 1 altColor 127,127,127\ baseColorDefault genomicCodons\ baseColorUseSequence lfExtra\ bigDataUrl /gbdb/hg38/bbi/lrgBigPsl.bb\ color 54,125,29\ exonNumbers on\ group genes\ html lrgTranscriptAli\ indelDoubleInsert on\ indelPolyA on\ indelQueryInsert on\ longLabel Locus Reference Genomic (LRG) / RefSeqGene Fixed Transcript Annotations\ searchIndex name\ shortLabel LRG Transcripts\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ skipEmptyFields on\ skipFields mouseOver\ track lrgTranscriptAli\ type bigPsl\ url http://ftp.ebi.ac.uk/pub/databases/lrgex/$<_lrgParent>.xml#transcripts_anchor\ urlLabel Link to LRG transcript\ urls ncbiTranscript=https://www.ncbi.nlm.nih.gov/nuccore/$$ ensemblTranscript=https://www.ensembl.org/Multi/Search/Results?site=ensembl_all;q=$$ ncbiProtein=https://www.ncbi.nlm.nih.gov/protein/$$ ensemblProtein=https://www.ensembl.org/Multi/Search/Results?site=ensembl_all;q=$$\ visibility hide\ gnomADPextLung Lung bigWig 0 1 gnomAD pext Lung 0 100 153 255 0 204 255 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Lung.bw\ color 153,255,0\ longLabel gnomAD pext Lung\ parent gnomadPext off\ shortLabel Lung\ track gnomADPextLung\ visibility hide\ lungAlveoMacro448 Lung Alveolar - Macrophages - Z00000448 bigWig Methylation Atlas: Lung Alveolar - Macrophages - Z00000448 2 100 244 164 96 249 209 175 0 0 0 regulation 0 alwaysZero on\ autoScale off\ bigDataUrl /gbdb/hg38/dnaMethylationAtlas/lungAlveoMacro448.bw\ color 244,164,96\ longLabel Methylation Atlas: Lung Alveolar - Macrophages - Z00000448\ maxHeightPixels 100:70:5\ parent humanMethylationAtlasSignals off\ priority 100\ shortLabel Lung Alveolar - Macrophages - Z00000448\ subGroups cellType=Blood-Mono-Macro dataType=Replicate\ track lungAlveoMacro448\ type bigWig\ viewLimits 0:1\ windowingFunction mean\ lungTravaglini2020CellType10x Lung Cells bigBarChart Lung cells 10x method binned by merged cell type from Travaglini et al 2020 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars smooth_muscle_(airway)_cell alveolar_Type_1_cell alveolar_Type_2_cell artery/vein_endothelial_cell airway_basal_cell basophil/mast_cell bronchial_vessel_cell capillary_endothelial_cell ciliated_cell club_cell dendritic_cell fibroblast goblet_cell lymphatic_cell lymphocyte macrophage/monocyte mucous_cell other/rare_cell pericyte smooth_muscle_(vascular)_cell\ barChartColors #be04bb #905d31 #0695bc #339a1b #4a4eb4 #c82c38 #c74050 #04bd03 #0371d4 #1451e7 #e41819 #af5022 #0950f5 #ab435d #fb344b #df2901 #2652d0 #3b4ebb #a05331 #bd05b9\ barChartLimit 5\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/cell_type.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by merged cell type from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Cells\ track lungTravaglini2020CellType10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ lungTravaglini2020CellTypeFacs Lung Cells FACS bigBarChart Lung cells FACS method binned by merged cell type from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars smooth_muscle_(airway)_cell alveolar_Type_1_cell alveolar_Type_2_cell artery/vein_endothelial_cell airway_basal_cell basophil/mast_cell bronchial_vessel_cell capillary_endothelial_cell ciliated_cell club_cell dendritic_cell fibroblast goblet_cell lymphatic_cell lymphocyte macrophage/monocyte mucous_cell other/rare_cell pericyte smooth_muscle_(vascular)_cell\ barChartColors #be04bb #a63276 #0497be #23a218 #a33b7b #e5171b #7a555b #02be01 #0272d5 #2450d5 #d02a1e #af5021 #0750f6 #7e5164 #fd334a #df2901 #c5341d #bd356d #b514a7 #be04bb\ barChartLimit 900\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/cell_type.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/cell_type.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by merged cell type from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Cells FACS\ track lungTravaglini2020CellTypeFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020Compartment10x Lung Compart bigBarChart Lung cells 10x method binned by compartment from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars endothelial epithelial immune stromal\ barChartColors #0ab906 #0894bb #dd2a03 #ad4d2d\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/compartment.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/compartment.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by compartment from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Compart\ track lungTravaglini2020Compartment10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020CompartmentFacs Lung Compart FACS bigBarChart Lung cells FACS method binned by compartment from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars endothelial epithelial immune stromal\ barChartColors #03bd02 #0497be #fc334b #b9149d\ barChartLimit 300\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/compartment.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/compartment.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by compartment from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Compart FACS\ track lungTravaglini2020CompartmentFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020DetailedCellType10x Lung Detail bigBarChart Lung cells 10x method binned by detailed cell type from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars Adventitial_Fibroblast_P1 Adventitial_Fibroblast_P2 Adventitial_Fibroblast_P3 Airway_Smooth_Muscle_P1 Airway_Smooth_Muscle_P2 Airway_Smooth_Muscle_P3 Alveolar_Epithelial_Type_1_P1 Alveolar_Epithelial_Type_1_P2 Alveolar_Epithelial_Type_1_P3 Alveolar_Epithelial_Type_2_P1 Alveolar_Epithelial_Type_2_P2 Alveolar_Epithelial_Type_2_P3 Alveolar_Fibroblast_P1 Alveolar_Fibroblast_P2 Alveolar_Fibroblast_P3 Artery_P1 Artery_P2 Artery_P3 B_P1 B_P2 B_P3 Basal_P1 Basal_P2 Basal_P3 Basophil/Mast_1_P1 Basophil/Mast_1_P2 Basophil/Mast_1_P3 Basophil/Mast_2_P3 Bronchial_Vessel_1_P1 Bronchial_Vessel_1_P3 Bronchial_Vessel_2_P1 Bronchial_Vessel_2_P3 CD4+_Memory/Effector_T_P1 CD4+_Memory/Effector_T_P2 CD4+_Memory/Effector_T_P3 CD4+_Naive_T_P1 CD4+_Naive_T_P2 CD4+_Naive_T_P3 CD8+_Memory/Effector_T_P1 CD8+_Memory/Effector_T_P2 CD8+_Memory/Effector_T_P3 CD8+_Naive_T_P1 CD8+_Naive_T_P2 CD8+_Naive_T_P3 Capillary_Aerocyte_P1 Capillary_Aerocyte_P2 Capillary_Aerocyte_P3 Capillary_Intermediate_1_P2 Capillary_Intermediate_2_P2 Capillary_P1 Capillary_P2 Capillary_P3 Ciliated_P1 Ciliated_P2 Ciliated_P3 Classical_Monocyte_P1 Classical_Monocyte_P2 Classical_Monocyte_P3 Club_P1 Club_P2 Club_P3 Differentiating_Basal_P1 Differentiating_Basal_P3 EREG+_Dendritic_P1 EREG+_Dendritic_P2 Fibromyocyte_P3 Goblet_P3 IGSF21+_Dendritic_P1 IGSF21+_Dendritic_P2 IGSF21+_Dendritic_P3 Intermediate_Monocyte_P2 Ionocyte_P3 Lipofibroblast_P1 Lymphatic_P1 Lymphatic_P2 Lymphatic_P3 Macrophage_P1 Macrophage_P2 Macrophage_P3 Mesothelial_P1 Mucous_P2 Mucous_P3 Myeloid_Dendritic_Type_1_P1 Myeloid_Dendritic_Type_1_P2 Myeloid_Dendritic_Type_1_P3 Myeloid_Dendritic_Type_2_P1 Myeloid_Dendritic_Type_2_P2 Myeloid_Dendritic_Type_2_P3 Myofibroblast_P1 Myofibroblast_P2 Myofibroblast_P3 Natural_Killer_T_P2 Natural_Killer_T_P3 Natural_Killer_P1 Natural_Killer_P2 Natural_Killer_P3 Neuroendocrine_P3 Nonclassical_Monocyte_P1 Nonclassical_Monocyte_P2 Nonclassical_Monocyte_P3 OLR1+_Classical_Monocyte_P2 Pericyte_P1 Pericyte_P2 Pericyte_P3 Plasma_P1 Plasma_P3 Plasmacytoid_Dendritic_P1 Plasmacytoid_Dendritic_P2 Plasmacytoid_Dendritic_P3 Platelet/Megakaryocyte_P1 Platelet/Megakaryocyte_P3 Proliferating_Basal_P1 Proliferating_Basal_P3 Proliferating_Macrophage_P1 Proliferating_Macrophage_P2 Proliferating_Macrophage_P3 Proliferating_NK/T_P2 Proliferating_NK/T_P3 Proximal_Basal_P3 Proximal_Ciliated_P3 Serous_P3 Signaling_Alveolar_Epithelial_Type_2_P3 TREM2+_Dendritic_P1 TREM2+_Dendritic_P3 Vascular_Smooth_Muscle_P2 Vascular_Smooth_Muscle_P3 Vein_P1 Vein_P2 Vein_P3\ barChartColors #c18a7a #d8b0a5 #aa502a #d15dca #b90eab #bc06b8 #dcd0c3 #965a2f #886036 #0596bc #0496bd #0695bc #e6cbc0 #ab5027 #ab5027 #b29a79 #36971e #588328 #ed7a8b #f2c5cc #e63750 #e4c4ce #695095 #b63c6a #e16d74 #c42f3c #cc2932 #c72e3a #d47e90 #d33a51 #96a27c #73bf65 #ea3750 #ee374c #f1374c #e53752 #f47989 #ea3750 #f3364c #ef374c #f87988 #f3364b #ed384b #f4364b #60cd5b #09b905 #15b00c #06bb04 #c18378 #0eb608 #0cb707 #10b409 #1b51de #0e6eca #0d6ecc #cf2733 #cf2530 #ca2936 #1851e2 #1851e2 #1d51dc #dea3b0 #1450e8 #f4c0b6 #ed6567 #cf65bb #0950f5 #f5bcbc #ea6769 #e96869 #dd1d21 #cad5df #d0afb4 #e3c8d0 #ab445d #c88194 #de2a02 #de2a02 #de2a02 #e0a3ad #2652cf #2452d2 #f3bfc1 #f09a9c #ec9d9f #ee9b9e #ef9b9e #dd1c21 #e6c7ca #d9adac #d27b8e #ef7c87 #ef7c87 #f0374b #ca4943 #ea3a4b #f2d9de #d92027 #d82128 #db1e24 #cc262e #d7aab5 #a25231 #b79276 #d1afb9 #974b62 #f5c3ca #f3a6b1 #f1a6b0 #eac3c3 #eca295 #e9d9e2 #8e88c3 #f0a190 #dd2a04 #dd2a04 #f2a8b0 #eaa6af #594ba7 #1b69c1 #a8b2e0 #0695bc #efa190 #db2b06 #d062c1 #bd06b8 #96aa71 #34991b #ab9b78\ barChartLimit 7\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/detailed_cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/detailed_cell_type.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by detailed cell type from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Detail\ track lungTravaglini2020DetailedCellType10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020DetailedCellTypeFacs Lung Detail FACS bigBarChart Lung cells FACS method binned by detailed cell type from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars Adventitial_Fibroblast_P1 Adventitial_Fibroblast_P2 Adventitial_Fibroblast_P3 Airway_Smooth_Muscle_P1 Airway_Smooth_Muscle_P2 Airway_Smooth_Muscle_P3 Alveolar_Epithelial_Type_1_P1 Alveolar_Epithelial_Type_1_P2 Alveolar_Epithelial_Type_1_P3 Alveolar_Epithelial_Type_2_P1 Alveolar_Epithelial_Type_2_P2 Alveolar_Epithelial_Type_2_P3 Alveolar_Fibroblast_P1 Alveolar_Fibroblast_P2 Alveolar_Fibroblast_P3 Artery_P1 Artery_P2 Artery_P3 B_P1 B_P2 B_P3 Basal_P1 Basal_P2 Basal_P3 Basophil/Mast_1_P1 Basophil/Mast_1_P2 Basophil/Mast_1_P3 Bronchial_Vessel_1_P1 CD4+_Memory/Effector_T_P1 CD4+_Naive_T_P1 CD4+_Naive_T_P2 CD8+_Memory/Effector_T_P1 CD8+_Naive_T_P1 CD8+_Naive_T_P2 Capillary_Aerocyte_P1 Capillary_Aerocyte_P2 Capillary_Aerocyte_P3 Capillary_Intermediate_1_P2 Capillary_P1 Capillary_P2 Capillary_P3 Ciliated_P1 Ciliated_P2 Ciliated_P3 Classical_Monocyte_P1 Club_P1 Club_P2 Club_P3 Dendritic_P1 Differentiating_Basal_P3 Fibromyocyte_P3 Goblet_P1 Goblet_P2 Goblet_P3 IGSF21+_Dendritic_P2 IGSF21+_Dendritic_P3 Intermediate_Monocyte_P2 Intermediate_Monocyte_P3 Ionocyte_P3 Lipofibroblast_P1 Lymphatic_P1 Lymphatic_P2 Lymphatic_P3 Macrophage_P2 Macrophage_P3 Myeloid_Dendritic_Type_2_P3 Myofibroblast_P2 Myofibroblast_P3 Natural_Killer_T_P2 Natural_Killer_T_P3 Natural_Killer_P1 Natural_Killer_P2 Natural_Killer_P3 Neuroendocrine_P1 Neuroendocrine_P3 Neutrophil_P1 Neutrophil_P2 Neutrophil_P3 Nonclassical_Monocyte_P1 Nonclassical_Monocyte_P2 Pericyte_P1 Pericyte_P2 Pericyte_P3 Plasma_P3 Plasmacytoid_Dendritic_P1 Plasmacytoid_Dendritic_P2 Plasmacytoid_Dendritic_P3 Proliferating_NK/T_P2 Proliferating_NK/T_P3 Signaling_Alveolar_Epithelial_Type_2_P1 Signaling_Alveolar_Epithelial_Type_2_P3 Vascular_Smooth_Muscle_P1 Vascular_Smooth_Muscle_P2 Vascular_Smooth_Muscle_P3 Vein_P2\ barChartColors #a84b36 #a44e34 #aa4f2c #bd06b8 #b80daf #bc08b5 #a93276 #983e69 #a83177 #0596bc #0497bd #0496bd #ac4d2c #aa4f2a #ad4f26 #379323 #1ea518 #b08b8d #a14746 #d27977 #d57679 #be3372 #78488f #b63670 #e3181d #ea676b #e01a1e #7a555b #d93357 #d83555 #ec788e #d13755 #f6344d #ed3550 #1ca912 #07ba05 #1ea712 #06bb04 #21a515 #06bb05 #22a415 #0d6ecc #0471d3 #0f6dca #c13628 #2b50ce #1951e0 #2651d2 #dd7170 #a93c75 #b714a0 #768adb #0850f5 #5a8af9 #de7561 #d77966 #d47a68 #e4735c #c67ca1 #a34252 #8d4960 #71566c #ae8a92 #d92c07 #e5735b #d97571 #b38793 #b32e6e #f37989 #f0344f #f9344c #f8344c #f9344c #ba366e #d67a9b #bd3728 #e1a69e #d87869 #aa4237 #df7561 #bb0fa9 #b018a4 #bb12a5 #b88493 #d57877 #d87573 #eaa2a5 #f2798a #f5788a #0596bd #0397be #bb08b4 #b217a0 #bc09b3 #41892b\ barChartLimit 1200\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/detailed_cell_type.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/detailed_cell_type.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by detailed cell type from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Detail FACS\ track lungTravaglini2020DetailedCellTypeFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020Donor10x Lung Donor bigBarChart Lung cells 10x method binned by organ donor from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars 1 2 3\ barChartColors #da2b07 #d12425 #ba352f\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/donor.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by organ donor from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Donor\ track lungTravaglini2020Donor10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020DonorFacs Lung Donor FACS bigBarChart Lung cells FACS method binned by organ donor from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars 1 2 3\ barChartColors #168cb3 #1f86aa #0b93b9\ barChartLimit 200\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/donor.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/donor.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by organ donor from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Donor FACS\ track lungTravaglini2020DonorFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020GatingFacs Lung Gating FACS bigBarChart Lung cells FACS method binned by gating from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars Bcell CD45+_Epcam- CD45-_Epcam+ CD45-_Epcam- NK cd4 cd8 monocyte nan wbc\ barChartColors #944c4a #f7334c #0a93b9 #a52b8a #d63852 #bc3b56 #da3654 #b63b31 #138eb3 #cf3651\ barChartLimit 600\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/gating.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/gating.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by gating from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Gating FACS\ track lungTravaglini2020GatingFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020HalfDetailedCellType10x Lung Half Det bigBarChart Lung cells 10x method binned by halfway detailed cell type from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars Adventitial_Fibroblast Airway_Smooth_Muscle Alveolar_Epithelial_Type_1 Alveolar_Epithelial_Type_2 Alveolar_Fibroblast Artery B Basal Basophil/Mast_1 Basophil/Mast_2 Bronchial_Vessel_1 Bronchial_Vessel_2 CD4+_Memory/Effector_T CD4+_Naive_T CD8+_Memory/Effector_T CD8+_Naive_T Capillary Capillary_Aerocyte Capillary_Intermediate_1 Capillary_Intermediate_2 Ciliated Classical_Monocyte Club Differentiating_Basal EREG+_Dendritic Fibromyocyte Goblet IGSF21+_Dendritic Intermediate_Monocyte Ionocyte Lipofibroblast Lymphatic Macrophage Mesothelial Mucous Myeloid_Dendritic_Type_1 Myeloid_Dendritic_Type_2 Myofibroblast Natural_Killer Natural_Killer_T Neuroendocrine Nonclassical_Monocyte OLR1+_Classical_Monocyte Pericyte Plasma Plasmacytoid_Dendritic Platelet/Megakaryocyte Proliferating_Basal Proliferating_Macrophage Proliferating_NK/T Proximal_Basal Proximal_Ciliated Serous Signaling_Alveolar_Epithelial_Type_2 TREM2+_Dendritic Vascular_Smooth_Muscle Vein\ barChartColors #aa4f2a #be04bb #905d31 #0695bc #ac5026 #37971d #e83750 #914680 #c72c38 #c72e3a #d03a52 #82b46d #f3364c #e93750 #f4364c #f4364b #0ab906 #0ab806 #06bb04 #c18378 #0471d3 #cd2734 #1451e7 #1750e5 #e4181b #cf65bb #0950f5 #e01b1d #dd1d21 #cad5df #d0afb4 #ab435d #df2901 #e0a3ad #2652d0 #dd1d21 #dd1c21 #b3414a #e03e48 #f27b87 #f2d9de #da1f25 #cc262e #a05331 #974a62 #ec7989 #eba295 #9088c2 #dd2a03 #e87c89 #594ba7 #1b69c1 #a8b2e0 #0695bc #dc2a05 #bd05b9 #35991b\ barChartLimit 6\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/half_merged.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/half_merged.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by halfway detailed cell type from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Half Det\ track lungTravaglini2020HalfDetailedCellType10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020HalfDetailedFacs Lung Half Det FACS bigBarChart Lung cells FACS method binned by merged cell type from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars Adventitial_Fibroblast Airway_Smooth_Muscle Alveolar_Epithelial_Type_1 Alveolar_Epithelial_Type_2 Alveolar_Fibroblast Artery B Basal Basophil/Mast_1 Bronchial_Vessel_1 CD4+_Memory/Effector_T CD4+_Naive_T CD8+_Memory/Effector_T CD8+_Naive_T Capillary Capillary_Aerocyte Capillary_Intermediate_1 Ciliated Classical_Monocyte Club Dendritic Differentiating_Basal Fibromyocyte Goblet IGSF21+_Dendritic Intermediate_Monocyte Ionocyte Lipofibroblast Lymphatic Macrophage Myeloid_Dendritic_Type_2 Myofibroblast Natural_Killer Natural_Killer_T Neuroendocrine Neutrophil Nonclassical_Monocyte Pericyte Plasma Plasmacytoid_Dendritic Proliferating_NK/T Signaling_Alveolar_Epithelial_Type_2 Vascular_Smooth_Muscle Vein\ barChartColors #ab4e2b #be04bb #a63276 #0497bd #ad4f25 #20a516 #ae3f3f #a13b7d #e5171b #7a555b #d93357 #dd3554 #d13755 #f7344d #06bb04 #06bb04 #06bb04 #0272d5 #c13628 #2450d5 #dd7170 #a93c75 #b714a0 #0750f6 #de7661 #e6725a #c67ca1 #a34252 #7e5164 #db2b06 #d97571 #aa3c5a #fb344b #f2344e #bc356d #c5341d #c93319 #b514a7 #b88493 #c82f2c #f1344e #0496bd #be04bb #41892b\ barChartLimit 900\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/half_merged.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/half_merged.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by merged cell type from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Half Det FACS\ track lungTravaglini2020HalfDetailedFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020LabelFacs Lung Label FACS bigBarChart Lung cells FACS method binned by label from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars Ecpam,_CD45 Epcam_(+) Epcam_(-) na\ barChartColors #138eb3 #0a93b9 #f03351 #c93952\ barChartLimit 600\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/label.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/label.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by label from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Label FACS\ track lungTravaglini2020LabelFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020Location10x Lung Locat bigBarChart Lung cells 10x method binned by location from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars blood distal medial proximal\ barChartColors #ec364e #d62c0d #d02426 #0b92b9\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/location.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/location.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by location from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Locat\ track lungTravaglini2020Location10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020LocationFacs Lung Locat FACS bigBarChart Lung cells FACS method binned by location from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars blood distal medial proximal\ barChartColors #c93952 #138eb4 #178bb1 #0497bd\ barChartLimit 400\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/location.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/location.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by location from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Locat FACS\ track lungTravaglini2020LocationFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020MagneticSelection10x Lung Mag Sel bigBarChart Lung cells 10x method binned by magnetic.selection from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars blood epithelial immune_and_endothelial stromal\ barChartColors #ec364e #2c7ea1 #de1c1e #dd2a04\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/magnetic.selection.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/magnetic.selection.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by magnetic.selection from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Mag Sel\ track lungTravaglini2020MagneticSelection10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020Organ10x Lung Organ bigBarChart Lung cells 10x method binned by organ from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars blood lung\ barChartColors #ec364e #d22a18\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/organ.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/organ.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by organ from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Organ\ track lungTravaglini2020Organ10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020OrganFacs Lung Organ FACS bigBarChart Lung cells FACS method binned by organ from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars blood lung\ barChartColors #c93952 #108fb5\ barChartLimit 600\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/organ.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/organ.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by organ from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Organ FACS\ track lungTravaglini2020OrganFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020Sample10x Lung Sample bigBarChart Lung cells 10x method binned by sample from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars blood_1 blood_3 distal_1a distal_2 distal_3 medial_2 proximal_3\ barChartColors #ed364e #eb364e #dc2a05 #d12424 #d42d0e #d02426 #0b92b9\ barChartLimit 4\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/sample.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/droplet/sample.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells 10x method binned by sample from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Sample\ track lungTravaglini2020Sample10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+droplet&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020SampleFacs Lung Sample FACS bigBarChart Lung cells FACS method binned by sample from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 1 barChartBars blood_1 distal_1a distal_1b distal_2 distal_3 medial_2 medial_3 proximal_3\ barChartColors #c93952 #0795bb #9c3c84 #2482a5 #1090b6 #188aaf #1d89af #0497bd\ barChartLimit 400\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/sample.stats\ barChartUnit count/cell\ bigDataUrl /gbdb/hg38/bbi/lungTravaglini2020/facs/sample.bb\ defaultLabelFields name\ html lungTravaglini2020\ labelFields name,name2\ longLabel Lung cells FACS method binned by sample from Travaglini et al 2020\ parent lungTravaglini2020\ shortLabel Lung Sample FACS\ track lungTravaglini2020SampleFacs\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=stanford-czb-hlca+facs&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ lungTravaglini2020 Lung Travaglini Lung cells from from Travaglini et al 2020 0 100 0 0 0 127 127 127 0 0 0\ This track displays data from A\ molecular cell atlas of the human lung from single-cell RNA\ sequencing. Using droplet-based and plate-based single-cell RNA\ sequencing (scRNA-seq), 58 lung cell type populations were identified: \ 15 epithelial, 9 endothelial, 9 stromal, and 25 immune. This dataset \ covers ~75,000 human cells across all lung tissue compartments and\ circulating blood.
\ \\ This track collection contains 19 bar chart tracks of RNA expression in the human lung where cells \ are grouped such as by cell type (Lung Cells, \ Lung Cells FACS), tissue compartments \ (Lung Compart, \ Lung Compart FACS), \ detailed cell type (Lung Detail, \ Lung Detail FACS), \ organ donor (Lung Donor, \ Lung Donor FACS), halfway detailed cell type \ (Lung Half Det, \ Lung Half Det FACS), \ sample location (Lung Locat, \ Lung Locat FACS), or organ \ (Lung Organ, \ Lung Organ FACS). \ The default track displayed is Lung Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the Lung Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Healthy lung tissue and peripheral blood was surgically removed from 2 male\ patients (ages 46 and 75) and 1 female patient (age 51) undergoing lobectomy\ for focal lung tumors. Lung tissue was sampled from the bronchi (proximal),\ bronchiole (medial), and alveolar (distal) regions. Lung samples were\ dissociated and enriched with magnetic columns before being sorted into\ epithelial, endothelial/immune, and stromal cell suspensions. Lung and\ peripheral blood libraries were prepared using the 10x Genomics 3' v2 kit. In\ parallel, Smart-Seq2 (SS2) cDNA libraries were prepared using the Nextera XT\ library kit. Both 10x and SS2 libraries were sequenced on a NovaSeq 6000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Kyle J. Travaglini, Ahmad N. Nabhan, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ singleCell 0 group singleCell\ longLabel Lung cells from from Travaglini et al 2020\ shortLabel Lung Travaglini\ superTrack on\ track lungTravaglini2020\ visibility hide\ mane MANE bigGenePred MANE Select Plus Clinical: Representative transcript from RefSeq & GENCODE 3 100 0 0 0 127 127 127 0 0 0\ The Matched Annotation from\ NCBI and EMBL-EBI (MANE) project aims to produce a matched set of \ high-confidence transcripts that are identically annotated between RefSeq (NCBI) and \ Ensembl/GENCODE (led by EMBL-EBI). Transcripts for MANE are chosen by a combination of \ automated and manual methods based on conservation, expression levels, clinical significance, \ and other factors. Transcripts are matched between the NCBI RefSeq and Ensembl/GENCODE annotations\ based on the GRCh38 genome assembly, with precise 5' and 3' ends defined by high-throughput\ sequencing or other available data.
\\ This track is automatically updated, see the source data version above for the current\ version number. MANE includes almost all human protein-coding genes and genes of clinical relevance,\ including genes in the\ American\ College of Medical Genetics and Genomics (ACMG) Secondary Findings list (SF) v3.0. It includes \ both MANE Select and MANE Plus Clinical transcripts. MANE\ Plus Clinical items are colored red.\
\ For more information on the different gene tracks, including MANE vs GENCODE or RefSeq,\ see our Genes FAQ.
\ \\ The raw data can be explored interactively with the Table Browser, or the Data Integrator. For computational analysis, genome annotations are stored in\ a bigGenePred file that can be downloaded from the\ download\ server. Regional or genome-wide annotations can be converted from binary data to human readable\ text using our command line utility bigBedToBed which can be compiled from source code or\ downloaded as a precompiled binary for your system. Files and instructions can be found in the\ utilities directory.\ \ The utility can be used to obtain features within a given range, for example:
\bigBedToBed -chrom=chr6 -start=0 -end=1000000 http://hgdownload.soe.ucsc.edu/gbdb/hg38/mane/mane.bb stdout\
\
\ Download links for MANE:\ ftp://ftp.ncbi.nlm.nih.gov/refseq/MANE\
\ \\ Previous MANE versions are also available on our download archive.
\ \\ Please refer to our Data Access FAQ\ for more information or our mailing list for archived user questions.
\ \\ Thank you to the RefSeq project at NCBI and the Ensembl/GENCODE project at EMBL-EBI.\ You can contact the authors directly at \ MANE-help@ncbi.nlm.nih.gov\ or \ mane-help@ebi.ac.uk.
\ \\ Morales J, Pujar S, Loveland JE, Astashyn A, Bennett R, Berry A, Cox E, Davidson C, Ermolaeva O,\ Farrell CM et al.\ \ A joint NCBI and EMBL-EBI transcript set for clinical genomics and research.\ Nature. 2022 Apr;604(7905):310-315.\ PMID: 35388217; PMC: PMC9007741\
\ genes 1 baseColorDefault genomicCodons\ bigDataUrl /gbdb/hg38/mane/mane.bb\ dataVersion /gbdb/hg38/mane/README_versions.txt\ defaultLabelFields geneName2\ group genes\ itemRgb on\ labelFields geneName2,name,ensemblProtAcc,geneName,ncbiId,ncbiProtAcc,ncbiGene\ longLabel MANE Select Plus Clinical: Representative transcript from RefSeq & GENCODE\ maxItems 5000\ mouseOver $ncbiId, $name\ searchIndex name\ searchTrix /gbdb/hg38/mane/mane.ix\ shortLabel MANE\ skipFields cdsStartStat,cdsEndStat,exonFrames,geneType,type\ track mane\ type bigGenePred\ urls name2="https://www.ensembl.org/Homo_sapiens/Transcript/Summary?t=$$" geneName="https://www.ensembl.org/homo_sapiens/Gene/Summary?g=$$&db=core" geneName2="https://www.genecards.org/cgi-bin/carddisp.pl?gene=$$" ensemblProtAcc="https://www.ensembl.org/Homo_sapiens/Transcript/Summary?t=$$" ncbiId="https://www.ncbi.nlm.nih.gov/nuccore/$$" ncbiProtAcc="https://www.ncbi.nlm.nih.gov/nuccore/$$"\ visibility pack\ mappability Mappability Hoffman Lab Umap and Bismap Mappability 0 100 0 0 0 127 127 127 0 0 0\ These tracks indicate regions with uniquely mappable reads of particular lengths before and after\ bisulfite conversion. Both Umap and Bismap tracks contain single-read mappability and multi-read\ mappability tracks for four different read lengths: 24 bp, 36 bp, 50 bp, and 100 bp.
\\ You can use these tracks for many purposes, including filtering unreliable signal from\ sequencing assays. The Bismap track can help filter unreliable signal from sequencing assays\ involving bisulfite conversion, such as whole-genome bisulfite sequencing or reduced representation\ bisulfite sequencing.
\ \ \These tracks mark any region of the bisulfite-converted genome that is uniquely mappable by\ at least one k-mer on the specified strand. Mappability of the forward strand was\ generated by converting all instances of cytosine to thymine. Similarly, mappability of the\ reverse strand was generated by converting all instances of guanine to adenine.
\To calculate the single-read mappability, you must find the overlap of a given region with\ the region that is uniquely mappable on both strands. Regions not uniquely mappable on both\ strands or have a low multi-read mappability might bias the downstream analysis.
These tracks represent the probability that a randomly selected k-mer which overlaps\ with a given position is uniquely mappable. Multi-read mappability track is calculated for\ k-mers that are uniquely mappable on both strands, and thus there is no strand\ specification.
These tracks mark any region of the genome that is uniquely mappable by at least one\ k-mer. To calculate the single-read mappability, you must find the overlap of a given\ region with this track.
These tracks represent the probability that a randomly selected k-mer which overlaps\ with a given position is uniquely mappable.
For greater detail and explanatory diagrams, see the\ preprint, the\ Umap and Bismap project website, or the\ Umap and Bismap software\ documentation.\ \
\ The raw data can be explored interactively with the Table Browser, or the Data Integrator. For automated analysis, genome annotation is stored in a bigBed\ or bigWig file that can be downloaded from the\ download\ server. Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed or bigWigToWig, which can be compiled from the source code or\ downloaded as a precompiled binary for your system. Instructions for downloading source code and\ binaries can be found here.\ The tool can also be used to obtain only features within a given range, for example:
\ bigBedToBed -chrom=chr6 -start=0 -end=1000000\ http://hgdownload.soe.ucsc.edu/gbdb/hg38/hoffmanMappability/k24.Unique.Mappability.bb stdout\\ Please refer to our mailing list archives for questions, or our\ Data Access FAQ for more\ information.
\ \\ Anshul Kundaje (Stanford\ University) created the original Umap software in MATLAB. The original Umap repository is available\ here.\ Mehran Karimzadeh (Michael Hoffman\ lab, Princess Margaret Cancer Centre) implemented the Python version of Umap and added features,\ including Bismap.
\ \\ Karimzadeh M, Ernst C, Kundaje A, Hoffman MM.,\ Umap and Bismap:\ quantifying genome and methylome mappability\ bioRxiv bioRxiv, p. 095463, 2016.; doi: https://doi.org/10.1101/095463.
\ map 0 group map\ longLabel Hoffman Lab Umap and Bismap Mappability\ shortLabel Mappability\ superTrack on\ track mappability\ mavedb MaveDB Experiments bigBed 12 + Heatmaps and Alignment for MaveDB 0 100 0 0 0 127 127 127 0 0 0\ This supertrack provides heatmaps of multiplexed assays of variant effects (MAVE) from\ MaveDB. Each heatmap presents the results of an\ experiment where many small substitutions were tested within a gene to examine their\ functional consequences. Accompanying tracks display alignments of each experiment sequence\ to the genome.\
\\ Direct access to the data files for these experiments can be obtained from\ MaveDB.\
\\ Rubin AF, Stone J, Bianchi AH, Capodanno BJ, Da EY, Dias M, Esposito D, Frazer J, Fu Y, Grindstaff\ SB et al.\ \ MaveDB 2024: a curated community database with over seven million variant effects from multiplexed\ functional assays.\ Genome Biol. 2025 Jan 21;26(1):13.\ PMID: 39838450; PMC: PMC11753097\
\ expression 1 group expression\ longLabel Heatmaps and Alignment for MaveDB\ shortLabel MaveDB Experiments\ superTrack on\ track mavedb\ type bigBed 12 +\ metamorf MetamORF bigGenePred ncORFs: MetamORF - meta-database of non-canonical ORFs 3 100 0 0 0 127 127 127 0 0 0\ This track displays 664,558 unique small open reading frames (sORFs) in the human\ genome from MetamORF, a\ meta-database that consolidates sORF data identified by both experimental and computational\ approaches. sORFs are defined as ORFs encoding fewer than 100 amino acids (excluding stop\ codons and introns).\
\ \\ MetamORF was built by gathering publicly available sORF data from multiple sources,\ normalizing it, and removing redundancy. From 2,594,154 source ORFs across human and mouse,\ MetamORF identified 1,162,675 unique ORFs (664,771 human, 497,904 mouse) associated with\ 153,553 unique transcripts. The database enables comparison of sORFs across distinct original\ data sources at the ORF, transcript, and gene levels. For full documentation, see the\ MetamORF documentation page.\
\ \\ The human sORFs in MetamORF were compiled from seven primary data sources and 46 individual\ ribosome profiling datasets from\ sORFs.org.\ The primary sources are:\
\ \| Source | \Description | \Reference | \
|---|---|---|
| Erhard et al. 2018 | \Union of ORFs detected by PRICE, RP-BP, ORF-RATER, or annotated in Ensembl v75 | \Nat Methods 2018 | \
| Johnstone et al. 2016 | \Location and translation data for analyzed transcripts and ORFs | \EMBO J 2016 | \
| Laumont et al. 2016 | \Cryptic MAPs (minor ORF-encoded peptides) with genomic and proteomic features | \Nat Commun 2016 | \
| Mackowiak et al. 2015 | \Systematic identification of sORFs across vertebrate genomes | \Genome Biol 2015 | \
| Samandi et al. 2017 | \Alternative protein predictions based on RefSeq GRCh38 | \eLife 2017 | \
| sORFs.org | \Repository of sORFs from 46 individual ribosome profiling experiments | \Olexiouk et al., Nucleic Acids Res 2018 | \
\ ORFs were identified using three main approaches: bioinformatic predictions, ribosome profiling\ experiments, and mass spectrometry (proteomics, peptidomics, and proteogenomics).\
\ \\ MetamORF classifies ORFs by their position relative to annotated coding sequences:\
\\ ORFs are also classified by the biotype of their host RNA: intergenic, ncRNA, pseudogene,\ NMD (nonsense-mediated decay), or readthrough transcripts.\
\ \\ Items are displayed in bigGenePred format. Each item is labeled with its MetamORF ORF\ ID. Color reflects the categorical Kozak consensus strength:\
\\
Strong – A/G at position −3 and G at position +4
\
Moderate – only one of those positions matches
\
Weak – neither position matches
\
non-ATG – near-cognate start codon; the Kozak rule does not apply
\
no context – chromosome edge or context unavailable\
\ Mouseover shows the ORF ID, ORF annotation, start codon, Kozak strength and TE,\ host transcripts, and the cell types where the ORF was reported.\
\ \Available filters: start codon, Kozak strength, Kozak TE.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. The data can be accessed from\ scripts through our API; the track name is\ "metamorf".\
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/ncOrfs/metamorf/MetamORF.kozak.bb -chrom=chr21 -start=0 -end=100000000 stdout\ \\ The original data and additional downloads are available from the\ MetamORF website.\ Source code is available on\ GitHub.\
\ \\ The MetamORF BED 12 data was obtained from the MetamORF\ track hub\ and converted to bigBed format at UCSC. Coordinates are on the GRCh38/hg38 assembly\ (based on Ensembl release 90).\
\ \\ Thanks to the MetamORF team at the TAGC (Theories and Approaches of Genomic Complexity)\ laboratory, Aix-Marseille University, for creating this resource and making it publicly\ available.\
\ \\ Erhard F, Halenius A, Zimmermann C, L'Hernault A, Kowalewski DJ, Weekes MP, Stevanovic S,\ Zimmer R, Dölken L.\ \ Improved Ribo-seq enables identification of cryptic translation events.\ Nat Methods. 2018 May;15(5):363-366.\ PMID: 29529017; PMC: PMC6152898\
\ \\ Johnstone TG, Bazzini AA, Giraldez AJ.\ \ Upstream ORFs are prevalent translational repressors in vertebrates.\ EMBO J. 2016 Apr 1;35(7):706-23.\ PMID: 26896445; PMC: PMC4818764\
\ \\ Laumont CM, Daouda T, Laverdure JP, Bonneil É, Caron-Lizotte O, Hardy MP, Granados DP, Durette C,\ Lemieux S, Thibault P et al.\ \ Global proteogenomic analysis of human MHC class I-associated peptides derived from non-canonical\ reading frames.\ Nat Commun. 2016 Jan 5;7:10238.\ PMID: 26728094; PMC: PMC4728431\
\ \\ Mackowiak SD, Zauber H, Bielow C, Thiel D, Kutz K, Calviello L, Mastrobuoni G, Rajewsky N, Kempa S,\ Selbach M et al.\ \ Extensive identification and analysis of conserved small ORFs in animals.\ Genome Biol. 2015 Sep 14;16:179.\ PMID: 26364619; PMC: PMC4568590\
\ \\ Olexiouk V, Van Criekinge W, Menschaert G.\ \ An update on sORFs.org: a repository of small ORFs identified by ribosome profiling.\ Nucleic Acids Res. 2018 Jan 4;46(D1):D497-D502.\ PMID: 29140531; PMC: PMC5753181\
\ \\ Samandi S, Roy AV, Delcourt V, Lucier JF, Gagnon J, Beaudoin MC, Vanderperre B, Breton MA, Motard J,\ Jacques JF et al.\ \ Deep transcriptome annotation enables the discovery and functional characterization of cryptic small\ proteins.\ Elife. 2017 Oct 30;6.\ PMID: 29083303; PMC: PMC5703645\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/ncOrfs/metamorf/MetamORF.kozak.bb\ filter.kozakTE -1:1.5\ filterByRange.kozakTE on\ filterLimits.kozakTE -1:1.5\ filterType.kozakStrength multipleListOr\ filterType.startCodon multipleListOr\ filterValues.kozakStrength Strong,Moderate,Weak,non-ATG,None\ filterValues.startCodon ATG,CTG,GTG,TTG,ACG,other,none\ itemRgb on\ longLabel ncORFs: MetamORF - meta-database of non-canonical ORFs\ mouseOver $name ($type)\ This track show alignments of human mRNAs from the\ Mammalian Gene Collection\ (MGC) having full-length open reading frames (ORFs) to the genome.\ The goal of the Mammalian Gene Collection is to provide researchers with\ unrestricted access to sequence-validated full-length protein-coding cDNA\ clones for human, mouse, rat, xenopus, and zerbrafish genes.\
\ \\ The track follows the display conventions for\ gene prediction\ tracks.\
\ \\ An optional codon coloring feature is available for quick\ validation and comparison of gene predictions.\ To display codon colors, select the genomic codons option from the\ Color track by codons pull-down menu. For more information\ about this feature, go to the\ \ Coloring Gene Predictions and Annotations by Codon page.\
\ \\ GenBank human MGC mRNAs identified as having full-length ORFs\ were aligned against the genome using blat. When a single mRNA\ aligned in multiple places, the alignment having the highest base identity was\ found. Only alignments having a base identity level within 1% of\ the best and at least 95% base identity with the genomic sequence\ were kept.\
\ \\ The human MGC full-length mRNA track was produced at UCSC from\ mRNA sequence data submitted to\ \ GenBank by the Mammalian Gene Collection project.\
\ \\ Mammalian Gene Collection project\ references.\
\ \\ Kent WJ.\ \ BLAT--the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ genes 1 baseColorDefault diffCodons\ baseColorUseCds genbank\ baseColorUseSequence genbank\ color 0,100,0\ group genes\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Mammalian Gene Collection Full ORF mRNAs\ parent mgcOrfeomeMrna\ shortLabel MGC Genes\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ track mgcFullMrna\ type psl\ visibility pack\ mgcOrfeomeMrna MGC/ORFeome Genes MGC/ORFeome Full ORF mRNA Clones 0 100 0 0 0 127 127 127 0 0 0\ These tracks show alignments of human mRNAs from the\ Mammalian Gene Collection\ (MGC) and ORFeome Collaboration having full-length open reading frames (ORFs) to the genome.\ The goal of the Mammalian Gene Collection is to provide researchers with\ unrestricted access to sequence-validated full-length protein-coding cDNA\ clones for human, mouse, and rat genes. The ORFeome project extended MGC to\ provide additional human, mouse, and zebrafish clones.\
\ \\ The track follows the display conventions for\ gene prediction\ tracks.\
\ \\ An optional codon coloring feature is available for quick\ validation and comparison of gene predictions.\ To display codon colors, select the genomic codons option from the\ Color track by codons pull-down menu. For more information\ about this feature, go to the\ \ Coloring Gene Predictions and Annotations by Codon page.\
\ \\ GenBank human MGC mRNAs identified as having full-length ORFs\ were aligned against the genome using blat. When a single mRNA\ aligned in multiple places, the alignment having the highest base identity was\ found. Only alignments having a base identity level within 1% of\ the best and at least 95% base identity with the genomic sequence\ were kept.\
\ \\ The human MGC full-length mRNA track was produced at UCSC from\ mRNA sequence data submitted to\ \ GenBank by the Mammalian Gene Collection project.\
\ \\ Visit the ORFeome Collaboration\ members page for a list of credits and references.\
\ \\ Mammalian Gene Collection project\ references.\
\ \\ Kent WJ.\ \ BLAT--the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ genes 0 cartVersion 4\ group genes\ longLabel MGC/ORFeome Full ORF mRNA Clones\ shortLabel MGC/ORFeome Genes\ superTrack on\ track mgcOrfeomeMrna\ visibility hide\ gnomADPextMinorSalivaryGland Minor Salivary Gland bigWig 0 1 gnomAD pext Minor Salivary Gland 0 100 153 187 136 204 221 195 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/MinorSalivaryGland.bw\ color 153,187,136\ longLabel gnomAD pext Minor Salivary Gland\ parent gnomadPext off\ shortLabel Minor Salivary Gland\ track gnomADPextMinorSalivaryGland\ visibility hide\ miRnaAtlas miRNA Tissue Atlas bigBarChart Tissue-Specific microRNA Expression from Two Individuals 0 100 0 0 0 127 127 127 0 0 0\ The Human miRNA Tissue Atlas is a\ catalog of tissue-specific microRNA (miRNA) expression across 62 tissues. This track contains\ quantile normalized miRNA expression data sampled from two individuals and mapped to\ miRBase v21 coordinates. The track contains two subtracks, one\ for each individual sampled.
\ \\ The Tissue Specificity Index (TSI) is analogous to the "tau" value for mRNA expression,\ and is calculated as described in the\ \ associated publication. Values closer to 0 indicate miRNAs expressed in many or all tissues,\ while values closer to 1 indicate miRNAs expressed only in a specific tissue or tissues. To\ browse miRNAs by TSI value, please see the\ miRNA Tissue Atlas.
\ \\ This track is formatted as a barChart track,\ similar to the GTEx or the\ TCGA Cancer Expression tracks, where the\ heights of each bar indicate the expression value for the miRNA in a specific tissue. The tissues\ sampled are described in the table below:\
\| Bar Color | Sample 1 | Sample 2 |
| Adipocyte | Adipocyte | |
| Artery | Artery | |
| Colon | Colon | |
| Dura mater | Dura mater | |
| Kidney | Kidney | |
| Liver | Liver | |
| Lung | Lung | |
| Muscle | Muscle | |
| Myocardium | Myocardium | |
| Skin | Skin | |
| Spleen | Spleen | |
| Stomach | Stomach | |
| Testis | Testis | |
| Thyroid | Thyroid | |
| Small intestine | ||
| Bone | ||
| Gallbladder | ||
| Fascia | ||
| Bladder | ||
| Epididymis | ||
| Tunica albuginea | ||
| Nervus intercostalis | ||
| Arachnoid mater | ||
| Brain | ||
| Small intestine duodenum | ||
| Small intestine jejunum | ||
| Pancreas | ||
| Kidney glandula suprarenalis | ||
| Kidney cortex renalis | ||
| Esophagus | ||
| Prostate | ||
| Bone marrow | ||
| Vein | ||
| Lymph node | ||
| Nerve not specified | ||
| Pleura | ||
| Pituitary gland | ||
| Spinal cord | ||
| Thalamus | ||
| Brain white matter | ||
| Nucleus caudatus | ||
| Kidney medulla renalis | ||
| Brain gray_matter | ||
| Cerebral cortex temporal | ||
| Cerebral cortex frontal | ||
| Cerebral cortex occipital | ||
| Cerebellum |
\ The 14 shared tissues sampled across both individuals are presented in the same order for easier comparison.\
\ \\ The underlying expression matrix and TSI values can be obtained from the\ miRNA tissue atlas website, in the\ data_matrix_quantile.txt and tsi_quantile.csv files.\
\ \\ Ludwig N, Leidinger P, Becker K, Backes C, Fehlmann T, Pallasch C, Rheinheimer S, Meder B,\ Stähler C, Meese E et al.\ \ Distribution of miRNA expression across human tissues.\ Nucleic Acids Res. 2016 May 5;44(8):3865-77.\ PMID: 26921406; PMC: PMC4856985\
\ expression 1 barChartLabel Tissue\ compositeTrack on\ configurable off\ group expression\ longLabel Tissue-Specific microRNA Expression from Two Individuals\ maxLimit 52000\ shortLabel miRNA Tissue Atlas\ subGroup1 view View a_A=Sample1 b_B=Sample2\ track miRnaAtlas\ type bigBarChart\ miRnaAtlasSample1 miRNA Tissue Atlas bigBarChart Tissue-Specific microRNA Expression from Two Individuals 3 100 0 0 0 127 127 127 0 0 0 expression 1 configurable on\ longLabel Tissue-Specific microRNA Expression from Two Individuals\ parent miRnaAtlas\ shortLabel miRNA Tissue Atlas\ track miRnaAtlasSample1\ type bigBarChart\ view a_A\ visibility pack\ miRnaAtlasSample2 miRNA Tissue Atlas bigBarChart Tissue-Specific microRNA Expression from Two Individuals 3 100 0 0 0 127 127 127 0 0 0 expression 1 configurable on\ longLabel Tissue-Specific microRNA Expression from Two Individuals\ parent miRnaAtlas\ shortLabel miRNA Tissue Atlas\ track miRnaAtlasSample2\ type bigBarChart\ view b_B\ visibility pack\ mitoMap MITOMAP bigBed 9 + MITOMAP: A human mitochondrial genome database 0 100 0 0 0 127 127 127 0 0 2 chrM,chrMT,NOTE: MITOMAP data is available for\
\ chrM on hg38 and chrMT on hg19.
\
\ This track shows annotations from MITOMAP.\ MITOMAP is a database of human mitochondrial DNA (mtDNA) information containing\ a compilation of mtDNA variation. It allows users to look up human mitochondrial gene \ loci, search for public mitochondrial sequences, and browse or search for reported \ general population nucleotide variants as well as those reported in clinical disease.\
\\ The data in these tracks are automatically updated from MitoMap weekly.
\ \\
These data are separated into two tracks:
\
MITOMAP Control and Coding Variants
\
This data track contains variants, including mini insertions and deletions, in the\
complete mtDNA. The item colors correspond to the variant type:\
control region vs.\
coding region.
\
MITOMAP Disease Mutations
\
This data track contains disease-annotated mutations (variants) in the\
complete mtDNA. The item colors correspond to the variant type:\
coding/control vs.\
rRNA/tRNA.
\ For both tracks, item names correspond to the\ variant nucleotide change, and mousing over features displays all available\ metadata for MITOMAP. Linkouts to the specific MITOMAP datasets are available from\ the item description pages, however, you must input the variant on MITOMAP.
\ \ Abbreviations and Definitions\Nucleotide changes are indicated as L-strand substitutions.
\Variant Classification:
\Disease Associations:
\Mutation Terminology:
\Mutation Status Definitions:
\\ MITOMAP collected the sequences from GenBank, aligned them to the rCRS using BLASTn, and\ haplotyped them with Haplogrep via the Mitomaster web service.
\\ The data were originally downloaded from the \ MITOMAP resource. For the Control and Coding Variants track, the following\ datasets were combined:\
\ And for the Disease Mutations track, the following two were combined:\ \\ These tracks have since been updated to automatically fetch files from the MitoMap server.\ For all the details on how the data are processed and combined, see\ the \ MITOMAP makedoc.\
\ \\ All source data can be found on the \ MITOMAP site.\
\ The MITOMAP data on the UCSC Genome Browser can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated download and analysis, the genome annotation is stored at UCSC in bigBed\ files that can be downloaded from the respective file, e.g.\ MITOMAP Variants, on our download server.\ The data may also be explored interactively using our\ REST API.
\ \\
The file for this track may also be locally explored using our tools bigBedToBed\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigBedToBed -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/mitoMapVars.bb stdout
\ Thanks to Shiping Zhang and the entire MITOMAP resource\ for making these annotations available.
\ \Lott MT, Leipzig JN, Derbeneva O, Xie HM, Chalkia D, Sarmady M, Procaccio V, Wallace DC. mtDNA Variation and Analysis\ Using Mitomap and Mitomaster. Curr Protoc Bioinformatics. 2013Dec;44(123):1.23.1-26.\ PMID: 25489354; PMC: PMC4257604
\ phenDis 1 chromosomes chrM,chrMT\ compositeTrack on\ dataVersion /gbdb/$D/bbi/mitoMapVersion.txt\ group phenDis\ longLabel MITOMAP: A human mitochondrial genome database\ noScoreFilter on\ shortLabel MITOMAP\ track mitoMap\ type bigBed 9 +\ visibility hide\ models_view Models bigBed Capture long-seq long-read lncRNAs 3 100 0 0 0 127 127 127 0 0 0 rna 1 longLabel Capture long-seq long-read lncRNAs\ noScoreFilter on\ parent clsLongReadRnaTrack\ shortLabel Models\ track models_view\ type bigBed\ view models_view\ visibility pack\ per_expr_models_view Models bigBed Capture long-seq long-read lncRNAs 1 100 0 0 0 127 127 127 0 0 0 rna 1 longLabel Capture long-seq long-read lncRNAs\ noScoreFilter on\ parent clsLongReadRnaTrack on\ shortLabel Models\ track per_expr_models_view\ type bigBed\ view per_expr_models_view\ visibility dense\ mpra MPRAs Massively Parallel Reporter Assays 0 100 0 0 0 127 127 127 0 0 0\ Massively Parallel Reporter Assays (MPRAs) are high-throughput methods that\ measure the regulatory activity of thousands of candidate DNA sequences in\ parallel. Each fragment is cloned next to a reporter gene and tagged with a\ unique barcode; sequencing the resulting reporter RNA quantifies how strongly\ each fragment drives expression. Most assays place the candidate fragment\ upstream of the reporter to measure transcriptional activation; some place it\ in the 3' untranslated region instead, where the readout reflects post-\ transcriptional effects on mRNA stability, decay, or translation. When matched\ reference and mutated versions of a sequence are tested side-by-side, the\ effect of a genetic variant on regulatory activity can be measured directly.\
\ \\ This track collection brings together results from two MPRA databases, one for\ the complete sequence fragments and one for the impact of variants in selected\ fragments:\
\ \\ Note on cell lines: The cell line shown for each element or variant is\ the reporter cell line in which the sequence was assayed. Most rows test human\ DNA in human cells. Several studies used mouse cell lines (Neuro-2a, N2A,\ NIH/3T3, MIN6) as reporter systems for human regulatory sequences. One MPRA\ Base study (Mattioli et al., 2020) tested mouse orthologous sequences\ in mouse embryonic stem cells (mESC); those items retain hg38 coordinates,\ derived from the orthologous human position by liftOver.\
\ \\ See the individual subtrack documentation pages linked above for detailed information\ on how to download and intersect the annotations.\
\ \\ Thanks to Weijia Jin and colleagues at the University of Florida for\ MPRAVarDB,\ and to Varda Singhal and the\ Ahituv Lab\ at the University of California San Francisco for\ MPRA Base.\
\ \\ Jin W, Xia Y, Nizomov J, Liu Y, Li Z, Lu Q, Chen L.\ \ MPRAVarDB: an online database and web server for exploring regulatory effects of genetic variants.\ Bioinformatics. 2024 Oct 1;40(10).\ PMID: 39325859; PMC: PMC11464417\
\ \\ Zhao J, Baltoumas FA, Konnaris MA, Mouratidis I, Liu Z, Sims J, Agarwal V, Pavlopoulos GA,\ Georgakopoulos-Soares I, Ahituv N.\ \ MPRAbase: A Massively Parallel Reporter Assay Database.\ bioRxiv. 2023 Nov 22;.\ PMID: 38045264; PMC: PMC10690217\
\ \ regulation 0 group regulation\ longLabel Massively Parallel Reporter Assays\ pennantIcon New red ../goldenPath/newsarch.html#060226 "Released Jun. 2, 2026"\ shortLabel MPRAs\ superTrack on\ track mpra\ visibility hide\ bismapBigWig Multi-read mappability bigWig Single-read and multi-read mappability after bisulfite conversion 2 100 0 0 0 127 127 127 0 0 0 map 0 longLabel Single-read and multi-read mappability after bisulfite conversion\ parent bismap on\ shortLabel Multi-read mappability\ track bismapBigWig\ type bigWig\ view MR\ viewLimits 0:1\ visibility full\ consHprc90way Multiple Alignment bed 4 Multiple Alignment on 90 human genome assemblies 0 100 0 0 0 127 127 127 0 0 0\ This track shows multiple alignments of 90 human genomes generated by the Minigraph-Cactus\ pangenome pipeline, which creates pangenomes directly from whole-genome alignments. This method\ builds graphs containing all forms of genetic variation while allowing use of current mapping and\ genotyping tools.\
\ \\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ size of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ Note that excluding species from the pairwise display does not alter the\ the conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. The following\ conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment. The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\ \\ The MAF was obtained from the HPRC v1.0 minigraph-cactus HAL file (renamed\ to replace all "." characters in sample names with "#" using\ halRenameGenomes) using cactus v2.6.4 as follows.\
\ cactus-hal2maf ./js ./hprc-v1.0-mc-grch38.h\ al hprc-v1.0-mc-grch38.maf.gz --noAncestors --refGenome GRCh38\ --filterGapCausingDupes --chunkSize 100000 --batchCores 96 --batchCount 1\ 0 --noAncestors --batchParallelTaf 32 --batchSystem slurm --logFile\ hprc-v1.0-mc-grch38.maf.gz.log\ \ zcat hprc-v1.0-mc-grch38.maf.gz | mafDuplicateFilter -m - -k | bgzip >\ hprc-v1.0-mc-grch38-single-copy.maf.gz\ \ \
\ Thank you to Glenn Hickey for providing the HAL file from the HPRC project.\
\ \\ Liao WW, Asri M, Ebler J, Doerr D, Haukness M, Hickey G, Lu S, Lucas JK, Monlong J, Abel HJ et\ al.\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ DOI: 10.1038/s41586-023-05896-x; PMID: 37165242; PMC: PMC10172123\
\ \\ Hickey G, Monlong J, Ebler J, Novak AM, Eizenga JM, Gao Y, Human Pangenome Reference Consortium,\ Marschall T, Li H, Paten B.\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nat Biotechnol. 2023 May 10;.\ DOI: 10.1038/s41587-023-01793-w; PMID: 37165083; PMC: PMC10638906\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q, Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ DOI: 10.1038/s41586-020-2871-y; PMID: 33177663; PMC: PMC7673649\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ DOI: 10.1101/gr.123356.111;\ PMID: 21665927; PMC: PMC3166836\
\ hprc 1 compositeTrack on\ dragAndDrop subTracks\ group hprc\ html hprc90way\ longLabel Multiple Alignment on 90 human genome assemblies\ shortLabel Multiple Alignment\ subGroup1 view Views align=Multiz_Alignment\ track consHprc90way\ type bed 4\ visibility hide\ cons447wayViewalign Multiz 447-way bed 4 Zoonomia+Primates 447 - 447 mammals, including 233 primates, aligned with Cactus, for Kuderna et al. 2023 3 100 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel Zoonomia+Primates 447 - 447 mammals, including 233 primates, aligned with Cactus, for Kuderna et al. 2023\ parent cons447way\ shortLabel Multiz 447-way\ track cons447wayViewalign\ view align\ viewUi on\ visibility pack\ cons470wayViewalign Multiz 470-way bed 4 Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group, 3 100 0 0 0 127 127 127 0 0 0 compGeno 1 longLabel Hiller Lab 470 Mammals - 470 mammalian genomes aligned with Multiz by Michael Hiller's Group,\ parent cons470way\ shortLabel Multiz 470-way\ track cons470wayViewalign\ view align\ viewUi on\ visibility pack\ multiz30way Multiz Align wigMaf 0.0 1.0 Multiz Alignments of 30 mammals (27 primates) 3 100 0 10 100 0 90 10 0 0 0 compGeno 1 altColor 0,90,10\ color 0, 10, 100\ frames multiz30wayFrames\ group compGeno\ irows on\ itemFirstCharCase noChange\ longLabel Multiz Alignments of 30 mammals (27 primates)\ noInherit on\ parent cons30wayViewalign on\ priority 100\ sGroup_Primates panTro5 panPan2 gorGor5 ponAbe2 nomLeu3 nasLar1 rhiBie1 rhiRox1 colAng1 macFas5 rheMac8 papAnu3 macNem1 cerAty1 chlSab2 manLeu1 saiBol1 aotNan1 calJac3 cebCap1 tarSyr2 eulFla1 eulMac1 proCoq1 micMur3 otoGar3 mm10 canFam3 dasNov3\ shortLabel Multiz Align\ speciesCodonDefault hg38\ speciesGroups Primates\ subGroups view=align\ summary multiz30waySummary\ track multiz30way\ treeImage phylo/hg38_30way.png\ type wigMaf 0.0 1.0\ multiz100way Multiz Align wigMaf 0.0 1.0 Multiz Alignments of 100 Vertebrates 3 100 0 10 100 0 90 10 0 0 0 compGeno 1 altColor 0,90,10\ color 0, 10, 100\ defaultMaf multiz100wayDefault\ frames multiz100wayFrames\ group compGeno\ irows on\ itemFirstCharCase noChange\ longLabel Multiz Alignments of 100 Vertebrates\ noInherit on\ parent cons100wayViewalign on\ priority 100\ sGroup_Afrotheria loxAfr3 eleEdw1 triMan1 chrAsi1 echTel2 oryAfe1\ sGroup_Birds falChe1 falPer1 ficAlb2 zonAlb1 geoFor1 taeGut2 pseHum1 melUnd1 amaVit1 araMac1 colLiv1 anaPla1 galGal4 melGal1\ sGroup_Euarchontoglires tupChi1 speTri2 jacJac1 micOch1 criGri1 mesAur1 mm10 rn6 hetGla2 cavPor3 chiLan1 octDeg1 oryCun2 ochPri3\ sGroup_Fish tetNig2 fr3 takFla1 oreNil2 neoBri1 hapBur1 mayZeb1 punNye1 oryLat2 xipMac1 gasAcu1 gadMor1 danRer10 astMex1 lepOcu1 petMar2\ sGroup_Laurasiatheria susScr3 vicPac2 camFer1 turTru2 orcOrc1 panHod1 bosTau8 oviAri3 capHir1 equCab2 cerSim1 felCat8 canFam3 musFur1 ailMel1 odoRosDiv1 lepWed1 pteAle1 pteVam1 myoDav1 myoLuc2 eptFus1 eriEur2 sorAra2 conCri1\ sGroup_Mammal dasNov3 monDom5 sarHar1 macEug2 ornAna1\ sGroup_Primate panTro4 gorGor3 ponAbe2 nomLeu3 rheMac3 macFas5 papAnu2 chlSab2 calJac3 saiBol1 otoGar3\ sGroup_Sarcopterygii allMis1 cheMyd1 chrPic2 pelSin1 apaSpi1 anoCar2 xenTro7 latCha1\ shortLabel Multiz Align\ speciesCodonDefault hg38\ speciesDefaultOff speTri2 micOch1 criGri1 mesAur1 rn6 hetGla2 cavPor3 chiLan1 octDeg1 oryCun2 ochPri3 susScr3 vicPac2 camFer1 turTru2 orcOrc1 panHod1 bosTau8 oviAri3 capHir1 equCab2 cerSim1 felCat8 musFur1 ailMel1 odoRosDiv1 lepWed1 pteAle1 pteVam1 myoDav1 myoLuc2 eptFus1 eriEur2 sorAra2 conCri1 eleEdw1 triMan1 chrAsi1 echTel2 oryAfe1 dasNov3 sarHar1 macEug2 ornAna1 falChe1 falPer1 ficAlb2 zonAlb1 geoFor1 taeGut2 pseHum1 melUnd1 amaVit1 araMac1 colLiv1 anaPla1 melGal1 allMis1 cheMyd1 chrPic2 pelSin1 apaSpi1 anoCar2 latCha1 tetNig2 fr3 takFla1 oreNil2 neoBri1 hapBur1 mayZeb1 punNye1 oryLat2 xipMac1 gasAcu1 gadMor1 astMex1 lepOcu1 calJac3 chlSab2 gorGor3 jacJac1 macFas5 monDom5 nomLeu3 otoGar3 panTro4 papAnu2 ponAbe2 saiBol1 tupChi1 petMar2\ speciesDefaultOn hg38 canFam3 loxAfr3 xenTro7 danRer10 galGal4 rheMac3 mm10\ speciesGroups Primate Euarchontoglires Laurasiatheria Afrotheria Mammal Birds Sarcopterygii Fish\ subGroups view=align\ summary multiz100waySummary\ track multiz100way\ treeImage phylo/hg38_100way.png\ type wigMaf 0.0 1.0\ muscleDeMicheliCellType Muscle Cells bigBarChart Muscle RNA binned by cell type from De Micheli et al 2020 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=muscle-cell-atlas&gene=$$\ This track displays data from A\ reference single-cell transcriptomic atlas of human skeletal muscle tissue\ reveals bifurcated muscle stem cell populations. Muscle tissue was\ analyzed using single-cell RNA-sequencing (scRNA-seq) and subsequent clustering\ distinguished 16 muscle-resident cell types based on their identified marker\ genes found in De Micheli et al., 2020. Muscle samples were from\ surgically discarded tissue taken from a wide variety of anatomical sites.
\ \\ This track collection contains two bar chart tracks of RNA expression in the\ human muscle where cells are grouped by cell type \ (Muscle Cells) or biosample\ (Muscle Sample). \ The default track displayed is \ Muscle Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| stem cell | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Muscle Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well. Note that the \ Muscle Sample subtrack is colored based on \ colors provided from Figure 1 from De Micheli et al., 2020.
\ \ \ \\ Muscle tissue cell type populations.\ \
\
\
\
\
De Micheli et al. Skelet\
Muscle. 2020. / CC BY 4.0\
\
\
\ Muscle samples were taken from 10 healthy donors of ages ranging from 41-81\ years old from different sections of the face (F), trunk (T), and leg (L).\ Excessive fat and connective tissue were removed from the muscle samples prior\ to enzymatic dissociation. Next, libraries were prepared using the 10x Genomics\ 3' v2 or v3 library kit and sequenced on the Illumina NextSeq 500. This\ resulted in libraries with 200-250 million reads which were processed using Cell\ Ranger version 3.1. In total, over 22,000 RNA transcriptomic profiles were\ generated from all of the samples after quality control filtering. The single\ cell transcriptomes from all 10 datasets were integrated using a scRNA-seq\ integration method called Scanorama as described in the reference below.\ \
The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Andrea De Micheli of the Cosgrove Laboratory at Cornell University\ and to the many authors who worked on producing and publishing this data set. The\ data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed Luis Nassar. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ De Micheli AJ, Spector JA, Elemento O, Cosgrove BD.\ \ A reference single-cell transcriptomic atlas of human skeletal muscle tissue reveals bifurcated\ muscle stem cell populations.\ Skelet Muscle. 2020 Jul 6;10(1):19.\ PMID: 32624006; PMC: PMC7336639
\ singleCell 1 barChartBars skeletal_muscle_cell_ACTA1+ smooth_muscle_cell_ACTA2+_MYH11+_MYL9+ adipocyte_APOD+_CFD+_PLAC9+ macrophage_C1QA+_CD74+ platelet_CD36+_VWF+ endothelial_cell_CLDN5+_PECAM1+ fibroblast_COL1A1+ fibroblast_DCN+_GSN+_MYOC+ fibroblast_FBN1+_MFAP5+_CD55+ erythroblast_HBA1+ endothelial_cell_HBA1+ B/T/NK_cell_IL7R+_PTPRC+_NKG7+ muscle_stem_cell_PAX7+_DLK1+_(MuSC1) muscle_stem_cell_PAX7-_MYF5+_(MuSC2) pericyte_RGS5+_MYL9+ macrophage_(inflammatory)_S100A9+_LYZ+\ barChartColors #d55acd #bb1b98 #fd8738 #da2f08 #b6513e #11b606 #b65928 #b35024 #b25023 #cf8b7e #419916 #fc344a #d33e3f #98672c #1dad0c #dc2c04\ barChartLabel Cell type\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/muscleDeMicheli/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/muscleDeMicheli/cell_type.bb\ defaultLabelFields name\ html muscleDeMicheli\ labelFields name,name2\ longLabel Muscle RNA binned by cell type from De Micheli et al 2020\ parent muscleDeMicheli\ shortLabel Muscle Cells\ track muscleDeMicheliCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=muscle-cell-atlas&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ muscleDeMicheli Muscle De Micheli Muscle single cell data from De Micheli et al 2020 0 100 0 0 0 127 127 127 0 0 0\ This track displays data from A\ reference single-cell transcriptomic atlas of human skeletal muscle tissue\ reveals bifurcated muscle stem cell populations. Muscle tissue was\ analyzed using single-cell RNA-sequencing (scRNA-seq) and subsequent clustering\ distinguished 16 muscle-resident cell types based on their identified marker\ genes found in De Micheli et al., 2020. Muscle samples were from\ surgically discarded tissue taken from a wide variety of anatomical sites.
\ \\ This track collection contains two bar chart tracks of RNA expression in the\ human muscle where cells are grouped by cell type \ (Muscle Cells) or biosample\ (Muscle Sample). \ The default track displayed is \ Muscle Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| stem cell | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Muscle Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well. Note that the \ Muscle Sample subtrack is colored based on \ colors provided from Figure 1 from De Micheli et al., 2020.
\ \ \ \ \\ Muscle samples were taken from 10 healthy donors of ages ranging from 41-81\ years old from different sections of the face (F), trunk (T), and leg (L).\ Excessive fat and connective tissue were removed from the muscle samples prior\ to enzymatic dissociation. Next, libraries were prepared using the 10x Genomics\ 3' v2 or v3 library kit and sequenced on the Illumina NextSeq 500. This\ resulted in libraries with 200-250 million reads which were processed using Cell\ Ranger version 3.1. In total, over 22,000 RNA transcriptomic profiles were\ generated from all of the samples after quality control filtering. The single\ cell transcriptomes from all 10 datasets were integrated using a scRNA-seq\ integration method called Scanorama as described in the reference below.\ \
The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Andrea De Micheli of the Cosgrove Laboratory at Cornell University\ and to the many authors who worked on producing and publishing this data set. The\ data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed Luis Nassar. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ De Micheli AJ, Spector JA, Elemento O, Cosgrove BD.\ \ A reference single-cell transcriptomic atlas of human skeletal muscle tissue reveals bifurcated\ muscle stem cell populations.\ Skelet Muscle. 2020 Jul 6;10(1):19.\ PMID: 32624006; PMC: PMC7336639
\ singleCell 0 group singleCell\ longLabel Muscle single cell data from De Micheli et al 2020\ shortLabel Muscle De Micheli\ superTrack on\ track muscleDeMicheli\ visibility hide\ muscleDeMicheliSample Muscle Sample bigBarChart Muscle RNA binned by biosample from De Micheli et al 2020 0 100 0 0 0 127 127 127 0 0 0 http://cells.ucsc.edu/?ds=muscle-cell-atlas&gene=$$\ This track displays data from A\ reference single-cell transcriptomic atlas of human skeletal muscle tissue\ reveals bifurcated muscle stem cell populations. Muscle tissue was\ analyzed using single-cell RNA-sequencing (scRNA-seq) and subsequent clustering\ distinguished 16 muscle-resident cell types based on their identified marker\ genes found in De Micheli et al., 2020. Muscle samples were from\ surgically discarded tissue taken from a wide variety of anatomical sites.
\ \\ This track collection contains two bar chart tracks of RNA expression in the\ human muscle where cells are grouped by cell type \ (Muscle Cells) or biosample\ (Muscle Sample). \ The default track displayed is \ Muscle Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| stem cell | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the \ Muscle Cells subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well. Note that the \ Muscle Sample subtrack is colored based on \ colors provided from Figure 1 from De Micheli et al., 2020.
\ \ \ \\ Details on sex, age, anatomical site, and single-cell transcriptomes after\ quality control (QC) filtering from 10 donors. Colors represent areas from\ which samples were taken from.
\ \\
\
\
\
De Micheli et al. Skelet\
Muscle. 2020. / CC BY 4.0
\ Cell type proportions across the 10 donors and grouped by leg (donors 02, 07,\ 08), trunk (donors 01, 05, 06, 09, 10), and face (donors 03, 04).
\ \\
\
\
\
De Micheli et al. Skelet\
Muscle. 2020. / CC BY 4.0
\ Muscle samples were taken from 10 healthy donors of ages ranging from 41-81\ years old from different sections of the face (F), trunk (T), and leg (L).\ Excessive fat and connective tissue were removed from the muscle samples prior\ to enzymatic dissociation. Next, libraries were prepared using the 10x Genomics\ 3' v2 or v3 library kit and sequenced on the Illumina NextSeq 500. This\ resulted in libraries with 200-250 million reads which were processed using Cell\ Ranger version 3.1. In total, over 22,000 RNA transcriptomic profiles were\ generated from all of the samples after quality control filtering. The single\ cell transcriptomes from all 10 datasets were integrated using a scRNA-seq\ integration method called Scanorama as described in the reference below.\ \
The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Andrea De Micheli of the Cosgrove Laboratory at Cornell University\ and to the many authors who worked on producing and publishing this data set. The\ data were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick \ then reviewed Luis Nassar. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ De Micheli AJ, Spector JA, Elemento O, Cosgrove BD.\ \ A reference single-cell transcriptomic atlas of human skeletal muscle tissue reveals bifurcated\ muscle stem cell populations.\ Skelet Muscle. 2020 Jul 6;10(1):19.\ PMID: 32624006; PMC: PMC7336639
\ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/muscleDeMicheli/sample.colors\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/muscleDeMicheli/sample.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/muscleDeMicheli/sample.bb\ defaultLabelFields name\ html muscleDeMicheli\ labelFields name,name2\ longLabel Muscle RNA binned by biosample from De Micheli et al 2020\ parent muscleDeMicheli\ shortLabel Muscle Sample\ track muscleDeMicheliSample\ transformFunc NONE\ type bigBarChart\ url http://cells.ucsc.edu/?ds=muscle-cell-atlas&gene=$$\ urlLabel UCSC Cell Browser:\ visibility hide\ gnomADPextMuscle_Skeletal Muscle-Skeletal bigWig 0 1 gnomAD pext Muscle-Skeletal 0 100 170 170 255 212 212 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Muscle_Skeletal.bw\ color 170,170,255\ longLabel gnomAD pext Muscle-Skeletal\ parent gnomadPext off\ shortLabel Muscle-Skeletal\ track gnomADPextMuscle_Skeletal\ visibility hide\ mutScore MutScore bigWig MutScore: Variant clustering in 3D protein structures 2 100 50 80 200 152 167 227 0 0 0\ The "Prediction Scores" container track contains subtracks showing the results of variant impact prediction\ scores. Usually these are prediction algorithms that use protein features, conservation, nucleotide composition and similar\ signals to determine if a genome variant is pathogenic or not.
\ \BayesDel is a deleteriousness meta-score for coding and \ non-coding variants, single nucleotide\ variants, and small insertion/deletions. The range of the score is from -1.29334 to 0.75731.\ The higher the score, the more likely the variant is pathogenic.
\\ MaxAF stands for maximum allele frequency. The old ACMG (American College of Medical Genetics and\ Genomics) rules utilize allele frequency to classify variants, so the "BayesDel without MaxAF"\ tracks were created to avoid double-dipping. However, new ACMG rules will not include allele\ frequency, so it is okay to use the "BayesDel with MaxAF" for variant classification in the future.\ For gene discovery research, it is better to use BayesDel with MaxAF.
\\ For gene discovery research, a universal cutoff value (0.0692655 with MaxAF, -0.0570105 without\ MaxAF) was obtained by maximizing sensitivity and specificity in classifying ClinVar variants;\ Version 1 (build date 2017-08-24).
\\ For clinical variant classification, Bayesdel thresholds have been calculated for a variant to\ reach various levels of evidence; please refer to Pejaver et al. 2022 for general application\ of these scores in clinical applications.\
\ \\ Interpretation: The authors define that at an M-CAP score > 0.025, 5% of \ pathogenic variants are misclassified as benign. 0.025 is the recommended cutoff.\
\ \\ The Mendelian Clinically Applicable Pathogenicity (M-CAP)\ score (Jagadeesh et al, Nat Genetics 2016) is a\ pathogenicity likelihood score that aims to misclassify no more than 5% of\ pathogenic variants while aggressively reducing the list of variants of\ uncertain significance. Much like allele frequency, M-CAP is readily\ interpreted; if it classifies a variant as benign, then that variant can be\ trusted to be benign with high confidence.
\ \\ At an M-CAP score > 0.025, 5% of pathogenic variants are misclassified as benign.\ The score varies from 0.0 - 1.0, following a geometric distribution with a mean of 0.09.\
\ \\ Interpretation: The authors defined the thresholds <0.140 for a variant\ to be benign, and > 0.730 for pathogenic with 95% confidence.
\\ The within-gene clustering of pathogenic and benign DNA changes is an important\ feature of the human exome.\ MutScore\ score (Quinodoz, AJHG 2022) integrates qualitative features of\ DNA substitutions with new additional information derived from \ positional clustering. Variants of unknown significance that are scored\ as benign by other algorithms but located close to known pathogenic variants\ should be weighted more pathogenic by MutScore. The score ranges from 0.0-1.0, resembles\ a negative binomial distribution with a maximum ~0.05, depending on the nucleotide.\ MutScore was seen to outperform other scores by papers Porretta et al and Brock et al.\
\ \\ Interpretation: Scores range from 0 to 1, with higher values indicating greater\ predicted pathogenicity. The authors suggest a clinical threshold of 0.821 for distinguishing\ pathogenic from benign missense variants. 75% of all possible missense variants are classified\ as benign, 25% as pathogenic.\
\\ PrimateAI-3D\ (Gao et al, Science 2023) is a semi-supervised 3D convolutional neural network trained on\ 4.5 million benign missense variants from 233 primate species and common human variants.\ It operates on voxelized protein structures at 2 Å resolution (from AlphaFold or\ homology models) combined with multiple sequence alignments from 592 species. The track\ contains pre-computed scores for all 70.7 million possible single nucleotide missense\ variants.\ Pathogenic variants are shown in red,\ benign in blue.\ Items can be filtered by prediction and by percentile score.\
\ \\ Interpretation: Scores range from -1 to 1. Positive scores indicate predicted\ disruption of promoter function, negative scores indicate the variant is tolerated.\
\\ PromoterAI\ predicts the impact of single nucleotide variants in gene\ promoter regions, scoring all possible substitutions within 500 bp of annotated\ transcription start sites. The track contains four bigWig subtracks (one per alternate\ allele) covering 39.5 million positions, plus a bigBed track for the 3.8% of positions\ where overlapping transcripts produce different scores.\
\ \\ Interpretation: Scores range from 0 to 1, with higher values indicating greater\ predicted likelihood of pathogenicity. The authors recommend a threshold of ≥ 0.5 to\ flag variants as likely disease-relevant.\
\\ ClinPred\ (Alirezaie et al, AJHG 2018) is a machine-learning predictor for nonsynonymous\ (missense) single-nucleotide variants. It combines existing pathogenicity scores\ with population allele frequency from gnomAD, and was trained on confidently\ annotated disease-causing and benign variants from ClinVar. The track contains\ four bigWig subtracks (one per alternate allele) with pre-computed scores for\ all possible human missense variants in the exome.\ Pathogenic variants are shown in red,\ benign in blue.\
\ \\ Interpretation: EVE scores range from 0 (benign) to 1 (pathogenic) and are\ normalized within each protein, so they are not directly comparable across proteins. A\ Class25 label assigns each variant to benign, uncertain, or pathogenic using a 25%\ uncertainty threshold.\
\\ EVE\ (Frazer et al, Nature 2021) is a deep generative model (a Bayesian variational\ autoencoder) trained per protein on evolutionary sequence alignments, without using\ clinical labels. The track shows scores for all possible missense substitutions in\ 2,949 disease-associated proteins as a heatmap (rows = amino acids, columns = protein\ positions), colored from benign (blue) through\ uncertain (white) to pathogenic (red).\
\ \\ Interpretation: popEVE scores are a continuous, proteome-wide measure of\ deleteriousness and, unlike most missense scores, are calibrated to be comparable across\ genes; lower (more negative) scores are more deleterious. The authors define a\ high-confidence severe threshold at −5.056 and a moderate threshold at −4.617.\
\\ popEVE\ (Orenbuch et al, Nature Genetics 2025) builds on EVE and the ESM-1v protein language model,\ calibrating their scores against human population variation (UK Biobank) with a Gaussian\ process to place variants across the whole proteome on a single scale. The track shows\ scores for all single-nucleotide-reachable missense substitutions across roughly 18,000\ proteins as a heatmap, colored on a global gradient from\ deleterious (red) to\ tolerated (blue).\
\ \There are eight subtracks for the BayesDel track: four include pre-computed MaxAF-integrated BayesDel\ scores for missense variants, one for each base. The other four are of the same format, but scores\ are not MaxAF-integrated.
\ \For SNVs, at each genome position, there are three values per position, one for every possible\ nucleotide mutation. The fourth value, "no mutation", representing the reference allele,\ (e.g. A to A) is always set to zero.
\ \Note: There are cases in which a genomic position will have one value missing.\
\ \When using this track, zoom in until you can see every base pair at the top of the display.\ Otherwise, there are several nucleotides per pixel under your mouse cursor and instead of an actual\ score, the tooltip text will show the average score of all nucleotides under the cursor. This is\ indicated by the prefix "~" in the mouseover.\
\ \\
Details on suggested ranges for BayesDel can be found in Bergquist et al Genet Med 2025, Table 2:\
\
There are four subtracks: one for each nucleotide.
\ \There are four subtracks: one for each alternate nucleotide. Each shows the\ ClinPred score for variants from the reference base to that nucleotide. Reference\ and synonymous alternates are set to 0; positions with no exome coverage appear as\ gaps. The track is colored at each position by the recommended threshold\ (≥ 0.5 = pathogenic, < 0.5 = benign).
\ \A single bigBed track containing all possible missense variants. Items are\ colored by prediction (red = pathogenic, blue = benign) and can be filtered by\ prediction or percentile score. See the per-track description page for details.
\ \Four bigWig subtracks (one per alternate nucleotide) covering positions within\ 500 bp of annotated transcription start sites, plus a bigBed track for\ positions where overlapping transcripts produce different scores.
\ \Each is a single bigBed track displayed as a heatmap: one column per amino acid\ position (placed at the codon's genomic coordinate) and one row per amino acid. Hover\ over a cell to see the substitution and its score. EVE is colored per protein from blue\ (benign) to red (pathogenic); popEVE uses a single global gradient (red = deleterious,\ blue = tolerated) so that cells are comparable across genes. These tracks are best viewed\ zoomed in to a single gene or exon.
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server, there is one subdirectory per score.\ The files for this track are called usually called by their alternate allele, e.g. mcapA.bw and mutScoreA.bw. Individual\ regions or the whole genome annotation can be obtained using our tool bigWigToBedGraph\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool\ can also be used to obtain only features within a given range, e.g. \ bigWigToBedGraph http://hgdownload.soe.ucsc.edu/gbdb/hg19/mcap/mcapA.bw -chrom=chr21 -start=0 -end=100000000 stdout
\ \ \The original BayesDel files are available at the\ BayesDel website.\
The other algorithms also have their own download formats, on the\ M-CAP website and the MutScore Website.\ \
BayesDel data was converted from the files provided on the\ BayesDel_170824 Database.\ The number 170824 is the date (2017-08-24) the scores were created. Both sets of BayesDel scores are\ available in this database, one integrated MaxAF (named BayesDel_170824_addAF) and one without\ (named BayesDel_170824_noAF). Data conversion was performed using\ \ custom Python scripts.\
\ \M-CAP data was converted using a custom Python script and converted to\ bigWig, as documented in the our makeDoc\ text file. MutScore was already available in bigWig format to download.
\ \Thanks to the BayesDel, MutScore, M-CAP, ClinPred, PrimateAI-3D and PromoterAI teams for\ providing precomputed data, and to Tiana Pereira, Christopher Lee, Gerardo Perez, and Anna\ Benet-Pages of the Genome Browser team.
\ \\ Alirezaie N, Kernohan KD, Hartley T, Majewski J, Hocking TD.\ \ ClinPred: Prediction Tool to Identify Disease-Relevant Nonsynonymous Single-Nucleotide Variants.\ Am J Hum Genet. 2018 Oct 4;103(4):474-483.\ PMID: 30220433; PMC: PMC6174354\
\ \\ Bergquist T, Stenton SL, Nadeau EAW, Byrne AB, Greenblatt MS, Harrison SM, Tavtigian SV,\ O'Donnell-Luria A, Biesecker LG, Radivojac P et al.\ \ Calibration of additional computational tools expands ClinGen recommendation options for variant\ classification with PP3/BP4 criteria.\ Genet Med. 2025 Mar 10;27(6):101402.\ PMID: 40084623\
\ \\ Feng BJ.\ \ PERCH: A Unified Framework for Disease Gene Prioritization.\ Hum Mutat. 2017 Mar;38(3):243-251.\ PMID: 27995669; PMC: PMC5299048\
\ \\ Gao H, Hamp T, Ede J, Schraiber JG, McRae J, Singer-Berk M, Yang Y, Dietrich ASD,\ Fiziev PP, Kuderna LFK et al.\ \ The landscape of tolerated genetic variation in humans and primates.\ Science. 2023 Jun 2;380(6648):eabn8197.\ PMID: 37262156; PMC: PMC10187174\
\ \\ Jagadeesh KA, Wenger AM, Berger MJ, Guturu H, Stenson PD, Cooper DN, Bernstein JA, Bejerano G.\ \ M-CAP eliminates a majority of variants of uncertain significance in clinical exomes at high\ sensitivity.\ Nat Genet. 2016 Dec;48(12):1581-1586.\ PMID: 27776117\
\ \\ Pejaver V, Byrne AB, Feng BJ, Pagel KA, Mooney SD, Karchin R, O'Donnell-Luria A, Harrison SM,\ Tavtigian SV, Greenblatt MS et al.\ \ Calibration of computational tools for missense variant pathogenicity classification and ClinGen\ recommendations for PP3/BP4 criteria.\ Am J Hum Genet. 2022 Dec 1;109(12):2163-2177.\ PMID: 36413997; PMC: PMC9748256\
\ \\ Quinodoz M, Peter VG, Cisarova K, Royer-Bertrand B, Stenson PD, Cooper DN, Unger S, Superti-Furga A,\ Rivolta C.\ \ Analysis of missense variants in the human genome reveals widespread gene-specific clustering and\ improves prediction of pathogenicity.\ Am J Hum Genet. 2022 Mar 3;109(3):457-470.\ PMID: 35120630; PMC: PMC8948164\
\ \\ Sundaram L, Gao H, Padigepati SR, McRae JF, Li Y, Kosmicki JA, Fritzilas N, Hakenberg J,\ Dutta A, Shon J et al.\ \ Predicting the clinical impact of human mutation with deep neural networks.\ Nat Genet. 2018 Aug;50(8):1161-1170.\ PMID: 30038395; PMC: PMC6237276\
\ \\ Tian Y, Pesaran T, Chamberlin A, Fenwick RB, Li S, Gau CL, Chao EC, Lu HM, Black MH, Qian D.\ \ REVEL and BayesDel outperform other in silico meta-predictors for clinical variant\ classification.\ Sci Rep. 2019 Sep 4;9(1):12752.\ PMID: 31484976; PMC: PMC6726608\
\ \ phenDis 0 color 50,80,200\ compositeTrack on\ group phenDis\ html predictionScoresSuper\ longLabel MutScore: Variant clustering in 3D protein structures\ parent predictionScoresSuper\ shortLabel MutScore\ track mutScore\ type bigWig\ visibility full\ gnomADPextNerve_Tibial Nerve-Tibial bigWig 0 1 gnomAD pext Nerve-Tibial 0 100 255 215 0 255 235 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Nerve_Tibial.bw\ color 255,215,0\ longLabel gnomAD pext Nerve-Tibial\ parent gnomadPext off\ shortLabel Nerve-Tibial\ track gnomADPextNerve_Tibial\ visibility hide\ hprcChainNetViewnet Nets bed 3 Human Genomes, Chain/Net pairwise alignments, as mapped by the HPRC project 1 100 0 0 0 255 255 0 0 0 0 hprc 1 longLabel Human Genomes, Chain/Net pairwise alignments, as mapped by the HPRC project\ parent hprcChainNet\ shortLabel Nets\ track hprcChainNetViewnet\ view net\ visibility dense\ nmd NMD Escape bed 4 NMD Escape: Predicted regions where premature termination codons escape NMD 0 100 0 0 0 127 127 127 0 0 0\ NMD is a cellular quality control mechanism that\ detects and degrades mRNAs containing premature termination codons (PTCs),\ preventing the accumulation of truncated, potentially harmful proteins.\ However, not all PTCs trigger NMD. PTCs in certain regions of a transcript are\ predicted to escape NMD, meaning the truncated mRNA may be translated into a\ protein with unpredictable functional consequences.\ The NMD Escape container includes several tracks that display putative regions where\ PTC variants are assumed to escape the NMD mechanism. These are typically located\ close to the first or last splice junction, within unusually long coding exons,\ or in transcripts without any junction.\
\ \\ Rule-based predictions of NMD escape regions, computed from transcript\ annotations. Three transcript sets are provided:\
\\ Click either of the links to the track details here or above to show the four rules\ that were used (50 bp, intronless, 100 bp, long exon >400 nt).\
\ \\ Machine-learning predictions of NMD efficiency from\ Lindeboom\ et al. 2016 (A and B models) and from Veiner et al.\ (NMDetective-AI, pre-print 2026). Positive scores indicate predicted NMD\ triggering; negative scores indicate predicted escape.\
\\ The ACMG guidelines say under PVS1:\
\\ \ (ii) One must also be cautious when interpreting truncating variants downstream of the most 3′ truncating variant established as pathogenic in the literature. This is especially true if the predicted stop codon occurs in the last exon or in the last 50 base pairs of the penultimate exon, such that nonsense-mediated decay would not be predicted, and there is a higher likelihood of an expressed protein.\ \
\ \\ The data underlying these tracks can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated analysis,\ the data may be queried from our\ REST API. Please refer to our\ mailing list archives for questions, or our\ Data Access FAQ for more\ information.\
\ \\ Thanks to Guido Neidhardt for suggesting this track at HUGO VEPTC 2025 and Andreas Lahner\ for feedback. Thanks to the Decipher Genome Browser team for introducing the idea of a\ track. Thanks to Rik Lindeboom for providing custom tracks.\
\ \\ Kurosaki T, Popp MW, Maquat LE.\ \ Quality and quantity control of gene expression by nonsense-mediated mRNA decay.\ Nat Rev Mol Cell Biol. 2019 Jul;20(7):406-420.\ PMID: 30992545; PMC: PMC6855384\
\ \\ Lindeboom RGH, Supek F, Lehner B.\ \ The rules and impact of nonsense-mediated mRNA decay in human cancers.\ Nat Genet. 2016 Oct;48(10):1112-8.\ PMID: 27618451; PMC: PMC5045715\
\ \\ Lindeboom RGH, Vermeulen M, Lehner B, Supek F.\ \ The impact of nonsense-mediated mRNA decay on genetic disease, gene editing and cancer\ immunotherapy.\ Nat Genet. 2019 Nov;51(11):1645-1651.\ PMID: 31659324; PMC: PMC6858879\
\ \\ Nagy E, Maquat LE.\ \ A rule for termination-codon position within intron-containing genes: when nonsense\ affects RNA abundance.\ Trends Biochem Sci. 1998 Jun;23(6):198-9.\ PMID: 9644970\
\ genes 1 group genes\ longLabel NMD Escape: Predicted regions where premature termination codons escape NMD\ pennantIcon New red ../goldenPath/newsarch.html#042226 "Released Apr. 22, 2026"\ shortLabel NMD Escape\ superTrack on\ track nmd\ type bed 4\ visibility hide\ ncOrfs Non-canonical ORFs Non-canonical Open Reading Frames 0 100 0 0 0 127 127 127 0 0 0\ The non-canonical ORFs supertrack contains tracks that display open reading frames (ORFs)\ found outside of annotated protein-coding sequences. While the human genome has approximately\ 20,000 annotated protein-coding genes, recent advances in ribosome profiling (Ribo-seq) and\ proteomics have revealed widespread translation of ORFs that do not correspond to known\ protein-coding genes. These non-canonical ORFs are found in regions previously considered\ non-coding, including 5' and 3' UTRs, long non-coding RNAs, pseudogenes, and alternative\ reading frames of known genes.\
\ \\ Several subtypes of non-canonical ORFs are commonly distinguished. Upstream ORFs (uORFs)\ are located in 5' UTRs and can regulate translation of the downstream main coding sequence;\ ribosomes that translate a uORF may fail to reinitiate at the main start codon, reducing\ protein output. Small ORFs (sORFs), generally defined as encoding fewer than 100 amino\ acids, have been systematically overlooked by gene annotation pipelines due to their short\ length, but many produce functional micropeptides involved in signaling, metabolism, and\ development. Other types include downstream ORFs (dORFs) in 3' UTRs,\ out-of-frame ORFs that overlap known coding sequences in an alternative reading frame,\ and ORFs in transcripts annotated as non-coding RNAs or pseudogenes.\
\This track collection imports various databases and annotates all ORFs with their Kozak strength, \ and colors the features by Kozak strength.
\ \Click any of the track names below to show their configuration/documentation page:
\ \| Track | \Description | \Items | \Genome Coverage | \
Exon Coverage | \
Start codon | \Kozak strength (ATG only) | \|||
|---|---|---|---|---|---|---|---|---|---|
| ATG | \non-ATG | \Strong | \Moderate | \Weak | \|||||
| UTRannotator uORFs | \Upstream ORFs in 5' UTRs from UTRannotator | \44,435 | \1.15% | \1.15% | \6,236 | \38,199 | \1,307 | \3,054 | \1,875 | \
| GENCODE ncORFs | \GENCODE non-canonical ORFs supported by Ribo-seq | \7,264 | \1.02% | \0.03% | \7,263 | \1 | \1,571 | \3,705 | \1,987 | \
| GENCODE ncORFs primary | \GENCODE non-canonical ORFs – primary set | \10,127 | \0.45% | \0.02% | \6,183 | \3,944 | \1,746 | \3,300 | \1,137 | \
| GENCODE ncORFs comprehensive | \GENCODE non-canonical ORFs – comprehensive set | \28,359 | \2.24% | \0.06% | \13,776 | \14,583 | \3,133 | \7,168 | \3,475 | \
| 5ULTRA uORFs | \uORFs in MANE Select transcripts from 5ULTRA (ATG only) | \22,567 | \2.44% | \0.09% | \22,567 | \0 | \3,768 | \10,862 | \7,472 | \
| nuORFdb | \Non-canonical ORFs from nuORFdb v1.2 | \229,251 | \22.14% | \0.83% | \51,080 | \178,171 | \10,905 | \25,539 | \14,636 | \
| MetamORF | \Meta-database of small ORFs (sORFs) | \664,558 | \33.53% | \1.19% | \147,490 | \517,068 | \33,481 | \74,267 | \39,742 | \
| OpenProt | \Alternative and reference proteins from OpenProt v2.2 | \921,170 | \49.85% | \3.36% | \906,942 | \14,228 | \202,199 | \446,288 | \258,455 | \
| OpenProt (MS>=2) | \OpenProt proteins with mass spectrometry evidence (≥2 peptides) | \377,916 | \40.29% | \1.85% | \367,257 | \10,659 | \106,148 | \181,817 | \79,292 | \
\ The three GENCODE ncORF tracks display non-canonical translated open reading frames\ identified from ribosome profiling (Ribo-seq) data and mapped to the GENCODE annotation by\ the GENCODE / TransCODE\ consortium.\
\\ See the\ GENCODE ncORFs Phase I subtrack page\ or the Phase II\ primary /\ comprehensive\ pages for download URLs, methods, and references.\
\ \\ 5ULTRA is a pipeline for\ prioritizing 5' UTR variants by their impact on protein translation. As part of the project,\ Chaldebas et al. compiled a reference set of 22,567 ATG-initiated uORFs from two databases\ (Ribo-uORF and uORFdb), mapped to MANE Select transcripts and classified into three functional\ types. See the 5ULTRA uORFs subtrack page\ for more details.\
\ \\ Created by the Whiffin lab,\ UTRannotator\ is a VEP plugin for annotating 5' UTR variants with respect to upstream open reading frames\ (uORFs). As part of the project, the authors compiled a curated reference set of uORFs in\ human 5' UTRs from\ sorfs.org, which contains ORFs supported\ by Ribo-Seq.\ See the UTRannotator uORFs subtrack page\ for more details. Data from sorfs.org is also part of the Metamorf track (see below).\ This track is useful if you have a prediction from the VEP plugin and want to see the context.\
\\ The UTRannotator source data is distributed as single-span features with no exon/intron\ structure, so the uORFs would appear as continuous blocks even across introns of their host\ transcripts. To recover the splicing structure we look up, for each uORF, a same-strand\ MANE Select / MANE Plus Clinical\ transcript whose coordinates overlap the uORF range. The host transcript's exons are\ clipped to the uORF range so that any MANE intron inside the overlap is preserved as an\ intron of the displayed bed12 record. A uORF that extends past either end of MANE keeps\ the MANE introns inside the overlap and gets a single bridging block for the orphan\ portion. If a uORF endpoint falls inside a MANE intron (i.e. UTRannotator originally used\ a transcript whose UTR exon boundaries differ from MANE's), we fall back to the full\ GENCODE comprehensive set and apply\ the same projection. If no donor in either pool can host the uORF, it stays single-block.\ The chosen donor transcript ID is recorded in the intronsSource field\ (or none if no host was found).\
\ \\ nuORFdb (novel\ unannotated ORF database) is a Broad Institute database of non-canonical open reading frames\ with evidence of translation from ribosome profiling (Ribo-seq). ORF types include uORFs,\ dORFs, out-of-frame ORFs, pseudogene ORFs, lincRNA ORFs, and others.\ See the nuORFdb subtrack page for more details.\ The nuORFdb database is a very consistent dataset, from a well-known paper.\
\ \\ MetamORF is a repository of\ small ORFs (sORFs) in the human genome, consolidated from several primary data sources and\ many individual ribosome profiling datasets. It integrates bioinformatic predictions,\ ribosome profiling experiments, and mass spectrometry studies into a unified format.\ See the MetamORF subtrack page for more details.\ Metamorf has many predictions, and not all may be relevant, but gives an example of\ a database with as many models as possible, and is the only complete archive of sorfs.org \ that we are aware of.\
\ \\ OpenProt is a comprehensive annotation\ of all possible protein-coding ORFs in eukaryotic genomes. It distinguishes RefProts (the\ known canonical proteins), Isoforms (alternative products of canonical genes), and AltProts\ (predicted from alternative reading frames in UTRs, frameshifted CDS overlaps, and\ non-coding RNAs). Each ORF is annotated with mass spectrometry and ribosome profiling\ evidence; a pre-filtered Mass-Spec-supported subset (≥2 unique peptides) is also available.\ See the OpenProt subtrack page for more details.\ OpenProt is widely known and has by far the most predictions, even more than Metamorf, which is\ why a subset exists with only the more reliable ORFs with Mass-Spec data.\
\ \\ Every ORF in every subtrack carries three additional annotation fields derived from the\ genomic sequence around its start codon:\
\-1 for non-ATG starts and rows with no lookup.\ Features in every subtrack are colored by the categorical kozakStrength\ field. The same legend applies to all subtracks:\
\\
Strong – A/G at position −3 and G at position +4
\
Moderate – only one of those two positions matches
\
Weak – neither position matches
\
non-ATG – near-cognate start codon; the Kozak rule does not apply
\
no context – chromosome edge or other case where the 11-base context could not be read\
\ Per-subtrack counts of ATG vs. non-ATG starts and the Strong / Moderate / Weak breakdown\ are shown in the table at the top of this page. The high non-ATG fraction in UTRannotator,\ MetamORF, and nuORFdb is inherent to those catalogs — they explicitly include\ non-canonical CTG/GTG/TTG starts. GENCODE Phase I restricted itself to ATG-only starts.\
\ \\ The 11-base Kozak context is fetched directly from the genome at the position of the start\ codon. For multi-exon ORFs with an intron immediately upstream of the start codon, the\ upstream bases of the context are genomic rather than the host transcript's true 5' UTR;\ in that uncommon case the computed Kozak value may be inaccurate.\
\ \\ The Kozak strength annotation and color coding are added by the script\ colorByKozak.py\ in the kent source tree\ (src/hg/makeDb/scripts/ncOrfs/),\ along with the per-track autoSql files, the cached Noderer 2014 TE table, and a\ helper script (addIntrons.py) that recovers exon/intron structure for the\ UTRannotator uORFs. The Kozak strength logic is a Python port of the corresponding\ R routines in the\ VuTR pipeline\ (Whiffin lab / Computational Rare-Disease Genomics, WHG Oxford); credit and thanks to\ the VuTR authors for the original implementation. Full build steps are recorded in\ the\ makedoc.\
\ \ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be\ accessed from scripts through our API. See the individual\ track pages for more details.\
\ \\ For automated download and analysis, each subtrack is stored as a bigBed file that can be\ downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain features within a given range, e.g. for the GENCODE\ Phase I ncORF subtrack:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/ncOrfs/gencNcOrf/Ribo-seq_ORFs.kozak.bb -chrom=chr21 -start=0 -end=100000000 stdout\ \\ File names for all eight subtracks (under\ /gbdb/hg38/ncOrfs/):\
\\ Please refer to each subtrack's description page for references.
\\ References for the Kozak / TE methodology:\
\\ Kozak M.\ \ An analysis of 5'-noncoding sequences from 699 vertebrate messenger RNAs.\ Nucleic Acids Res. 1987 Oct 26;15(20):8125-48.\ DOI: 10.1093/nar/15.20.8125;\ PMID: 3313277; PMC: PMC306349\
\ \\ Noderer WL, Flockhart RJ, Bhaduri A, Diaz de Arce AJ, Zhang J, Khavari PA, Wang CL.\ \ Quantitative analysis of mammalian translation initiation sites by FACS-seq.\ Mol Syst Biol. 2014 Aug 28;10(8):748.\ DOI: 10.15252/msb.20145136;\ PMID: 25170020; PMC: PMC4299517\
\ genes 0 group genes\ longLabel Non-canonical Open Reading Frames\ shortLabel Non-canonical ORFs\ superTrack on\ track ncOrfs\ nonCodingRNAs Non-coding RNA RNA sequences that do not code for a protein 0 100 0 0 0 127 127 127 0 0 0\ This is a super track for non-coding RNA data, subtracks represent some form of non-coding RNA data. \
\The body map RNA-Seq data was kindly provided by the Gene Expression\ Applications research group at Illumina.
\ \ Genome coordinates for the sno/miRNA track were obtained from the miRBase sequences\ FTP site and from \ \ snoRNABase coordinates download page.\\ \
\ When making use of these data, please cite the folowing articles in addition to\ the primary sources of the miRNA sequences:
\\ Griffiths-Jones S, Saini HK, van Dongen S, Enright AJ.\ miRBase: tools for microRNA genomics.\ Nucleic Acids Res. 2008 Jan 1;36(Database issue):D154-8.
\\ Griffiths-Jones S, Grocock RJ, van Dongen S, Bateman A, Enright AJ.\ miRBase: microRNA sequences, targets and gene nomenclature.\ Nucleic Acids Res. 2006 Jan 1;34(Database issue):D140-4.
\\ Griffiths-Jones S.\ The microRNA Registry.\ Nucleic Acids Res. 2004 Jan 1;32(Database issue):D109-11.
\\ Weber MJ.\ New human and mouse microRNA genes found by homology search.\
\ You may also want to cite The Wellcome Trust Sanger Institute \ miRBase and The Laboratoire de Biologie Moleculaire \ Eucaryote snoRNABase.
\\ The following publication provides guidelines on miRNA annotation:\ Ambros V. et al., \ A uniform system for microRNA annotation. \ RNA. 2003;9(3):277-9.
\\ \ \
\ Cabili MN, Trapnell C, Goff L, Koziol M, Tazon-Vega B, Regev A, Rinn JL.\ \ Integrative annotation of human large intergenic noncoding RNAs reveals global properties and\ specific subclasses.\ Genes Dev. 2011 Sep 15;25(18):1915-27.\ PMID: 21890647; PMC: PMC3185964\
\ \\ Trapnell C, Williams BA, Pertea G, Mortazavi A, Kwan G, van Baren MJ, Salzberg SL, Wold BJ, Pachter\ L.\ \ Transcript assembly and quantification by RNA-Seq reveals unannotated transcripts and isoform\ switching during cell differentiation.\ Nat Biotechnol. 2010 May;28(5):511-5.\ PMID: 20436464; PMC: PMC3146043\
\ \ \ genes 0 group genes\ longLabel RNA sequences that do not code for a protein\ shortLabel Non-coding RNA\ superTrack on\ track nonCodingRNAs\ nuorfdb nuORFdb bigGenePred ncORFs: nuORFdb - non-canonical ORFs from nuORFdb v1.2 3 100 0 0 0 127 127 127 0 0 0\ This track displays 229,251 non-canonical open reading frames (ORFs) from\ nuORFdb v1.2\ (novel unannotated ORF database), a database of ORFs with evidence of translation detected by\ ribosome profiling (Ribo-seq). nuORFdb was developed at the Broad Institute of MIT and Harvard as a resource for\ identifying non-canonical peptides in immunopeptidomic mass spectrometry datasets.\
\ \\ The ORFs were predicted using a hierarchical pipeline that aggregates ribosome profiling signal\ across 29 primary healthy and cancer tissue samples and cell lines. The pipeline operates at\ multiple levels—individual samples, tissues, and combined across all samples—to predict\ lowly translated ORFs while maintaining sensitivity for tissue-specific variants.\ All ORFs have a minimum length of 8 amino acids.\
\ \\ Items are displayed in bigGenePred format. Each item is labeled with the nuORFdb ORF\ identifier, which encodes the source Ensembl transcript and ORF number (e.g.\ ENST00000488147.1_1_1). Color reflects the categorical\ Kozak consensus strength:\
\\
Strong – A/G at position −3 and G at position +4
\
Moderate – only one of those positions matches
\
Weak – neither position matches
\
non-ATG – near-cognate start codon; the Kozak rule does not apply
\
no context – chromosome edge or context unavailable\
\ Mouseover shows the ORF ID in its host gene, gene biotype, start codon, Kozak\ strength and TE, predictor type, and the simplified plotType category.\
\ \\ Available filters: start codon, Kozak strength, Kozak TE, ORF category\ (plotType: 8 broad classes; or type: 25 finer\ categories).\
\ \\ The track includes the following ORF categories (by type):\
\\ Each item also includes the predicted protein sequence and additional classification fields\ (predictorType, plotType, geneType) from the nuORFdb annotations.\
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. The data can be accessed from\ scripts through our API; the track name is\ "nuorfdb".\
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/ncOrfs/nuorfdb/nuorfdb.kozak.bb -chrom=chr21 -start=0 -end=100000000 stdout\ \\ The original data files can be downloaded from the\ nuORFdb website\ at the Broad Institute.\
\ \\ The nuORFdb v1.2 data files (BED12 coordinates, Excel annotations, and protein FASTA sequences)\ were downloaded from the Broad Institute. The BED12 file was combined with the annotation\ spreadsheet (keyed on ORF_ID_hg38) and protein FASTA (keyed on sequence header ID) to\ produce a bigGenePred+ format file with 23 fields (12 standard BED fields, 8 bigGenePred fields,\ and 3 extended fields: predictorType, plotType, and proteinSequence).\
\ \\ A small number of entries (176 out of 229,251) used non-standard chromosome names\ (e.g. chrGL000008.2, chrMT) which were mapped to UCSC standard names\ (e.g. chr4_GL000008v2_random, chrM).\
\ \\ Thanks to Tamara Ouspenskaia, Travis Law, Karl Clauser, and colleagues at the Broad Institute\ of MIT and Harvard for creating nuORFdb and making the data publicly available.\ Thanks to Eric Malekos, UCSC, for suggesting this database.\
\ \\ Ouspenskaia T, Law T, Clauser KR, Klaeger S, Sarkizova S, Aguet F, Li B, Christian E, Knisbacher BA,\ Le PM et al.\ \ Unannotated proteins expand the MHC-I-restricted immunopeptidome in cancer.\ Nat Biotechnol. 2022 Feb;40(2):209-217.\ DOI: 10.1038/s41587-021-01021-3; PMID: 34663921; PMC: PMC10198624\
\ \ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/ncOrfs/nuorfdb/nuorfdb.kozak.bb\ filter.kozakTE -1:1.5\ filterByRange.kozakTE on\ filterLimits.kozakTE -1:1.5\ filterType.kozakStrength multipleListOr\ filterType.plotType multipleListOr\ filterType.startCodon multipleListOr\ filterType.type multipleListOr\ filterValues.kozakStrength Strong,Moderate,Weak,non-ATG,None\ filterValues.plotType lincRNA,Out-of-Frame,5' uORF,3' dORF,5' Overlap uORF,3' Overlap dORF,Pseudogene,Other\ filterValues.startCodon ATG,CTG,GTG,TTG,ACG,other,none\ filterValues.type Out-of-Frame,5' uORF,3' dORF,lincRNA,5' Overlap uORF,ncRNA Retained Intron,3' Overlap dORF,ncRNA Processed Transcript,Pseudogene,Antisense,TUCP,Nonsense Mediated Decay,TEC,Sense Overlapping,snoRNA\ itemRgb on\ longLabel ncORFs: nuORFdb - non-canonical ORFs from nuORFdb v1.2\ mouseOver $name in $geneName2 ($geneType)\ This track displays 921,170 protein-coding ORFs from\ OpenProt v2.2, a database that\ provides a comprehensive annotation of all possible protein-coding ORFs in the human genome.\ In addition to currently annotated coding sequences (CDSs) and their reference proteins\ (RefProts), OpenProt predicts alternative ORFs (AltORFs) and their corresponding alternative\ proteins (AltProts) that are hidden within transcripts previously considered to encode only\ a single protein.\
\ \\ A pre-filtered subtrack (OpenProt MS>=2) is also available, containing only the\ 377,916 ORFs with at least 2 unique mass spectrometry peptides detected across studies,\ matching the MS-evidence threshold used by OpenProt for their curated downloads.\
\ \\ OpenProt classifies proteins into three types:\
\\ AltORFs are further classified by their localization relative to the annotated CDS:\
\\ Items are displayed in bigGenePred format. Items are labeled with the protein accession\ number: IDs starting with IP_ are predicted AltProts, II_ are novel\ isoforms, and other IDs (e.g. NP_, ENSP) are RefProts from existing\ annotations. Color reflects the categorical Kozak consensus strength:\
\\
Strong – A/G at position −3 and G at position +4
\
Moderate – only one of those positions matches
\
Weak – neither position matches
\
non-ATG – near-cognate start codon; the Kozak rule does not apply
\
no context – chromosome edge or context unavailable\
\ Mouseover shows the protein accession in its host gene, protein type and ORF\ localization, start codon, Kozak strength and TE, MS score, TE score, and InterPro\ domain count.\
\ \The track includes the following filter options:
\\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. The data can be accessed from\ scripts through our API; the track name is\ "openprot" (all ORFs) or "openprotMs" (MS-filtered).\
\ \\ For automated download and analysis, the genome annotations are stored in bigBed files that\ can be downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/ncOrfs/openprot/openprot.kozak.bb -chrom=chr21 -start=0 -end=100000000 stdout\ \\ The original data files can be downloaded from the\ OpenProt download page.\
\ \\ The OpenProt v2.2 BED12 and TSV annotation files were downloaded from the OpenProt API.\ The BED file (2,846,289 rows) contains genomic coordinates for all predicted ORFs; since the\ same protein can be mapped through multiple transcripts to identical genomic coordinates,\ deduplication reduced this to 921,170 unique genomic features (3 entries with overlapping\ BED blocks were excluded).\
\ \\ Each BED entry was annotated with metadata from the TSV file by joining on protein accession.\ For proteins with multiple transcript entries in the TSV, the annotation with the highest\ MS score was retained. Extended fields include protein type (AltProt/RefProt/Isoform),\ ORF localization, MS score, TE (Translation Event) score, Kozak motif status, InterPro domain\ count, and reading frame.\
\ \\ The annotation is based on GRCh38.p13, Ensembl release 106, and UniProt release 2022_06_01.\
\ \\ Thanks to Xavier Roucou and the OpenProt team at the Université de Sherbrooke for\ creating OpenProt and making the data publicly available.\
\ \\ Brunet MA, Brunelle M, Lucier JF, Delcourt V, Levesque M, Grenier F, Samandi S, Leblanc S, Aguilar\ JD, Dufour P et al.\ \ OpenProt: a more comprehensive guide to explore eukaryotic coding potential and proteomes.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D403-D410.\ PMID: 30299502; PMC: PMC6323990\
\ \\ Brunet MA, Lucier JF, Levesque M, Leblanc S, Jacques JF, Al-Saedi HRH, Guilloy N, Grenier F, Avino\ M, Fournier I et al.\ \ OpenProt 2021: deeper functional annotation of the coding potential of eukaryotic genomes.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D380-D388.\ PMID: 33179748; PMC: PMC7779043\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/ncOrfs/openprot/openprot.kozak.bb\ filter.kozakTE -1:1.5\ filterByRange.kozakTE on\ filterByRange.msScore on\ filterLimits.kozakTE -1:1.5\ filterType.kozakMotif multipleListOr\ filterType.kozakStrength multipleListOr\ filterType.startCodon multipleListOr\ filterValues.kozakMotif +|Kozak motif present,-|No Kozak motif\ filterValues.kozakStrength Strong,Moderate,Weak,non-ATG,None\ filterValues.localization 5'UTR|5'UTR,3'UTR|3'UTR,CDS|CDS,ncRNA|ncRNA,multiple|multiple\ filterValues.startCodon ATG,CTG,GTG,TTG,ACG,other,none\ filterValues.type AltProt|AltProt,RefProt|RefProt,Isoform|Isoform\ itemRgb on\ longLabel ncORFs: OpenProt - alternative and reference proteins v2.2\ mouseOver $name in $geneName2 ($type, $localization)\ This track displays 921,170 protein-coding ORFs from\ OpenProt v2.2, a database that\ provides a comprehensive annotation of all possible protein-coding ORFs in the human genome.\ In addition to currently annotated coding sequences (CDSs) and their reference proteins\ (RefProts), OpenProt predicts alternative ORFs (AltORFs) and their corresponding alternative\ proteins (AltProts) that are hidden within transcripts previously considered to encode only\ a single protein.\
\ \\ A pre-filtered subtrack (OpenProt MS>=2) is also available, containing only the\ 377,916 ORFs with at least 2 unique mass spectrometry peptides detected across studies,\ matching the MS-evidence threshold used by OpenProt for their curated downloads.\
\ \\ OpenProt classifies proteins into three types:\
\\ AltORFs are further classified by their localization relative to the annotated CDS:\
\\ Items are displayed in bigGenePred format. Items are labeled with the protein accession\ number: IDs starting with IP_ are predicted AltProts, II_ are novel\ isoforms, and other IDs (e.g. NP_, ENSP) are RefProts from existing\ annotations. Color reflects the categorical Kozak consensus strength:\
\\
Strong – A/G at position −3 and G at position +4
\
Moderate – only one of those positions matches
\
Weak – neither position matches
\
non-ATG – near-cognate start codon; the Kozak rule does not apply
\
no context – chromosome edge or context unavailable\
\ Mouseover shows the protein accession in its host gene, protein type and ORF\ localization, start codon, Kozak strength and TE, MS score, TE score, and InterPro\ domain count.\
\ \The track includes the following filter options:
\\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. The data can be accessed from\ scripts through our API; the track name is\ "openprot" (all ORFs) or "openprotMs" (MS-filtered).\
\ \\ For automated download and analysis, the genome annotations are stored in bigBed files that\ can be downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/ncOrfs/openprot/openprot.kozak.bb -chrom=chr21 -start=0 -end=100000000 stdout\ \\ The original data files can be downloaded from the\ OpenProt download page.\
\ \\ The OpenProt v2.2 BED12 and TSV annotation files were downloaded from the OpenProt API.\ The BED file (2,846,289 rows) contains genomic coordinates for all predicted ORFs; since the\ same protein can be mapped through multiple transcripts to identical genomic coordinates,\ deduplication reduced this to 921,170 unique genomic features (3 entries with overlapping\ BED blocks were excluded).\
\ \\ Each BED entry was annotated with metadata from the TSV file by joining on protein accession.\ For proteins with multiple transcript entries in the TSV, the annotation with the highest\ MS score was retained. Extended fields include protein type (AltProt/RefProt/Isoform),\ ORF localization, MS score, TE (Translation Event) score, Kozak motif status, InterPro domain\ count, and reading frame.\
\ \\ The annotation is based on GRCh38.p13, Ensembl release 106, and UniProt release 2022_06_01.\
\ \\ Thanks to Xavier Roucou and the OpenProt team at the Université de Sherbrooke for\ creating OpenProt and making the data publicly available.\
\ \\ Brunet MA, Brunelle M, Lucier JF, Delcourt V, Levesque M, Grenier F, Samandi S, Leblanc S, Aguilar\ JD, Dufour P et al.\ \ OpenProt: a more comprehensive guide to explore eukaryotic coding potential and proteomes.\ Nucleic Acids Res. 2019 Jan 8;47(D1):D403-D410.\ PMID: 30299502; PMC: PMC6323990\
\ \\ Brunet MA, Lucier JF, Levesque M, Leblanc S, Jacques JF, Al-Saedi HRH, Guilloy N, Grenier F, Avino\ M, Fournier I et al.\ \ OpenProt 2021: deeper functional annotation of the coding potential of eukaryotic genomes.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D380-D388.\ PMID: 33179748; PMC: PMC7779043\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/ncOrfs/openprot/openprot.ms2.kozak.bb\ filter.kozakTE -1:1.5\ filterByRange.kozakTE on\ filterByRange.msScore on\ filterLimits.kozakTE -1:1.5\ filterType.kozakMotif multipleListOr\ filterType.kozakStrength multipleListOr\ filterType.startCodon multipleListOr\ filterValues.kozakMotif +|Kozak motif present,-|No Kozak motif\ filterValues.kozakStrength Strong,Moderate,Weak,non-ATG,None\ filterValues.localization 5'UTR|5'UTR,3'UTR|3'UTR,CDS|CDS,ncRNA|ncRNA,multiple|multiple\ filterValues.startCodon ATG,CTG,GTG,TTG,ACG,other,none\ filterValues.type AltProt|AltProt,RefProt|RefProt,Isoform|Isoform\ html openprot\ itemRgb on\ longLabel ncORFs: OpenProt - proteins with at least 2 MS peptides v2.2\ mouseOver $name in $geneName2 ($type, $localization)\ This track displays literature-curated regulatory regions, transcription\ factor binding sites, and regulatory polymorphisms from\ ORegAnno (Open Regulatory Annotation). For more detailed\ information on a particular regulatory element, follow the link to ORegAnno\ from the details page. \ \
\ \The display may be filtered to show only selected region types, such as:
\ \To exclude a region type, uncheck the appropriate box in the list at the top of \ the Track Settings page.
\ \\ An ORegAnno record describes an experimentally proven and published regulatory\ region (promoter, enhancer, etc.), transcription factor binding site, or\ regulatory polymorphism. Each annotation must have the following attributes:\
\ ORegAnno core team and principal contacts: Stephen Montgomery, Obi Griffith, \ and Steven Jones from Canada's Michael Smith Genome Sciences Centre, Vancouver, \ British Columbia, Canada.
\\ The ORegAnno community (please see individual citations for various\ features): ORegAnno Citation.\ \
\ Lesurf R, Cotto KC, Wang G, Griffith M, Kasaian K, Jones SJ, Montgomery SB, Griffith OL, Open\ Regulatory Annotation Consortium..\ \ ORegAnno 3.0: a community-driven resource for curated regulatory annotation.\ Nucleic Acids Res. 2016 Jan 4;44(D1):D126-32.\ PMID: 26578589; PMC: PMC4702855\
\ \\ Griffith OL, Montgomery SB, Bernier B, Chu B, Kasaian K, Aerts S, Mahony S, Sleumer MC, Bilenky M,\ Haeussler M et al.\ \ ORegAnno: an open-access community-driven resource for regulatory annotation.\ Nucleic Acids Res. 2008 Jan;36(Database issue):D107-13.\ PMID: 18006570; PMC: PMC2239002\
\ \\ Montgomery SB, Griffith OL, Sleumer MC, Bergman CM, Bilenky M, Pleasance ED, \ Prychyna Y, Zhang X, Jones SJ. \ ORegAnno: an open access database and curation system for \ literature-derived promoters, transcription factor binding sites and regulatory variation.\ Bioinformatics. 2006 Mar 1;22(5):637-40.\ PMID: 16397004\
\ \ regulation 1 color 102,102,0\ group regulation\ longLabel Regulatory elements from ORegAnno\ shortLabel ORegAnno\ track oreganno\ type bed 4 +\ visibility hide\ orfeomeMrna ORFeome Clones psl ORFeome Collaboration Gene Clones 3 100 34 139 34 144 197 144 0 0 0\ This track show alignments of human clones from the\ ORFeome Collaboration. The goal of the project is to be an\ "unrestricted source of fully sequence-validated full-ORF human cDNA\ clones in a format allowing easy transfer of the ORF sequences into\ virtually any type of expression vector. A major goal is to provide\ at least one fully-sequenced full-ORF clone for each human, mouse, and zebrafish gene.\ This track is updated automatically as new clones become available.\
\ \\ The track follows the display conventions for\ gene prediction\ tracks.
\ \\ ORFeome human clones were obtained from GenBank and aligned against the\ genome using the blat program. When a single clone aligned in multiple\ places, the alignment having the highest base identity was found. Only alignments\ having a base identity level within 0.5% of the best and at least 96% base\ identity with the genomic sequence were kept.\
\ \\ Visit the ORFeome Collaboration\ members page for a list of credits and references.\
\ genes 1 baseColorDefault diffCodons\ baseColorUseCds genbank\ baseColorUseSequence genbank\ color 34,139,34\ group genes\ indelDoubleInsert on\ indelQueryInsert on\ longLabel ORFeome Collaboration Gene Clones\ parent mgcOrfeomeMrna\ shortLabel ORFeome Clones\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ track orfeomeMrna\ type psl\ visibility pack\ orphadata Orphanet bigBed 9 + Orphadata: Aggregated Data From Orphanet 0 100 0 0 0 127 127 127 0 0 0 http://www.orpha.net/consor/cgi-bin/OC_Exp.php?lng=en&Expert=$$\ The Orphadata: Aggregated data from Orphanet (Orphanet) track shows genomic positions \ of genes and their association to human disorders, related epidemiological data, and phenotypic\ annotations. As a consortium of 40 countries throughout the world, \ Orphanet\ gathers and improves knowledge regarding rare diseases and maintains the Orphanet rare disease \ nomenclature (ORPHAcode), essential in improving the visibility of rare diseases in health and\ research information systems. The data is updated monthly by Orphanet and updated monthly \ on the UCSC Genome Browser.\
\ \Mouseover on items shows the gene name, disorder name, modes of inheritance(s) (if available), \ and age(s) of onset (if available). Tracks can be filtered according to gene-disorder association \ types, modes of inheritance, and ages of onset. Clicking an item from the browser will return \ the complete entry, including gene linkouts to Ensembl, OMIM, and HGNC, as well as phenotype information \ using HPO (human phenotype ontology) terms.\ \ For more information on the use of this data, see \ the Orphadata FAQs.
\ \The raw data can be explored interactively with the Table Browser, \ or the Data Integrator. \ For automated analysis, the data may be queried from our REST API. \ Please refer to our mailing list archives \ for questions, or our Data Access FAQ \ for more information.\ \
Data is also freely available through \ Orphadata datasets.
\ \Orphadata files were reformatted at UCSC to the \ bigBed format.
\ \Thank you to the Orphanet and Orphadata team and to Tiana Pereira, Christopher Lee, \ Daniel Schmelter, and Anna Benet-Pages of the Genome Browser team.
\ \\ Pavan S, Rommel K, Mateo Marquina ME, Höhn S, Lanneau V, Rath A.\ \ Clinical Practice Guidelines for Rare Diseases: The Orphanet Database.\ PLoS One. 2017;12(1):e0170365.\ PMID: 28099516; PMC: PMC5242437\
\ \\ Nguengang Wakap S, Lambert DM, Olry A, Rodwell C, Gueydan C, Lanneau V, Murphy D, Le Cam Y, Rath A.\ \ Estimating cumulative point prevalence of rare diseases: analysis of the Orphanet database.\ Eur J Hum Genet. 2020 Feb;28(2):165-173.\ PMID: 31527858; PMC: PMC6974615\
\ phenDis 1 bedNameLabel OrphaCode\ bigDataUrl /gbdb/hg38/bbi/orphanet/orphadata.bb\ dataVersion /gbdb/$D/bbi/orphanet/version.txt\ filterValues.assnType Biomarker tested in,Candidate gene tested in,Disease-causing germline mutation(s) (gain of function) in,Disease-causing germline mutation(s) (loss of function) in,Disease-causing germline mutation(s) in,Disease-causing somatic mutation(s) in,Major susceptibility factor in,Modifying germline mutation in,Part of a fusion gene in,Role in the phenotype of\ filterValues.inheritance Autosomal dominant,Autosomal recessive,Mitochondrial inheritance,Multigenic/multifactorial,No data available,Not applicable,Oligogenic,Semi-dominant,Unknown,X-linked dominant,X-linked recessive,Y-linked\ filterValues.onsetList Adolescent,Adult,All ages,Antenatal,Childhood,Elderly,Infancy,Neonatal,No data available\ group phenDis\ itemRgb on\ longLabel Orphadata: Aggregated Data From Orphanet\ mouseOver Gene: $geneSymbol, Disorder: $disorder, Inheritance(s): $inheritance, Onset: $onsetList\ shortLabel Orphanet\ skipEmptyFields on\ skipFields name,score,itemRgb\ track orphadata\ type bigBed 9 +\ url http://www.orpha.net/consor/cgi-bin/OC_Exp.php?lng=en&Expert=$$\ urlLabel OrphaNet Phenotype Link:\ urls ensemblID="https://ensembl.org/Homo_sapiens/Gene/Summary?db=core;g=$$" pmid="https://pubmed.ncbi.nlm.nih.gov/$$" orphaCode="http://www.orpha.net/consor/cgi-bin/OC_Exp.php?lng=en&Expert=$$" omim="https://www.omim.org/entry/$$?search=$$&highlight=$$" hgnc="https://www.genenames.org/data/gene-symbol-report/#!/hgnc_id/HGNC:$$"\ xenoEst Other ESTs psl xeno Non-Human ESTs from GenBank 0 100 0 0 0 127 127 127 1 0 0 https://www.ncbi.nlm.nih.gov/htbin-post/Entrez/query?form=4&db=n&term=$$\ This track displays translated blat alignments of expressed sequence tags \ (ESTs) in GenBank from organisms other than human.\ ESTs are single-read sequences, typically about 500 bases in length, that \ usually represent fragments of transcribed genes.
\ \\ This track follows the display conventions for \ PSL alignment tracks. In dense display mode, the items that\ are more darkly shaded indicate matches of better quality.
\\ The strand information (+/-) for this track is in two parts. The\ first + or - indicates the orientation of the query sequence whose\ translated protein produced the match. The second + or - indicates the\ orientation of the matching translated genomic sequence. Because the two\ orientations of a DNA sequence give different predicted protein sequences,\ there are four combinations. ++ is not the same as --, nor is +- the same\ as -+.
\\ The description page for this track has a filter that can be used to change \ the display mode, alter the color, and include/exclude a subset of items \ within the track. This may be helpful when many items are shown in the track \ display, especially when only some are relevant to the current task.
\\ To use the filter:\
\ This track may also be configured to display base labeling, a feature that\ allows the user to display all bases in the aligning sequence or only those\ that differ from the genomic sequence. For more information about this option,\ go to the\ \ Base Coloring for Alignment Tracks page.\ Several types of alignment gap may also be colored;\ for more information, go to the\ \ Alignment Insertion/Deletion Display Options page.\
\ \\ To generate this track, the ESTs were aligned against the genome using \ blat. When a single EST aligned in multiple places, the \ alignment having the highest base identity was found. Only alignments \ having a base identity level within 0.5% of the best and at least 96% base \ identity with the genomic sequence were kept.
\ \\ This track was produced at UCSC from EST sequence data submitted to the \ international public sequence databases by scientists worldwide.
\ \\ Benson DA, Cavanaugh M, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW.\ \ GenBank.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D36-42.\ PMID: 23193287; PMC: PMC3531190\
\ \\ Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Wheeler DL.\ GenBank: update.\ Nucleic Acids Res. 2004 Jan 1;32(Database issue):D23-6.\ PMID: 14681350; PMC: PMC308779\
\ \\ Kent WJ.\ BLAT - the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ rna 1 baseColorUseSequence genbank\ group rna\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Non-Human ESTs from GenBank\ shortLabel Other ESTs\ spectrum on\ track xenoEst\ type psl xeno\ url https://www.ncbi.nlm.nih.gov/htbin-post/Entrez/query?form=4&db=n&term=$$\ visibility hide\ xenoMrna Other mRNAs psl xeno Non-Human mRNAs from GenBank 0 100 0 0 0 127 127 127 1 0 0\ This track displays translated blat alignments of vertebrate and\ invertebrate mRNA in\ \ GenBank from organisms other than human.\
\ \\ This track follows the display conventions for\ \ PSL alignment tracks. In dense display mode, the items that\ are more darkly shaded indicate matches of better quality.\
\ \\ The strand information (+/-) for this track is in two parts. The\ first + indicates the orientation of the query sequence whose\ translated protein produced the match (here always 5' to 3', hence +).\ The second + or - indicates the orientation of the matching\ translated genomic sequence. Because the two orientations of a DNA\ sequence give different predicted protein sequences, there are four\ combinations. ++ is not the same as --, nor is +- the same as -+.\
\ \\ The description page for this track has a filter that can be used to change\ the display mode, alter the color, and include/exclude a subset of items\ within the track. This may be helpful when many items are shown in the track\ display, especially when only some are relevant to the current task.\
\ \\ To use the filter:\
\ This track may also be configured to display codon coloring, a feature that\ allows the user to quickly compare mRNAs against the genomic sequence. For more\ information about this option, go to the\ \ Codon and Base Coloring for Alignment Tracks page.\ Several types of alignment gap may also be colored;\ for more information, go to the\ \ Alignment Insertion/Deletion Display Options page.\
\ \\ The mRNAs were aligned against the human genome using translated blat.\ When a single mRNA aligned in multiple places, the alignment having the\ highest base identity was found. Only those alignments having a base\ identity level within 1% of the best and at least 25% base identity with the\ genomic sequence were kept.\
\ \\ The mRNA track was produced at UCSC from mRNA sequence data\ submitted to the international public sequence databases by\ scientists worldwide.\
\ \\ Benson DA, Cavanaugh M, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW.\ \ GenBank.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D36-42.\ PMID: 23193287; PMC: PMC3531190\
\ \\ Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Wheeler DL.\ GenBank: update.\ Nucleic Acids Res. 2004 Jan 1;32(Database issue):D23-6.\ PMID: 14681350; PMC: PMC308779\
\ \\ Kent WJ.\ BLAT - the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ rna 1 baseColorUseCds genbank\ baseColorUseSequence genbank\ group rna\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Non-Human mRNAs from GenBank\ shortLabel Other mRNAs\ showDiffBasesAllScales .\ spectrum on\ track xenoMrna\ type psl xeno\ visibility hide\ xenoRefGene Other RefSeq genePred xenoRefPep xenoRefMrna Non-Human RefSeq Genes 0 100 12 12 120 133 133 187 0 0 0\ This track shows known protein-coding and non-protein-coding genes \ for organisms other than human, taken from the NCBI RNA reference \ sequences collection (RefSeq). The data underlying this track are \ updated weekly.
\ \\ This track follows the display conventions for \ gene prediction \ tracks.\ The color shading indicates the level of review the RefSeq record has \ undergone: predicted (light), provisional (medium), reviewed (dark).
\\ The item labels and display colors of features within this track can be\ configured through the controls at the top of the track description page. \
\ The RNAs were aligned against the human genome using blat; those\ with an alignment of less than 15% were discarded. When a single RNA aligned \ in multiple places, the alignment having the highest base identity was \ identified. Only alignments having a base identity level within 0.5% of \ the best and at least 25% base identity with the genomic sequence were kept.\
\ \\ This track was produced at UCSC from RNA sequence data\ generated by scientists worldwide and curated by the \ NCBI RefSeq project.
\ \\ Kent WJ.\ \ BLAT--the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ \\ Pruitt KD, Brown GR, Hiatt SM, Thibaud-Nissen F, Astashyn A, Ermolaeva O, Farrell CM, Hart J,\ Landrum MJ, McGarvey KM et al.\ \ RefSeq: an update on mammalian reference sequences.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D756-63.\ PMID: 24259432; PMC: PMC3965018\
\ \\ Pruitt KD, Tatusova T, Maglott DR.\ \ NCBI Reference Sequence (RefSeq): a curated non-redundant sequence database of genomes, transcripts and proteins.\ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4.\ PMID: 15608248; PMC: PMC539979\
\ genes 1 color 12,12,120\ group genes\ longLabel Non-Human RefSeq Genes\ shortLabel Other RefSeq\ track xenoRefGene\ type genePred xenoRefPep xenoRefMrna\ visibility hide\ gnomADPextOvary Ovary bigWig 0 1 gnomAD pext Ovary 0 100 255 170 255 255 212 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Ovary.bw\ color 255,170,255\ longLabel gnomAD pext Ovary\ parent gnomadPext off\ shortLabel Ovary\ track gnomADPextOvary\ visibility hide\ hprcChainNet Pairwise Alignments bed 3 Human Genomes, Chain/Net pairwise alignments, as mapped by the HPRC project 0 100 0 0 0 255 255 0 0 0 0\ This track shows regions of the human genome that are alignable to other Homo sapiens genomes.\ The alignable parts are shown with thick blocks that look like exons.\ Non-alignable parts between these are shown with thin lines like introns.\ More description on this display can be found below.\
\ \\ Other assemblies included in this track are from the\ HPRC project.\
\ \\ The chain track shows alignments of the human genome to other\ Homo sapiens genomes using a gap scoring system that allows longer gaps\ than traditional affine gap scoring systems. It can also tolerate gaps in both\ source and target assemblies simultaneously. These\ "double-sided" gaps can be caused by local inversions and\ overlapping deletions in both species.\
\ The chain track displays boxes joined together by either single or\ double lines. The boxes represent aligning regions.\ Single lines indicate gaps that are largely due to a deletion in the\ query assembly or an insertion in the target assembly.\ assembly. Double lines represent more complex gaps that involve substantial\ sequence in both species. This may result from inversions, overlapping\ deletions, an abundance of local mutation, or an unsequenced gap in one\ species. In cases where multiple chains align over a particular region of\ the target genome, the chains with single-lined gaps are often\ due to processed pseudogenes, while chains with double-lined gaps are more\ often due to paralogs and unprocessed pseudogenes.
\\ In the "pack" and "full" display\ modes, the individual feature names indicate the chromosome, strand, and\ location (in thousands) of the match for each matching alignment.
\ \By default, the chains to chromosome-based assemblies are colored\ based on which chromosome they map to in the aligning organism. To turn\ off the coloring, check the "off" button next to: Color\ track based on chromosome.
\\ To display only the chains of one chromosome in the aligning\ organism, enter the name of that chromosome (e.g. chr4) in box next to:\ Filter by chromosome.
\ \\ The bigChain files were obtained from the\ HPRC S3 bucket (Amazon Web Services). For more\ information about how the bigChain files were generated, please refer to the HPRC publication below.\
\ \\ Thank you to Glenn Hickey for providing the HAL file from the HPRC project.\
\ \\ Liao WW, Asri M, Ebler J, Doerr D, Haukness M, Hickey G, Lu S, Lucas JK, Monlong J, Abel HJ et\ al.\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ DOI: 10.1038/s41586-023-05896-x; PMID: 37165242; PMC: PMC10172123\
\ \\ Hickey G, Monlong J, Ebler J, Novak AM, Eizenga JM, Gao Y, Human Pangenome Reference Consortium,\ Marschall T, Li H, Paten B.\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nat Biotechnol. 2023 May 10;.\ DOI: 10.1038/s41587-023-01793-w; PMID: 37165083; PMC: PMC10638906\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q, Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ DOI: 10.1038/s41586-020-2871-y; PMID: 33177663; PMC: PMC7673649\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ DOI: 10.1101/gr.123356.111;\ PMID: 21665927; PMC: PMC3166836\
\ hprc 1 altColor 255,255,0\ color 0,0,0\ compositeTrack on\ configurable on\ dimensions dimensionX=subpop dimensionY=sample\ dragAndDrop subTracks\ group hprc\ html hprcChains\ longLabel Human Genomes, Chain/Net pairwise alignments, as mapped by the HPRC project\ noInherit on\ shortLabel Pairwise Alignments\ sortOrder subpop=+ population=+ hap=+ sample=+\ subGroup1 view Views chain=Chains net=Nets\ subGroup2 sample Sample s001=HG02622.mat s002=HG02622.pat s003=HG02717.mat s004=HG02630.pat s005=HG02630.mat s006=HG02717.pat s007=HG02572.pat s008=HG02572.mat s009=HG02886.mat s010=HG02886.pat s011=HG03540.mat s012=HG03540.pat s013=HG02818.pat s014=HG02818.mat s015=HG02723.mat s016=HG02723.pat s017=HG02257.pat s018=HG02257.mat s019=HG02559.pat s020=HG02559.mat s021=HG02486.pat s022=HG02486.mat s023=HG01891.mat s024=HG01891.pat s025=HG02109.mat s026=HG02055.pat s027=HG02109.pat s028=HG02055.mat s029=HG02145.mat s030=HG02145.pat s031=HG03579.mat s032=HG03579.pat s033=HG03453.mat s034=HG03453.pat s035=HG03486.pat s036=HG03486.mat s037=HG03098.pat s038=HG03098.mat s039=NA18906.mat s040=NA18906.pat s041=NA20129.pat s042=NA20129.mat s043=HG03516.pat s044=HG03516.mat s045=HG01175.pat s046=HG01106.pat s047=HG01175.mat s048=HG00741.mat s049=HG00741.pat s050=HG01106.mat s051=HG01071.mat s052=HG00735.pat s053=HG01071.pat s054=HG00735.mat s055=HG01243.pat s056=HG01109.mat s057=HG01243.mat s058=HG01109.pat s059=HG00733.pat s060=HG00733.mat s061=HG02148.pat s062=HG02148.mat s063=HG01952.mat s064=HG01952.pat s065=HG01928.mat s066=HG01928.pat s067=HG01978.pat s068=HG01978.mat s069=HG01258.mat s070=HG01123.mat s071=HG01258.pat s072=HG01361.mat s073=HG01123.pat s074=HG01361.pat s075=HG01358.mat s076=HG01358.pat s077=HG00438.mat s078=HG00673.mat s079=HG00621.pat s080=HG00673.pat s081=HG00438.pat s082=HG00621.mat s083=HG02080.pat s084=HG02080.mat s085=NA21309.mat s086=NA21309.pat s087=T2T-CHM13v2.0 s088=HG03492.pat s089=HG03492.mat\ subGroup3 subpop Subpopulation gwd=Gambian acb=Afr_Carib_Barbados msl=Mende_Sierra_Leone yri=Yoruba_Nigeria asw=African_SW_USA esn=Esan_Nigeria pur=Puerto_Rico pel=Peru_Lima clm=Columbia_Medellin chs=Han_SoChina khv=Vietnam_Kinh pjl=Punjabo_Pakist hapmap=HAPMAP t2t=T2T\ subGroup4 population Population afr=African amr=American eas=East_Asian eur=European sas=South_Asian other=other\ subGroup5 hap Haplotype mat=maternal pat=paternal pri=primary\ track hprcChainNet\ type bed 3\ visibility hide\ gnomADPextPancreas Pancreas bigWig 0 1 gnomAD pext Pancreas 0 100 153 85 34 204 170 144 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Pancreas.bw\ color 153,85,34\ longLabel gnomAD pext Pancreas\ parent gnomadPext off\ shortLabel Pancreas\ track gnomADPextPancreas\ visibility hide\ pancreasBaron Pancreas Baron Pancreas single cell sequencing from Baron et al 2016 0 100 0 0 0 127 127 127 0 0 0\ This track shows data from A Single-Cell Transcriptomic Map of the Human and Mouse\ Pancreas Reveals Inter- and Intra-cell Population Structure. Pancreas\ tissue was analyzed using droplet-based single-cell RNA-sequencing (scRNA-seq)\ and subsequent clustering distinguished 14 pancreas-resident cell types based\ on their identified marker genes found in Baron et al., 2016.
\ \\ There are four bar chart tracks in this track collection with pancreas cells\ grouped by either batch (Pancreas Batch),\ cell type (Pancreas Cells), detailed\ cell type (Pancreas Details) and\ donor (Pancreas Donor). The default track\ displayed is pancreas cells grouped by cell type.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| secretory | |
| endothelial | |
| epithelial | |
| fibroblast |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the\ Pancreas Cells\ subtrack, where the bars represent relatively pure cell types. They can give an\ overview of the cell composition within other categories in other subtracks as\ well.
\ \\ Human islets were obtained from two female cadaveric donors ages 51 (human2)\ and 59 (human4) and two male cadaveric donors ages 17 (human1) and 38 (human3).\ The samples collected from human 1-3 were non-diabetic and human 4 had type 2\ diabetes mellitus. Using single-cell RNA-sequencing ~10,000 human pancreatic\ cells were isolated and sequenced. For each donor, several separate batches of\ ~800 cells were prepared and sequenced to obtain an average of about 100,000\ reads per cell. Cells were barcoded using the inDrop platform which follows the\ CEL-Seq protocol for library construction. Paired end sequencing was done on\ the Illumina Hiseq 2500. After filtering out cells with limited numbers of\ detected genes, the dataset contained 8,629 cells from the four donors.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Mayaan Baron, Adrian Veres, Samuel L. Wolock, Aubrey L. Faust, and to\ the many authors who worked on producing and publishing this data set. The data\ were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick then\ reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Baron M, Veres A, Wolock SL, Faust AL, Gaujoux R, Vetere A, Ryu JH, Wagner BK, Shen-Orr SS, Klein AM\ et al.\ \ A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell\ Population Structure.\ Cell Syst. 2016 Oct 26;3(4):346-360.e4.\ PMID: 27667365; PMC: PMC5228327
\ singleCell 0 group singleCell\ longLabel Pancreas single cell sequencing from Baron et al 2016\ shortLabel Pancreas Baron\ superTrack on\ track pancreasBaron\ visibility hide\ pancreasBaronBatch Pancreas Batch bigBarChart Pancreas cells binned by batch from Baron et al 2016 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ This track shows data from A Single-Cell Transcriptomic Map of the Human and Mouse\ Pancreas Reveals Inter- and Intra-cell Population Structure. Pancreas\ tissue was analyzed using droplet-based single-cell RNA-sequencing (scRNA-seq)\ and subsequent clustering distinguished 14 pancreas-resident cell types based\ on their identified marker genes found in Baron et al., 2016.
\ \\ There are four bar chart tracks in this track collection with pancreas cells\ grouped by either batch (Pancreas Batch),\ cell type (Pancreas Cells), detailed\ cell type (Pancreas Details) and\ donor (Pancreas Donor). The default track\ displayed is pancreas cells grouped by cell type.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| secretory | |
| endothelial | |
| epithelial | |
| fibroblast |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the\ Pancreas Cells\ subtrack, where the bars represent relatively pure cell types. They can give an\ overview of the cell composition within other categories in other subtracks as\ well.
\ \\ Human islets were obtained from two female cadaveric donors ages 51 (human2)\ and 59 (human4) and two male cadaveric donors ages 17 (human1) and 38 (human3).\ The samples collected from human 1-3 were non-diabetic and human 4 had type 2\ diabetes mellitus. Using single-cell RNA-sequencing ~10,000 human pancreatic\ cells were isolated and sequenced. For each donor, several separate batches of\ ~800 cells were prepared and sequenced to obtain an average of about 100,000\ reads per cell. Cells were barcoded using the inDrop platform which follows the\ CEL-Seq protocol for library construction. Paired end sequencing was done on\ the Illumina Hiseq 2500. After filtering out cells with limited numbers of\ detected genes, the dataset contained 8,629 cells from the four donors.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Mayaan Baron, Adrian Veres, Samuel L. Wolock, Aubrey L. Faust, and to\ the many authors who worked on producing and publishing this data set. The data\ were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick then\ reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Baron M, Veres A, Wolock SL, Faust AL, Gaujoux R, Vetere A, Ryu JH, Wagner BK, Shen-Orr SS, Klein AM\ et al.\ \ A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell\ Population Structure.\ Cell Syst. 2016 Oct 26;3(4):346-360.e4.\ PMID: 27667365; PMC: PMC5228327
\ singleCell 1 barChartBars human1_lib1 human1_lib2 human1_lib3 human2_lib1 human2_lib2 human2_lib3 human3_lib1 human3_lib2 human3_lib3 human3_lib4 human4_lib1 human4_lib3\ barChartColors #1e56cc #1e57cb #1c56d0 #2b5cb7 #2d5ab7 #275cbc #1256e0 #1055e2 #0f55e5 #0e55e6 #225ac4 #215ac6\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/pancreasBaron/batch.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/pancreasBaron/batch.bb\ defaultLabelFields name\ html pancreasBaron\ labelFields name,name2\ longLabel Pancreas cells binned by batch from Baron et al 2016\ parent pancreasBaron\ shortLabel Pancreas Batch\ track pancreasBaronBatch\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ pancreasBaronCellType Pancreas Cells bigBarChart Pancreas cells binned by cell type from Baron et al 2016 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ This track shows data from A Single-Cell Transcriptomic Map of the Human and Mouse\ Pancreas Reveals Inter- and Intra-cell Population Structure. Pancreas\ tissue was analyzed using droplet-based single-cell RNA-sequencing (scRNA-seq)\ and subsequent clustering distinguished 14 pancreas-resident cell types based\ on their identified marker genes found in Baron et al., 2016.
\ \\ There are four bar chart tracks in this track collection with pancreas cells\ grouped by either batch (Pancreas Batch),\ cell type (Pancreas Cells), detailed\ cell type (Pancreas Details) and\ donor (Pancreas Donor). The default track\ displayed is pancreas cells grouped by cell type.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| secretory | |
| endothelial | |
| epithelial | |
| fibroblast |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the\ Pancreas Cells\ subtrack, where the bars represent relatively pure cell types. They can give an\ overview of the cell composition within other categories in other subtracks as\ well.
\ \\ Human islets were obtained from two female cadaveric donors ages 51 (human2)\ and 59 (human4) and two male cadaveric donors ages 17 (human1) and 38 (human3).\ The samples collected from human 1-3 were non-diabetic and human 4 had type 2\ diabetes mellitus. Using single-cell RNA-sequencing ~10,000 human pancreatic\ cells were isolated and sequenced. For each donor, several separate batches of\ ~800 cells were prepared and sequenced to obtain an average of about 100,000\ reads per cell. Cells were barcoded using the inDrop platform which follows the\ CEL-Seq protocol for library construction. Paired end sequencing was done on\ the Illumina Hiseq 2500. After filtering out cells with limited numbers of\ detected genes, the dataset contained 8,629 cells from the four donors.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Mayaan Baron, Adrian Veres, Samuel L. Wolock, Aubrey L. Faust, and to\ the many authors who worked on producing and publishing this data set. The data\ were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick then\ reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Baron M, Veres A, Wolock SL, Faust AL, Gaujoux R, Vetere A, Ryu JH, Wagner BK, Shen-Orr SS, Klein AM\ et al.\ \ A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell\ Population Structure.\ Cell Syst. 2016 Oct 26;3(4):346-360.e4.\ PMID: 27667365; PMC: PMC5228327
\ singleCell 1 barChartBars acinar_cell stellate_(activated)_cell islet_alpha_cell islet_beta_cell islet_delta_cell ductal_cell endothelial_cell islet_epsilon_cell islet_gamma_cell other stellate_(quiescent)_cell\ barChartColors #0d55e6 #c68c6e #2a58bc #1754d9 #2457c4 #0298be #57d457 #c2cfe7 #7290d0 #f9b9b9 #c58c6e\ barChartLimit 2.5\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/pancreasBaron/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/pancreasBaron/cell_type.bb\ defaultLabelFields name\ html pancreasBaron\ labelFields name,name2\ longLabel Pancreas cells binned by cell type from Baron et al 2016\ parent pancreasBaron\ shortLabel Pancreas Cells\ track pancreasBaronCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ pancreasBaronDetailedCellType Pancreas Details bigBarChart Pancreas cells binned by detailed cell type from Baron et al 2016 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ This track shows data from A Single-Cell Transcriptomic Map of the Human and Mouse\ Pancreas Reveals Inter- and Intra-cell Population Structure. Pancreas\ tissue was analyzed using droplet-based single-cell RNA-sequencing (scRNA-seq)\ and subsequent clustering distinguished 14 pancreas-resident cell types based\ on their identified marker genes found in Baron et al., 2016.
\ \\ There are four bar chart tracks in this track collection with pancreas cells\ grouped by either batch (Pancreas Batch),\ cell type (Pancreas Cells), detailed\ cell type (Pancreas Details) and\ donor (Pancreas Donor). The default track\ displayed is pancreas cells grouped by cell type.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| secretory | |
| endothelial | |
| epithelial | |
| fibroblast |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the\ Pancreas Cells\ subtrack, where the bars represent relatively pure cell types. They can give an\ overview of the cell composition within other categories in other subtracks as\ well.
\ \\ Human islets were obtained from two female cadaveric donors ages 51 (human2)\ and 59 (human4) and two male cadaveric donors ages 17 (human1) and 38 (human3).\ The samples collected from human 1-3 were non-diabetic and human 4 had type 2\ diabetes mellitus. Using single-cell RNA-sequencing ~10,000 human pancreatic\ cells were isolated and sequenced. For each donor, several separate batches of\ ~800 cells were prepared and sequenced to obtain an average of about 100,000\ reads per cell. Cells were barcoded using the inDrop platform which follows the\ CEL-Seq protocol for library construction. Paired end sequencing was done on\ the Illumina Hiseq 2500. After filtering out cells with limited numbers of\ detected genes, the dataset contained 8,629 cells from the four donors.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Mayaan Baron, Adrian Veres, Samuel L. Wolock, Aubrey L. Faust, and to\ the many authors who worked on producing and publishing this data set. The data\ were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick then\ reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Baron M, Veres A, Wolock SL, Faust AL, Gaujoux R, Vetere A, Ryu JH, Wagner BK, Shen-Orr SS, Klein AM\ et al.\ \ A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell\ Population Structure.\ Cell Syst. 2016 Oct 26;3(4):346-360.e4.\ PMID: 27667365; PMC: PMC5228327
\ singleCell 1 barChartBars acinar activated_stellate alpha beta delta ductal endothelial epsilon gamma macrophage mast quiescent_stellate schwann t_cell\ barChartColors #0d55e6 #c68c6e #2a58bc #1754d9 #2457c4 #0298be #57d457 #c2cfe7 #7290d0 #f5bcbc #edc0c0 #c58c6e #dfcac6 #eadadb\ barChartLimit 2.5\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/pancreasBaron/detailed_cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/pancreasBaron/detailed_cell_type.bb\ defaultLabelFields name\ html pancreasBaron\ labelFields name,name2\ longLabel Pancreas cells binned by detailed cell type from Baron et al 2016\ parent pancreasBaron\ shortLabel Pancreas Details\ track pancreasBaronDetailedCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ pancreasBaronDonor Pancreas Donor bigBarChart Pancreas cells binned by organ donor from Baron et al 2016 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ This track shows data from A Single-Cell Transcriptomic Map of the Human and Mouse\ Pancreas Reveals Inter- and Intra-cell Population Structure. Pancreas\ tissue was analyzed using droplet-based single-cell RNA-sequencing (scRNA-seq)\ and subsequent clustering distinguished 14 pancreas-resident cell types based\ on their identified marker genes found in Baron et al., 2016.
\ \\ There are four bar chart tracks in this track collection with pancreas cells\ grouped by either batch (Pancreas Batch),\ cell type (Pancreas Cells), detailed\ cell type (Pancreas Details) and\ donor (Pancreas Donor). The default track\ displayed is pancreas cells grouped by cell type.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| secretory | |
| endothelial | |
| epithelial | |
| fibroblast |
\ Cells that fall into multiple classes will be colored by blending the colors\ associated with those classes. The colors will be purest in the\ Pancreas Cells\ subtrack, where the bars represent relatively pure cell types. They can give an\ overview of the cell composition within other categories in other subtracks as\ well.
\ \\ Human islets were obtained from two female cadaveric donors ages 51 (human2)\ and 59 (human4) and two male cadaveric donors ages 17 (human1) and 38 (human3).\ The samples collected from human 1-3 were non-diabetic and human 4 had type 2\ diabetes mellitus. Using single-cell RNA-sequencing ~10,000 human pancreatic\ cells were isolated and sequenced. For each donor, several separate batches of\ ~800 cells were prepared and sequenced to obtain an average of about 100,000\ reads per cell. Cells were barcoded using the inDrop platform which follows the\ CEL-Seq protocol for library construction. Paired end sequencing was done on\ the Illumina Hiseq 2500. After filtering out cells with limited numbers of\ detected genes, the dataset contained 8,629 cells from the four donors.
\ \\ The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Mayaan Baron, Adrian Veres, Samuel L. Wolock, Aubrey L. Faust, and to\ the many authors who worked on producing and publishing this data set. The data\ were integrated into the UCSC Genome Browser by Jim Kent and Brittney Wick then\ reviewed by Jairo Navarro. The UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Baron M, Veres A, Wolock SL, Faust AL, Gaujoux R, Vetere A, Ryu JH, Wagner BK, Shen-Orr SS, Klein AM\ et al.\ \ A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell\ Population Structure.\ Cell Syst. 2016 Oct 26;3(4):346-360.e4.\ PMID: 27667365; PMC: PMC5228327
\ singleCell 1 barChartBars human1 human2 human3 human4\ barChartColors #1d56cf #2a5bba #0f55e4 #225ac5\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/pancreasBaron/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/pancreasBaron/donor.bb\ defaultLabelFields name\ html pancreasBaron\ labelFields name,name2\ longLabel Pancreas cells binned by organ donor from Baron et al 2016\ parent pancreasBaron\ shortLabel Pancreas Donor\ track pancreasBaronDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-pancreas&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ panelApp PanelApp bigBed 9 + Genomics England and Australia PanelApp Diagnostics 0 100 0 0 0 127 127 127 0 0 0\ The PanelApp tracks show regions that are related to human disorders. These can be either\ genes, short tandem repeats, or copy number variants. The regions were curated by groups of\ specialists collaborating using the PanelApp web tool. The primary website is Genomics England PanelApp.\ Another deployment of the website, with different data, is \ PanelApp Australia.\
\ \\ Originally, PanelApp was developed to aid interpretation of participant genomes in the\ \ 100,000 Genomes Project.\ Genomics England PanelApp\ is now being used as the platform for achieving consensus on gene panels in the NHS\ Genomic Medicine Service (GMS). Later, the same platform was also deployed by\ Australian\ Genomics.\
\ \\ Genes and genomic\ entities, so short tandem repeats/STRs and copy number variants/CNVs,\ have been reviewed by experts to enable a community consensus to be reached on which\ genes and genomic entities should appear on a diagnostics grade panel for each disorder.\ A rating system (confidence level 0 - 3) is used to classify the level of evidence\ supporting association with phenotypes covered by the gene panel in question.\
\ \\ There are six subtracks in total: Three different types (genes, STRs, and CNVs), and these \ three exist for both countries, England and Australia. The three types of tracks are:
\ \\
There are a few differences between the Genomics England and the Australian Genomics tracks:
\
Genomics England\
\ The individual tracks are colored by confidence level:\ \
\ Mouseover on items shows the gene name, panel associated, mode of inheritance \ (if known), phenotypes related to the gene, and confidence level. Tracks can \ be filtered according to the confidence \ level of disease association evidence. For more information on \ the use of this data, see the PanelApp\ FAQs.\
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API.\
\\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ The files for this track are called genes.bb, tandRep.bb, and cnv.bb. Individual\ regions or the whole genome annotation can be obtained using our tool bigBedToBed,\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool\ can also be used to obtain only features within a given range, e.g. \ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/panelApp/genes.bb -chrom=chr21 -start=0 -end=100000000 stdout
\ \\ Please refer to our\ \ mailing list archives for questions, or our\ \ Data Access FAQ for more information.\
\\ Data is also freely available on the\ Genomics England PanelApp API\ and the Australia PanelApp API.\
\ \\ This track is updated automatically every week. If you need to access older releases of the data,\ you can download them from our archive directory on the download server. To load them into the browser, select a week on the archive directory, copy the link to a file, go to My Data > Custom Tracks, click "Add custom track", paste the link into the box, and click "Submit".\
\ \\ PanelApp files were reformatted at UCSC to the bigBed format. The script that updates the track is called \ doPanelApp.py and can be found in our GitHub repository.\
\ \\ Thank you to Genomics England PanelApp, especially Catherine Snow for technical\ coordination and consultation, and Zornitza Stark from Australia PanelApp.\ Thanks to Beagan Nguy, Lou Nassar, Christopher Lee, Daniel Schmelter, Ana\ Benet-Pagès and Maximilian Haeussler of the Genome Browser team for the\ creation of the tracks.\
\ \\ Martin AR, Williams E, Foulger RE, Leigh S, Daugherty LC, Niblock O, Leong IUS, Smith KR,\ Gerasimenko O, Haraldsdottir E et al.\ \ PanelApp crowdsources expert knowledge to establish consensus diagnostic gene panels.\ Nat Genet. 2019 Nov;51(11):1560-1565.\ PMID: 31676867\
\ phenDis 1 compositeTrack on\ dataVersion /gbdb/$D/panelApp/version.txt\ group phenDis\ longLabel Genomics England and Australia PanelApp Diagnostics\ noParentConfig on\ shortLabel PanelApp\ showCfg on\ track panelApp\ type bigBed 9 +\ visibility hide\ panmask151b Panmask Easy 151b bigBed 3 Panmask Easy 151b Regions: High accuracy for variant calling 0 100 0 0 0 127 127 127 0 0 0\ This container track helps call out sections of the genome that often cause problems or\ confusion when working with the genome. The hg19 genome has a track with the same name, but with\ more subtracks, as the GeT-RM and Genome-in-a-Bottle artifact variants do not exist \ for hg38.\ \
\ The Problematic Regions track contains the following subtracks:\
\ The Highly Reproducible Regions track highlights regions and variants\ from eight samples that can be used to assess variant detection pipelines. The\ "Highly Reproducible Regions" subtrack comprises the intersection of the reproducible\ regions across all eight samples, while the "Variants" subtracks contain the reproducible\ variants from each assayed sample. Both tracks contain data from the following samples:\
\The Genome in a Bottle (GIAB) Problematic Regions tracks provide stratifications of the\ genome to evaluate variant calls in complex regions. It is designed for use with Global Alliance\ for Genomic Health (GA4GH) benchmarking tools like\ hap.py\ and includes regions with low complexity, segmental duplications, functional regions,\ and difficult-to-sequence areas. Developed in collaboration with GA4GH, the\ Genome in a Bottle (GIAB) consortium, and the\ Telomere-to-Telomere Consortium (T2T), the dataset aims to standardize the\ analysis of genetic variation by offering pre-defined BED files for stratifying true and false\ positives in genomic studies, facilitating accurate assessments in complex areas of the genome.
\ \\ The creation of the GIAB Problematic Regions tracks involves using a pipeline and configuration to\ generate stratification BED files that categorize genomic regions based on specific challenges,\ such as low complexity or difficult mapping, to facilitate accurate benchmarking of variant calls.\ For more information on the pipeline and configuration used, please visit the following webpage:\ \ https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/genome-stratifications/v3.5/README.md.\ If you have questions or comments, please write to Justin Zook (jzook@nist.gov).
\ \\ The Panmask Easy 151b Regions subtrack contains a set of sample-agnostic easy regions where\ short-read variant calling reaches high accuracy. Easy regions are derived for variant filtration\ agnostic to individual samples. They are genomic intervals where general variant callers achieve\ high accuracy without sophisticated filtering.
\\ A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that\ could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples\ were inspected to exclude regions with excessively high or low coverage or those enriched with\ low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure\ are referred to as pm151:lenient, where "pm" stands for panmask. In addition, low\ complexity regions identified by SDUST were removed.
\The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and\ other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38,\ 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter\ variant calls for clinical or research human samples. Like the HighRepro track in this container\ (see above), it shows regions that are easy to sequence, not those that are problematic. The data\ was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.
\ \\ Each track contains a set of regions of varying length with no special configuration options. \ The UCSC Unusual Regions track has a mouse-over description, all other tracks have at most\ a name field, which can be shown in pack mode. The tracks are usually kept in dense mode.\
\ \\ The Hide empty subtracks control hides subtracks with no data in the browser window.\ Changing the browser window by zooming or scrolling may result in the display of a different\ selection of tracks.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator.\ \
\
For automated download and analysis, the genome annotation is stored in bigBed files that\
can be downloaded from\
our download server.\
Individual\
regions or the whole genome annotation can be obtained using our tool bigBedToBed\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g. \
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/problematic/comments.bb -chrom=chr21 -start=0 -end=100000000 stdout
\
\ Files were downloaded from the respective databases and converted to bigBed format.\ The procedure is documented in our\ hg38 makeDoc file.\
\ \\ Thanks to Anna Benet-Pagès, Max Haeussler, Angie Hinrichs, Daniel Schmelter, and Jairo\ Navarro at the UCSC Genome Browser for planning, building, and testing these tracks. The\ underlying data comes from the\ ENCODE Blacklist and some parts were copied manually from the HGNC and NCBI\ RefSeq tracks.\
\ \\ Amemiya HM, Kundaje A, Boyle AP.\ \ The ENCODE Blacklist: Identification of Problematic Regions of the Genome.\ Sci Rep. 2019 Jun 27;9(1):9354.\ PMID: 31249361; PMC: PMC6597582\
\ \\ Dwarshuis N, Kalra D, McDaniel J, Sanio P, Alvarez Jerez P, Jadhav B, Huang WE, Mondal R, Busby B,\ Olson ND et al.\ \ The GIAB genomic stratifications resource for human reference genomes.\ Nat Commun. 2024 Oct 19;15(1):9029.\ PMID: 39424793; PMC: PMC11489684\
\ \\ Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, Gonzalez-Porta M, Eberle MA,\ Tezak Z, Lababidi S et al.\ \ Best practices for benchmarking germline small-variant calls in human genomes.\ Nat Biotechnol. 2019 May;37(5):555-560.\ PMID: 30858580; PMC: PMC6699627\
\ \\ Li H.\ \ Finding easy regions for short-read variant calling from pangenome data.\ ArXiv. 2025 Aug 8;.\ PMID: 40799803; PMC: PMC12340882\
\ \\ Pan B, Ren L, Onuchic V, Guan M, Kusko R, Bruinsma S, Trigg L, Scherer A, Ning B, Zhang C et\ al.\ \ Assessing reproducibility of inherited variants detected with short-read whole genome\ sequencing.\ Genome Biol. 2022 Jan 3;23(1):2.\ PMID: 34980216; PMC: PMC8722114\
\ map 1 bigDataUrl /gbdb/hg38/problematic/hg38.pm151b-v3.easy.bb\ dataVersion pm151b-v3.easy.bed.gz (Panmask v1.4, Aug 6 2025, MD5: 2f59a43dab0b463bafcf3b59fc62)\ html problematic\ longLabel Panmask Easy 151b Regions: High accuracy for variant calling\ parent problematicSuper on\ shortLabel Panmask Easy 151b\ track panmask151b\ type bigBed 3\ visibility hide\ ucscGenePfam Pfam in GENCODE bed 12 Pfam Domains in GENCODE Genes 0 100 20 0 250 137 127 252 0 0 0 https://www.ebi.ac.uk/interpro/search/text/$$/?page=1#table\ Most proteins are composed of one or more conserved functional regions called\ domains. This track shows the high-quality, manually-curated\ \ Pfam-A\ domains found in transcripts located in the GENCODE Genes track by the software HMMER3.\
\ \\ This track follows the display conventions for\ gene\ tracks.\
\ \\ The sequences from the knownGenePep table (see \ GENCODE Genes description page)\ are submitted to the set of Pfam-A HMMs which annotate regions within the\ predicted peptide that are recognizable as Pfam protein domains. These regions\ are then mapped to the transcripts themselves using the\ \ pslMap utility. A complete shell script log for every version of UCSC genes can be found in \ our GitHub repository under \ \ hg/makeDb/doc/ucscGenes, e.g. \ \ mm10.knownGenes17.csh is for the database mm10 and version 17 of UCSC known genes.\
\ \\ Of the several options for filtering out false positives, the "Trusted cutoff (TC)" \ threshold method is used in this track to determine significance. For more information regarding \ thresholds and scores, see the HMMER \ documentation and\ results interpretation pages.\
\ \\ Note: There is currently an undocumented but known HMMER problem which results in lessened \ sensitivity and possible missed searches for some zinc finger domains. Until a fix is released for \ HMMER /PFAM thresholds, please also consult the "UniProt Domains" subtrack of the UniProt\ track for more comprehensive zinc finger annotations.\
\ \\ pslMap was written by Mark Diekhans at UCSC.\
\ \\ Finn RD, Mistry J, Tate J, Coggill P, Heger A, Pollington JE, Gavin OL, Gunasekaran P, Ceric G,\ Forslund K et al.\ The Pfam protein families database.\ Nucleic Acids Res. 2010 Jan;38(Database issue):D211-22.\ PMID: 19920124; PMC: PMC2808889\
\ genes 1 color 20,0,250\ group genes\ html gencodePfam\ longLabel Pfam Domains in GENCODE Genes\ shortLabel Pfam in GENCODE\ track ucscGenePfam\ type bed 12\ url https://www.ebi.ac.uk/interpro/search/text/$$/?page=1#table\ phasedVars Phased Variants bed 12 Phased Variants from various sequencing projects 0 100 0 0 0 127 127 127 0 0 0\ This tracks contains variants of individual genotypes, usually phased, from the projects\ Human Diversity Genome Project, Simons Genome Diversity Project, gnomad's HGDP+1000 Genomes callset,\ and the Mexico Biobank.\ The original release of 1000 Genomes has its own, separate track.\ Projects where the released variants are not phased can be found in the container track "SNV Frequencies".\
\ \\ Available on hg19 and hg38:
\\ Available only on hg38:
\\ Full haplotype display:\ In "pack" mode, this track sorts the haplotypes. This can be\ useful for determining the similarity between the samples and inferring\ inheritance at a particular locus.\ Each sample's phased and/or homozygous genotypes are split into haplotypes,\ clustered by similarity around a central variant (in pink), and sorted for\ display by their position in the clustering tree. Click a variant to center on it.\ The tree (as space allows) is drawn in the label area next to the track image.\ Leaf clusters, in which all haplotypes are identical (at least for the variants\ used in clustering), are colored purple. \
\\ For a full description of how the display works, please see our \ Haplotype Display help page.\ \
\ MXB: Allele frequencies by geographical state and ancestry are available via\ the MexVar platform.\ Raw genotype data are available under controlled access at the\ EGA (Study: EGAS00001005797; Dataset: EGAD00010002361). For the VCFs, email\ andres.moreno@cinvestav.mx.\
\ \\ SGDP: The version used was\ https://sharehost.hms.harvard.edu/genetics/reich_lab/sgdp/vcf_variants/,\ merged with bcftools and lifted to hg38 with CrossMap. \
\ \\ MXB: We thank the Center for Research and Advanced Studies (Cinvestav) of Mexico for\ generating and providing the frequency data, the National Institute of Medical\ Sciences and Nutrition (INCMNSZ) for DNA extraction, and the Ministry of Health\ together with the National Institute of Public Health (INSP) for the design and\ implementation of the National Health Survey 2000 (ENSA 2000). We also thank\ the ENSA-Genomics Consortium for their contributions to sample collection and\ data processing that made possible the construction of the MXB genomic\ resource.\
\\ SGDP: This project was funded by the Simons Foundation. Thanks to David Reich and Swapan \ Mallick for help with importing the data.\
\ \\ Barberena-Jonas C, Medina-Muñoz SG, Cedillo-Castelán V, Sepúlveda-Morales T,\ Gonzaga-Jáuregui C, ENSA Genomics Consortium, García-García L, Ioannidis AG,\ Moreno-Estrada A.\ \ Clinical genetic variation across Hispanic populations in the Mexican Biobank.\ Nat Med. 2026 Jan 21;.\ DOI: 10.1038/s41591-025-04100-z; PMID: 41566040\
\ \\ Sohail M, Moreno-Estrada A.\ \ The Mexican Biobank Project promotes genetic discovery, inclusive science and local capacity\ building.\ Dis Model Mech. 2024 Jan 1;17(1).\ PMID: 38299665; PMC: PMC10855211\
\ \\ Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Corés CD, Barberena-Jonas C, Medina-Muñoz SG,\ Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L et al.\ \ Mexican Biobank advances population and medical genomics of diverse ancestries.\ Nature. 2023 Oct;622(7984):775-783.\ PMID: 37821706; PMC: PMC10600006\
\ \\ Bergström A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, Chen Y, Felkel S, Hallast P, Kamm J\ et al.\ \ Insights into human genetic variation and population history from 929 diverse genomes.\ Science. 2020 Mar 20;367(6484).\ PMID: 32193295; PMC: PMC7115999\
\ \\ Koenig Z, Yohannes MT, Nkambule LL, Zhao X, Goodrich JK, Kim HA, Wilson MW, Tiao G, Hao SP, Sahakian\ N et al.\ \ A harmonized public resource of deeply sequenced diverse human genomes.\ Genome Res. 2024 Jun 25;34(5):796-809.\ PMID: 38749656; PMC: PMC11216312\
\ \\ Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S,\ Tandon A et al.\ \ The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.\ Nature. 2016 Oct 13;538(7624):201-206.\ PMID: 27654912; PMC: PMC5161557\
\ \ varRep 1 group varRep\ longLabel Phased Variants from various sequencing projects\ shortLabel Phased Variants\ superTrack on\ track phasedVars\ type bed 12\ visibility hide\ gnomADPextPituitary Pituitary bigWig 0 1 gnomAD pext Pituitary 0 100 170 255 153 212 255 204 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Pituitary.bw\ color 170,255,153\ longLabel gnomAD pext Pituitary\ parent gnomadPext off\ shortLabel Pituitary\ track gnomADPextPituitary\ visibility hide\ placentaVentoTormoCellType10x Placenta Cells bigBarChart Placenta and decidua cells binned by cell type 10x from Vento-Tormo et al 2018 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars T_cell_CD4+ T_cell_CD8+ extravillous_trophoblast_(EVT) endothelial_cell T_cell_mucosal_(MAIT) myeloid_cell natural_killer_cell_(NK) other_immune_cell syncytiotrophoblast_(SCT) villous_cytotrophoblast_(VCT) decidual_perivascular_cell_(dP) decidual_stromal_cell_(dS) fetal_fibroblast_(fFB)\ barChartColors #f63247 #fa3248 #6026c2 #06bb03 #f73247 #de2903 #f03142 #ee1313 #5823d1 #5923cf #a1288a #be03bb #af4f22\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/cell_type.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by cell type 10x from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Cells\ track placentaVentoTormoCellType10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ placentaVentoTormoCellTypeSs2 Placenta Cells Ss2 bigBarChart Placenta and decidua cells binned by cell type smart-seq2 from Vento-Tormo et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars T_cell_CD4+ T_cell_CD8+ extravillous_trophoblast_(EVT) endothelial_cell T_cell_mucosal_(MAIT) myeloid_cell natural_killer_cell_(NK) other_immune_cell syncytiotrophoblast_(SCT) villous_cytotrophoblast_(VCT) decidual_perivascular_cell_(dP) decidual_stromal_cell_(dS) fetal_fibroblast_(fFB)\ barChartColors #f83147 #fa3249 #906de0 #90e28f #fa7685 #df2902 #f63248 #f46162 #cebef2 #e66b76 #c76bb1 #d456d3 #efdcd3\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/cell_type.stats\ barChartUnit units/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/cell_type.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by cell type smart-seq2 from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Cells Ss2\ track placentaVentoTormoCellTypeSs2\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ placentaVentoTormoCellDetailed10x Placenta Detail bigBarChart Placenta and decidua cells binned by detailed cell type 10x from Vento-Tormo et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars DC1 DC2 EVT Endo_(f) Endo_(m) Endo_L Granulocytes HB ILC3 MAIT MO NK_CD16+ NK_CD16- PB_Naive_CD4_ PB_Naive_CD8 PB_clonal_CD8 Plasma SCT Treg VCT dM1 dM2 dM3 dNK_p dNK1 dNK2 dNK3 dP1 dP2 dS1 dS2 dS3 dT_CD4 dT_CD8 fFB1 fFB2\ barChartColors #ef6665 #ef6565 #6026c2 #78b768 #0db506 #6bc361 #ee6e73 #ce2e17 #f4737d #f73247 #e22016 #f23144 #f97684 #f53246 #f43246 #f73247 #ef6668 #5823d1 #f6737d #5923cf #db2a07 #dc2a08 #d72b0d #e32d36 #ea303d #ef3142 #f03142 #8f3b75 #ad1a9a #bd05b8 #bd05b7 #b2169d #f83247 #f43042 #af4f22 #c48778\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/detailed_cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/detailed_cell_type.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by detailed cell type 10x from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Detail\ track placentaVentoTormoCellDetailed10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ placentaVentoTormoCellDetailedSs2 Placenta Detail Ss2 bigBarChart Placenta and decidua cells binned by detailed cell type smart-seq2 from Vento-Tormo et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars DC1 DC2 EVT Endo_(m) Endo_L Granulocytes HB ILC3 MAIT MO NK_CD16+ NK_CD16- PB_Naive_CD4_ PB_Naive_CD8 PB_clonal_CD8 Plasma SCT Treg VCT dM1 dM2 dM3 dNK_p dNK1 dNK2 dNK3 dP1 dP2 dS1 dS2 dS3 dT_CD4 dT_CD8 fFB1\ barChartColors #f6bcbd #f6bcbc #906de0 #90e18f #dbe7d5 #f2bfc4 #f4c0b6 #fac0c5 #fa7685 #e67061 #f77684 #fbc2c8 #f73146 #f87684 #f97685 #f6bcbd #cebef2 #fabec2 #e66b76 #db2a06 #e6715b #f4c0b7 #f8c2c8 #f27684 #f77685 #f97685 #e4c0d8 #e8bcdf #d458d0 #d458d1 #eab9e4 #fa7684 #f83248 #efdcd3\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/detailed_cell_type.stats\ barChartUnit units/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/detailed_cell_type.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by detailed cell type smart-seq2 from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Detail Ss2\ track placentaVentoTormoCellDetailedSs2\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ placentaVentoTormoLocation10x Placenta Loc bigBarChart Placenta and decidua cells binned by cell location 10x from Vento-Tormo et al 2018 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars Blood Decidua Placenta\ barChartColors #f73246 #c6294e #5923cf\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/Location.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/Location.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by cell location 10x from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Loc\ track placentaVentoTormoLocation10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ placentaVentoTormoLocationSs2 Placenta Loc Ss2 bigBarChart Placenta and decidua cells binned by cell location smart-seq2 from Vento-Tormo et al 2018 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars Blood Decidua\ barChartColors #f22532 #e9222c\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/Location.stats\ barChartUnit units/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/Location.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by cell location smart-seq2 from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Loc Ss2\ track placentaVentoTormoLocationSs2\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ placentaVentoTormoMatFet10x Placenta Mat/Fet bigBarChart Placenta and decidua cells binned by maternal/fetal 10x from Vento-Tormo et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars fetal maternal unknown\ barChartColors #5823d1 #e32935 #6bc361\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/mom_child.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/mom_child.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by maternal/fetal 10x from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Mat/Fet\ track placentaVentoTormoMatFet10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ placentaVentoTormoMatFetSs2 Placenta Mat/Fet Ss2 bigBarChart Placenta and decidua cells binned by maternal/fetal smart-seq2 from Vento-Tormo et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars fetal maternal\ barChartColors #936ddc #f0232e\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/mom_child.stats\ barChartUnit units/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/ss2/mom_child.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by maternal/fetal smart-seq2 from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Mat/Fet Ss2\ track placentaVentoTormoMatFetSs2\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+ss2&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ placenta_placenta_models Placenta models bigBed 12 + Placenta transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-Placenta.bb\ longLabel Placenta transcript models\ parent sample_models_view on\ shortLabel Placenta models\ subGroups view=sample_models_view sample=placenta_placenta type=models\ track placenta_placenta_models\ type bigBed 12 +\ visibility squish\ placenta_placenta_ont_post_models Placenta ONT post models bigBed 12 + Placenta ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Placenta01Rep1.bb\ itemRgb on\ longLabel Placenta ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Placenta ONT post models\ subGroups view=per_expr_models_view sample=placenta_placenta type=post_capture_ont_models\ track placenta_placenta_ont_post_models\ type bigBed 12 +\ visibility hide\ placenta_placenta_ont_post_reads Placenta ONT post reads bam Placenta ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Placenta01Rep1.bam\ longLabel Placenta ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Placenta ONT post reads\ subGroups view=per_expr_reads_view sample=placenta_placenta type=post_capture_ont_reads\ track placenta_placenta_ont_post_reads\ type bam\ visibility hide\ placenta_placenta_ont_pre_models Placenta ONT pre models bigBed 12 + Placenta ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Placenta01Rep1.bb\ itemRgb on\ longLabel Placenta ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Placenta ONT pre models\ subGroups view=per_expr_models_view sample=placenta_placenta type=pre_capture_ont_models\ track placenta_placenta_ont_pre_models\ type bigBed 12 +\ visibility hide\ placenta_placenta_ont_pre_reads Placenta ONT pre reads bam Placenta ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Placenta01Rep1.bam\ longLabel Placenta ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Placenta ONT pre reads\ subGroups view=per_expr_reads_view sample=placenta_placenta type=pre_capture_ont_reads\ track placenta_placenta_ont_pre_reads\ type bam\ visibility hide\ placenta_placenta_pacbio_post_models Placenta PB post models bigBed 12 + Placenta PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Placenta01Rep1.bb\ itemRgb on\ longLabel Placenta PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Placenta PB post models\ subGroups view=per_expr_models_view sample=placenta_placenta type=post_capture_pacbio_models\ track placenta_placenta_pacbio_post_models\ type bigBed 12 +\ visibility hide\ placenta_placenta_pacbio_post_reads Placenta PB post reads bam Placenta PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Placenta01Rep1.bam\ longLabel Placenta PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Placenta PB post reads\ subGroups view=per_expr_reads_view sample=placenta_placenta type=post_capture_pacbio_reads\ track placenta_placenta_pacbio_post_reads\ type bam\ visibility hide\ placenta_placenta_pacbio_pre_models Placenta PB pre models bigBed 12 + Placenta PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Placenta01Rep1.bb\ itemRgb on\ longLabel Placenta PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Placenta PB pre models\ subGroups view=per_expr_models_view sample=placenta_placenta type=pre_capture_pacbio_models\ track placenta_placenta_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ placenta_placenta_pacbio_pre_reads Placenta PB pre reads bam Placenta PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Placenta01Rep1.bam\ longLabel Placenta PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Placenta PB pre reads\ subGroups view=per_expr_reads_view sample=placenta_placenta type=pre_capture_pacbio_reads\ track placenta_placenta_pacbio_pre_reads\ type bam\ visibility hide\ placentaVentoTormoStage10x Placenta Stage bigBarChart Placenta and decidua cells binned by placental stage 10x from Vento-Tormo et al 2018 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 1 barChartBars 12_+_1_LMP_(12_+_1_PCW) 12+2_LMP(10+2_PCW) 6_GW_/_LMP_(4_PCW) 8_+_2_LMP_(6_+_2_PCW) 9_+_2GW_(7_+_2_PCW) 9+2_GW_/_LMP_(7_PCW) 9+4_LMP(7+4_PCW)\ barChartColors #ed2c3a #ec2f3b #6026c3 #d72835 #a62c71 #6226c0 #bd06b6\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/Stage.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/placentaVentoTormo/10x/Stage.bb\ defaultLabelFields name\ html placentaVentoTormo\ labelFields name,name2\ longLabel Placenta and decidua cells binned by placental stage 10x from Vento-Tormo et al 2018\ parent placentaVentoTormo\ shortLabel Placenta Stage\ track placentaVentoTormoStage10x\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=placenta-decidua+10x&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ placentaVentoTormo Placenta Vento-Tormo Placenta and decidua cells from from Vento-Tormo et al 2018 0 100 0 0 0 127 127 127 0 0 0\ This track displays data from Single-cell reconstruction of the early maternal-fetal\ interface in humans. Using droplet-based 10x and plate-based\ Smart-seq2 single cell RNA-sequencing (scRNA-seq) ~70,000 cells were profiled\ from first-trimester placentas with matched decidual cells and maternal\ peripheral blood mononuclear cells (PBMC).
\ \\ This track collection contains nine bar chart tracks of RNA expression in the\ human placenta, decidua, and maternal PBMCs\ where cells are grouped by cell type (Placenta\ Cells, Placenta Cells Ss2), detailed\ cell type (Placenta Detail,\ Placenta Detail Ss2), cell location\ (Placenta Loc,\ Placenta Loc Ss2), stage\ (Placenta Stage), and placenta and\ decidua cells (Placenta Mat/Fet,\ Placenta Mat/Fet Ss2). The default tracks\ displayed are Placenta Cells,\ Placenta Loc,\ Placenta Loc Ss2, and\ Placenta Mat/Fet Ss2.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| muscle | |
| trophoblast | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Placenta Cells and\ Placenta Cells Ss2\ subtracks, where the bars represent relatively pure cell types. They can give an overview of \ the cell composition within other categories in other subtracks as well.
\ \\ Tissue was collected from 5 placentas (6-14 gestational weeks) and 11 deciduas.\ Additionally, blood was drawn from 6 of the donors (D4-D9) and enriched for\ PBMCs using a Ficoll-Paque gradient. Decidual and placental tissue were both\ first macroscopically separated. Decidual tissue was then chopped before\ enzymatic dissociation. Placental villi was scraped from the chorionic membrane\ before enzymatic dissociation. Decidual and blood cells were enriched for\ certain populations using an antibody panel prior to Smart-seq2 library\ preparation. Cells from blood decidua and placenta were enriched using FACS\ prior to 10x Genomics v2 library preparation. Smart-seq2 libraries were\ sequenced on an Illumina HiSeq2000. 10x libraries were sequenced on an Illumina\ HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Roser Vento-Tormo, Mirjana Efremova, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Jairo Navarro. The UCSC \ work was paid for by the Chan Zuckerberg Initiative.
\ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \ singleCell 0 group singleCell\ longLabel Placenta and decidua cells from from Vento-Tormo et al 2018\ shortLabel Placenta Vento-Tormo\ superTrack on\ track placentaVentoTormo\ visibility hide\ platinumGenomes Platinum Genomes vcfTabix Platinum genome variants 0 100 0 0 0 127 127 127 0 0 0\ These tracks show high-confidence "Platinum Genome" variant calls for two individuals,\ NA12877 and NA12878, part of a sequenced 17 member pedigree for family number\ 1463, from the Centre d'Etude du Polymorphisme Humain (CEPH). The hybrid\ track displays a merging of the NA12878 results with variant calls produced by Genome in a\ Bottle, discussed further below. CEPH is an international genetic research center that provides\ a resource of immortalized cell cultures used to map genetic markers, and pedigree 1463\ represents a family lineage from Utah of four grandparents, two parents, and 11 children.\ The whole pedigree was sequenced to 50x depth on a HiSeq 2000 Illumina system, which is\ considered a platinum standard, where platinum refers to the quality and completeness of\ the resulting assembly, such as providing full chromosome scaffolds with phasing and\ haplotypes resolved across the entire genome.
\
\ This figure depicts the pedigree of the family sequenced for this study, where the ID for each\ sample is defined by adding the prefix NA128 to each numbered individual, so that 77 = NA12877\ and 78 = NA12878, corresponding to the VCF tracks available in this track set. The dark orange\ individuals indicate sequences used in the analysis methods, whereas the blue represent the\ founder generations (grandparents), which were also sequenced and used in validation steps.\ The genomes of the parent-child trio on the top right side, 91-92-78, were also sequenced\ during Phase I of the 1000 Genomes Project.
\\ These tracks represent a comprehensive genome-wide set of phased small variants that have been\ validated to high confidence. Sequencing and phasing a larger pedigree, beyond the two parents\ and one child, increases the ability to detect errors and assess the accuracy of more of the\ variants compared to a standard trio analysis. The genetic inheritance data enables creating a more\ comprehensive catalog of "platinum variants" that reflects both high accuracy and\ completeness. These results are significant as a comprehensive set of valid\ single-nucleotide variants (SNVs) and insertions and deletions (indels),\ in both the easy and difficult parts of the genome, provides a vital resource for software\ developers creating the next generation of variant callers, because these are the areas where\ the current methods most need training data to improve their methods. Since every one of the\ variants in this catalog is phased, this data set provides a resource to better assess emerging\ technologies designed to generate valid phasing information. To generate the calls, six analysis\ pipelines to call SNVs and indels were used and merged into one catalog, where the sensitivity of\ the genetic inheritance aided to detect genotyping errors and maximize the chance of only\ including true variants, that might otherwise be removed by suboptimal filtering. Read more\ about the detailed methods in the referenced paper, further describing this variant catalog\ of 4.7 million SNVs plus 0.7 million small (1-50 bp) indels, that are all consistent with\ the pattern of inheritance in the parents and 11 children of this pedigree.
\\ The hybrid track in this set extends the characterization of NA12878\ by incorporating high confidence calls produced by Genome in a Bottle analysis.\ The resulting merged files contain more comprehensive coverage of variation than either\ set independently, for instance, the hg19 version contains over 80,000 more indels than\ either input set. Read more about the hybrid methods at the following link:\ https://github.com/Illumina/PlatinumGenomes/wiki/Hybrid-truthset
\ \\
The VCF files for this track can be obtained from the download server:\
\
https://hgdownload.soe.ucsc.edu/gbdb/hg38/platinumGenomes/.
\
These files were obtained from the Platinum genomes source archive:\
https://s3.eu-central-1.amazonaws.com/platinum-genomes/2017-1.0/ReleaseNotes.txt.\
\ Eberle MA, Fritzilas E, Krusche P, Källberg M, Moore BL, Bekritsky MA, Iqbal Z, Chuang HY,\ Humphray SJ, Halpern AL et al.\ \ A reference data set of 5.4 million phased human variants validated by genetic inheritance from\ sequencing a three-generation 17-member pedigree.\ Genome Res. 2017 Jan;27(1):157-164.\ PMID: 27903644; PMC: PMC5204340\
\ \ varRep 1 compositeTrack on\ configureByPopup off\ dataVersion Release 2017-1.0\ group varRep\ html ../platinumGenomes\ longLabel Platinum genome variants\ shortLabel Platinum Genomes\ track platinumGenomes\ type vcfTabix\ vcfDoFilter off\ vcfDoMaf off\ genePredArchive Prediction Archive genePred Gene Prediction Archive 0 100 0 0 0 127 127 127 0 0 0\ This supertrack is a collection of gene prediction tracks and is composed of the following tracks:\
\\ More information about display conventions, methods, credits, and references can be found on each\ subtrack's description page.
\ genes 1 cartVersion 2\ group genes\ html ../genePredArchive\ longLabel Gene Prediction Archive\ shortLabel Prediction Archive\ superTrack on\ track genePredArchive\ type genePred\ visibility hide\ primateAi PrimateAI-3D bigBed 9 + PrimateAI-3D Pathogenicity Predictions for Missense Variants 1 100 0 0 0 127 127 127 0 0 0\ PrimateAI-3D is a\ semi-supervised 3D convolutional neural network that predicts the pathogenicity of all\ possible missense variants in the human genome. It was trained on 4.5 million benign\ missense variants: 4.3 million common variants from 809 non-human primate individuals\ across 233 species, plus common human variants (>0.1% allele frequency) from gnomAD,\ TOPMed, and UK Biobank. These represent about 6% of all possible human missense variants.\
\ \\ The model operates on voxelized protein structures at 2 Å resolution (from\ AlphaFold or homology models) combined with multiple sequence alignments from 592 species.\ It uses three complementary loss functions: benign variant classification, 3D\ fill-in-the-blank prediction on masked amino acids, and a language model ranking component.\ This track shows 70.7 million scored variants across all protein-coding genes.\
\ \\
Each variant is colored blue (benign) or\
red (pathogenic) based on the Illumina-provided\
Prediction field. Because the three possible alternate bases at a given\
position sometimes produce the same amino acid change (codon degeneracy),\
each item is labeled by default with its nucleotide change (e.g. C>T)\
rather than its amino acid change. The label can be switched to the amino acid\
change via the "Label fields" control in the Track Settings.\
\ Hovering over a variant shows:\
\\ Items can be filtered by prediction (benign/pathogenic), by raw PrimateAI-3D\ score, or by percentile.\
\ \\
Due to the data license, the Table Browser, Data Integrator, and the REST API's\
getData endpoint are disabled for this track. The source data can be\
downloaded from the\
PrimateAI-3D website\
(requires registration). The primate variant database is available at\
PrimAD.\
Our Zoonomia 447-way Mammal/Primate alignment\
track displays the primate variants used in training PrimateAI-3D.\
\ The PrimateAI-3D hg38 site list was downloaded from the Illumina BaseSpace website.\ The tab-separated file contains pre-computed scores for all possible single nucleotide\ missense variants. Positions were formatted as bigBed. The percentile score was put into\ the track score field (scaled to 0-1000). No filtering was applied; all 70.7 million\ scored variants are included.\ A conversion script is available from\ our Github.\
\ \\ Thanks to Illumina, in particular Gao Hong, for making PrimateAI-3D predictions publicly available.\
\ \\ Gao H, Hamp T, Ede J, Schraiber JG, McRae J, Singer-Berk M, Yang Y, Dietrich ASD, Fiziev PP, Kuderna\ LFK et al.\ \ The landscape of tolerated genetic variation in humans and primates.\ Science. 2023 Jun 2;380(6648):eabn8153.\ PMID: 37262156; PMC: PMC10713091\
\ \\ Sundaram L, Gao H, Padigepati SR, McRae JF, Li Y, Kosmicki JA, Fritzilas N, Hakenberg J, Dutta A,\ Shon J et al.\ \ Predicting the clinical impact of human mutation with deep neural networks.\ Nat Genet. 2018 Aug;50(8):1161-1170.\ PMID: 30038395; PMC: PMC6237276\
\ phenDis 1 bigDataUrl /gbdb/hg38/_primateAi/primateAi.bb\ defaultLabelFields name\ filter.percentile 0\ filter.scorePAI3D 0\ filterByRange.percentile on\ filterByRange.scorePAI3D on\ filterLabel.percentile Percentile score\ filterLabel.prediction Prediction\ filterLabel.scorePAI3D PrimateAI-3D raw score (clinical threshold 0.821)\ filterLimits.percentile 0:1\ filterLimits.scorePAI3D 0:1\ filterValues.prediction benign|Benign,pathogenic|Pathogenic\ itemRgb on\ labelFields name,aaChange\ longLabel PrimateAI-3D Pathogenicity Predictions for Missense Variants\ maxWindowToDraw 2000000\ mouseOverField _mouseOver\ parent predictionScoresSuper\ pennantIcon New red ../goldenPath/newsarch.html#050126 "Released May 1, 2026"\ scoreFilter 0\ scoreFilterLimits 0:1000\ shortLabel PrimateAI-3D\ tableBrowser off\ track primateAi\ type bigBed 9 +\ urls gene="https://www.ensembl.org/Homo_sapiens/Transcript/Summary?t=$$" refSeq="https://www.ncbi.nlm.nih.gov/nuccore/$$"\ visibility dense\ problematicSuper Problematic Regions Problematic/special genomic regions for sequencing or very variable regions 0 100 0 0 0 127 127 127 0 0 0\ This container track helps call out sections of the genome that often cause problems or\ confusion when working with the genome. The hg19 genome has a track with the same name, but with\ more subtracks, as the GeT-RM and Genome-in-a-Bottle artifact variants do not exist \ for hg38.\ \
\ The Problematic Regions track contains the following subtracks:\
\ The Highly Reproducible Regions track highlights regions and variants\ from eight samples that can be used to assess variant detection pipelines. The\ "Highly Reproducible Regions" subtrack comprises the intersection of the reproducible\ regions across all eight samples, while the "Variants" subtracks contain the reproducible\ variants from each assayed sample. Both tracks contain data from the following samples:\
\The Genome in a Bottle (GIAB) Problematic Regions tracks provide stratifications of the\ genome to evaluate variant calls in complex regions. It is designed for use with Global Alliance\ for Genomic Health (GA4GH) benchmarking tools like\ hap.py\ and includes regions with low complexity, segmental duplications, functional regions,\ and difficult-to-sequence areas. Developed in collaboration with GA4GH, the\ Genome in a Bottle (GIAB) consortium, and the\ Telomere-to-Telomere Consortium (T2T), the dataset aims to standardize the\ analysis of genetic variation by offering pre-defined BED files for stratifying true and false\ positives in genomic studies, facilitating accurate assessments in complex areas of the genome.
\ \\ The creation of the GIAB Problematic Regions tracks involves using a pipeline and configuration to\ generate stratification BED files that categorize genomic regions based on specific challenges,\ such as low complexity or difficult mapping, to facilitate accurate benchmarking of variant calls.\ For more information on the pipeline and configuration used, please visit the following webpage:\ \ https://ftp-trace.ncbi.nlm.nih.gov/ReferenceSamples/giab/release/genome-stratifications/v3.5/README.md.\ If you have questions or comments, please write to Justin Zook (jzook@nist.gov).
\ \\ The Panmask Easy 151b Regions subtrack contains a set of sample-agnostic easy regions where\ short-read variant calling reaches high accuracy. Easy regions are derived for variant filtration\ agnostic to individual samples. They are genomic intervals where general variant callers achieve\ high accuracy without sophisticated filtering.
\\ A set of easy regions for ancient DNA variant filtering was generated by selecting 35-mers that\ could not be mapped elsewhere within one mismatch or gap. Read alignments from multiple samples\ were inspected to exclude regions with excessively high or low coverage or those enriched with\ low mapping quality alignments. The easy regions generated through this k-mer uniqueness procedure\ are referred to as pm151:lenient, where "pm" stands for panmask. In addition, low\ complexity regions identified by SDUST were removed.
\The pm151 regions are used to filter spurious variant calls in centromeres, long repeats, and\ other genomic regions where short-read mapping is often problematic. They cover 88.2% of hg38,\ 92.2% of coding regions, and 96.3% of ClinVar pathogenic variants. The track can be used to filter\ variant calls for clinical or research human samples. Like the HighRepro track in this container\ (see above), it shows regions that are easy to sequence, not those that are problematic. The data\ was derived from the HPRC assemblies, and this track presents the 151b-easy panmask set.
\ \\ Each track contains a set of regions of varying length with no special configuration options. \ The UCSC Unusual Regions track has a mouse-over description, all other tracks have at most\ a name field, which can be shown in pack mode. The tracks are usually kept in dense mode.\
\ \\ The Hide empty subtracks control hides subtracks with no data in the browser window.\ Changing the browser window by zooming or scrolling may result in the display of a different\ selection of tracks.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator.\ \
\
For automated download and analysis, the genome annotation is stored in bigBed files that\
can be downloaded from\
our download server.\
Individual\
regions or the whole genome annotation can be obtained using our tool bigBedToBed\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g. \
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/problematic/comments.bb -chrom=chr21 -start=0 -end=100000000 stdout
\
\ Files were downloaded from the respective databases and converted to bigBed format.\ The procedure is documented in our\ hg38 makeDoc file.\
\ \\ Thanks to Anna Benet-Pagès, Max Haeussler, Angie Hinrichs, Daniel Schmelter, and Jairo\ Navarro at the UCSC Genome Browser for planning, building, and testing these tracks. The\ underlying data comes from the\ ENCODE Blacklist and some parts were copied manually from the HGNC and NCBI\ RefSeq tracks.\
\ \\ Amemiya HM, Kundaje A, Boyle AP.\ \ The ENCODE Blacklist: Identification of Problematic Regions of the Genome.\ Sci Rep. 2019 Jun 27;9(1):9354.\ PMID: 31249361; PMC: PMC6597582\
\ \\ Dwarshuis N, Kalra D, McDaniel J, Sanio P, Alvarez Jerez P, Jadhav B, Huang WE, Mondal R, Busby B,\ Olson ND et al.\ \ The GIAB genomic stratifications resource for human reference genomes.\ Nat Commun. 2024 Oct 19;15(1):9029.\ PMID: 39424793; PMC: PMC11489684\
\ \\ Krusche P, Trigg L, Boutros PC, Mason CE, De La Vega FM, Moore BL, Gonzalez-Porta M, Eberle MA,\ Tezak Z, Lababidi S et al.\ \ Best practices for benchmarking germline small-variant calls in human genomes.\ Nat Biotechnol. 2019 May;37(5):555-560.\ PMID: 30858580; PMC: PMC6699627\
\ \\ Li H.\ \ Finding easy regions for short-read variant calling from pangenome data.\ ArXiv. 2025 Aug 8;.\ PMID: 40799803; PMC: PMC12340882\
\ \\ Pan B, Ren L, Onuchic V, Guan M, Kusko R, Bruinsma S, Trigg L, Scherer A, Ning B, Zhang C et\ al.\ \ Assessing reproducibility of inherited variants detected with short-read whole genome\ sequencing.\ Genome Biol. 2022 Jan 3;23(1):2.\ PMID: 34980216; PMC: PMC8722114\
\ map 0 group map\ html problematic\ longLabel Problematic/special genomic regions for sequencing or very variable regions\ shortLabel Problematic Regions\ superTrack on show\ track problematicSuper\ promoterAi PromoterAI bigWig PromoterAI Promoter Variant Impact Scores (zoom for exact score) 0 100 0 0 0 127 127 127 0 0 0\ PromoterAI is a deep neural network from Illumina that predicts the\ expression-altering impact of single nucleotide variants in gene promoter regions.\ It scores all possible substitutions within 500 bp of annotated transcription start\ sites (TSS), covering approximately 39.5 million genomic positions across all\ protein-coding genes.\
\ \\ Scores range from -1 to 1. A negative score is a predicted decrease in\ expression of the target gene; a positive score is a predicted increase\ in expression. Scores near zero indicate the variant is predicted to leave expression\ unchanged. Variants at either end of the range (large |score|) are dysregulating and\ are the ones enriched among patients with rare disease in the PromoterAI paper.\
\ \\ Illumina's PromoterAI\ GitHub page recommends three tiered thresholds for interpretation:\ |score| ≥ 0.1, |score| ≥ 0.2, and |score| ≥ 0.5.\ Higher absolute thresholds select progressively smaller, higher-confidence sets of\ predicted expression-altering variants.\
\ \\ This track is a composite with four bigWig subtracks, one for each possible alternate\ allele (A, C, G, T). When zoomed in, the exact PromoterAI score for each possible\ mutation is shown on mouseover. At wider zooms multiple data points fall into a single\ pixel and averaging scores is not biologically meaningful, so the mouseover displays\ "zoom in to see values" until you zoom in far enough that individual values\ can be shown.\
\ \\ A fifth subtrack ("PromoterAI overlaps") shows positions where overlapping\ transcripts produce different scores for the same variant. At these positions, the\ bigWig subtracks show the score with the largest absolute value, while the overlap\ track lists every per-transcript score. About 3.8% of variant positions have\ overlapping transcripts with differing scores; for more than 60% of these, the\ difference is smaller than 0.01. A filter, active by default, hides entries whose\ per-transcript score range is smaller than 0.01. The filter can be adjusted or turned\ off on the track configuration page.\
\ \\ Across all subtracks, coloring follows the direction of the predicted effect:\ red (bars above the zero line in the bigWigs,\ or filled boxes in the overlap subtrack) indicates predicted over-expression (positive\ score), and blue (bars below zero or filled\ boxes) indicates predicted under-expression (negative score).\
\ \\ The PromoterAI predictions are distributed by Illumina under a license that does not\ permit redistribution, so this track is not available for bulk download from UCSC and\ is excluded from the Table Browser and public API. The original prediction files are\ available for academic and non-commercial research use directly from Illumina:\ complete the license agreement linked from the\ PromoterAI GitHub\ page, and a download link is emailed after submission.\
\ \\ The PromoterAI hg38 TSS-500 file was downloaded from Illumina via the PromoterAI\ license agreement. The file\ contains pre-computed scores for all possible single nucleotide substitutions within\ 500 bp of annotated TSS positions. For positions covered by multiple transcripts,\ the score with the largest absolute value was used for the bigWig tracks. Positions\ where transcripts produced different scores (4.45M of 118.6M unique variants, 3.8%)\ were additionally written to a bigBed overlap track with per-transcript detail\ (transcript IDs, per-transcript scores, strand, and the maximum pairwise score\ difference). The conversion script is available from\ our Github.\
\ \\ Thanks to Kishore Jaganathan and colleagues at Illumina for making the PromoterAI\ predictions publicly available for academic and non-commercial research.\
\ \\ Jaganathan K, Ersaro N, Novakovsky G, Wang Y, James T, Schwartzentruber J, Fiziev P,\ Kassam I, Cao F, Hawe J et al.\ \ Predicting expression-altering promoter mutations with deep learning.\ Science. 2025 Aug 7;389(6760):eads7373.\ PMID: 40440429\
\ phenDis 0 compositeTrack on\ longLabel PromoterAI Promoter Variant Impact Scores (zoom for exact score)\ parent predictionScoresSuper\ pennantIcon New red ../goldenPath/newsarch.html#050126 "Released May 1, 2026"\ shortLabel PromoterAI\ tableBrowser off\ track promoterAi\ type bigWig\ visibility hide\ gnomADPextProstate Prostate bigWig 0 1 gnomAD pext Prostate 0 100 221 221 221 238 238 238 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Prostate.bw\ color 221,221,221\ longLabel gnomAD pext Prostate\ parent gnomadPext off\ shortLabel Prostate\ track gnomADPextProstate\ visibility hide\ wgEncodeReg4TxnAllProstateMinus Prostate - (all biosamples) bigWig Avg. - strand total RNA-seq level of 4 prostate experiments (all biosamples) 0 100 140 140 140 197 197 197 0 0 0 regulation 0 bigDataUrl /gbdb/hg38/encode4/regulation/organAve/prostateMinus.bw\ color 140,140,140\ longLabel Avg. - strand total RNA-seq level of 4 prostate experiments (all biosamples)\ negateValues on\ parent wgEncodeReg4Txn off\ priority 100\ shortLabel Prostate - (all biosamples)\ track wgEncodeReg4TxnAllProstateMinus\ type bigWig\ pseudogenes Pseudogenes bigbed Pseudogenes and Parents 0 100 0 0 0 127 127 127 0 0 0\ These tracks contain pseudogene predictions and their parents as identified by PseudoPipe.\ PseudoPipe is a homology-based\ computational pipeline that can search a mammalian genome and identify pseudogene sequences\ comprehensively and consistently.\
\\ Pseudogenes are genomic sequences that bear similarity to specific protein-coding genes, but are\ unable to produce functional proteins due to the existence of frameshifts, premature stop codons, or\ other deleterious mutations. They arise from gene duplication or retrotransposition events and are\ important resources in understanding the evolutionary history of genes and genomes.
\ \This composite track consists of two subtracks: the Pseudogenes track and the Pseudogene\ Parents track.
\\ The Pseudogene Parents track displays parent genes and pseudogenes\ labeled with their HUGO\ IDs, which were derived from Ensembl gene IDs provided by the Gerstein lab after dataset creation. It includes indicators for pseudogenes. \ These indicators do not show pseudogene locations directly but instead indicate how many pseudogenes\ are associated with each gene and link to their genomic regions in the Pseudogenes track.
\\ The Pseudogenes track shows pseudogenes labeled with their parent HUGO ID and colored\ according to pseudogene type. The authors assigned PGOHUMG IDs to genes and PGOHUMT IDs to\ transcripts. Note: Not all PseudoPipe IDs could be mapped back to their original Ensembl\ IDs. In these cases, the gene ID is listed as NA.
\ \ Pseudogene types:\Each parent gene is shown with associated pseudogenes represented as grey blocks. These blocks\ do not reflect actual pseudogene locations but rather indicate the count of pseudogenes linked to\ the gene.\
\\ If a parent gene has four grey blocks beneath it, this indicates the presence of four pseudogenes\ elsewhere in the genome. Hovering over an item displays the gene type, ID (Ensembl transcript ID\ or PseudoPipe transcript ID), and the genome position of the gene or pseudogene, with a link to\ that genomic region.\
\ \Pseudogenes are colored by type.
\\ Hovering over a pseudogene item shows the pseudogene type, parent HUGO gene symbol, and the Ensembl\ parent transcript ID, which links to the genome position of the parent gene.
\ \\ The PseudoPipe pipeline identifies pseudogenes through a series of steps. It first uses BLAST to\ rapidly cross-reference potential parent proteins against the intergenic regions of the genome. The\ resulting raw hits are then processed by removing redundancies, clustering neighboring sequences,\ and aligning each cluster with a unique parent gene. Finally, pseudogenes are classified based on a\ combination of criteria, including homology, intron-exon structure, and the presence of stop codons\ or frameshifts. This method is designed to detect pseudogenes that are unable to be translated into\ proteins.
\\ These tracks were generated using a Bash script that processes a GTF file with pseudogene\ annotations by removing duplicates, correcting overlapping exons, and converting the data to BED\ format with pseudoPipeToBed.py. This script extracts gene and transcript IDs, merges overlapping\ exons, assigns colors based on pseudogene type, and outputs a BED file with gene and parent\ annotations. PseudoPipeParents.py then links pseudogenes to their functional genes by determining\ parent gene coordinates, updating pseudogene entries with interactive browser links and generating a\ parent BED file. The final data are formatted into pseudoPipePgenes.bb and pseudoPipeParents.bb BigBed\ files. The detailed documentation (makeDoc) and \ Python scripts are available in our GitHub repository.\
\ \The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ The data may also be explored interactively using our\ REST API.
\For automated download and analysis, the genome annotation is stored at UCSC in bigBed files\ that can be downloaded from the\ download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed which can be compiled from the source code or downloaded as a precompiled\ binary for your system.
\\ Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/hg38/pseudogenes/pseudoPipePgenes.bb -chrom=chr21 -start=0 -end=10000000 stdout\ \ \Thanks to the Gerstein lab at Yale University for making this data available, and to Cristina\ Sisu for providing data in GTF format with parent annotations.
\ \\ Zhang Z, Carriero N, Zheng D, Karro J, Harrison PM, Gerstein M.\ \ PseudoPipe: an automated pseudogene identification pipeline.\ Bioinformatics. 2006 Jun 15;22(12):1437-9.\ PMID: 16574694\
\ genes 1 compositeTrack on\ group genes\ html pseudogenes.html\ longLabel Pseudogenes and Parents\ noScoreFilter on\ shortLabel Pseudogenes\ track pseudogenes\ type bigbed\ visibility hide\ pubtator PubTator Variants bigBed 9 + dbSNP variants and other genetic variants grounded to dbSNP by tmVar; collected by PubTator3 1 100 0 0 0 127 127 127 0 0 0The tracks that are listed here contain genetic variants and links to scientific publications that \ mention them.
\\ For additional information please click on the hyperlink of the respective track above.\
\ By default, each variant is labeled with the nucleotide change. Hover over the\ feature to see more information, explained on the track details page of the particular track\ or when clicking onto the feature.
\\ For data provenance, access and descriptions, please click the documentation via the link above.\
\ phenDis 1 bigDataUrl /gbdb/hg38/pubs2/pubtatorDbSnp.bb\ exonNumbers off\ html varsInPubs\ itemRgb on\ longLabel dbSNP variants and other genetic variants grounded to dbSNP by tmVar; collected by PubTator3\ mouseOver $name found in ${numPubmedIds} PubMed articles\ noScoreFilter on\ parent varsInPubs pack\ shortLabel PubTator Variants\ track pubtator\ type bigBed 9 +\ urls pubmedIds="https://www.ncbi.nlm.nih.gov/pubmed/$$"\ visibility dense\ per_expr_reads_view Reads bam Capture long-seq long-read lncRNAs 4 100 0 0 0 127 127 127 0 0 0 rna 1 longLabel Capture long-seq long-read lncRNAs\ parent clsLongReadRnaTrack on\ shortLabel Reads\ track per_expr_reads_view\ type bam\ view per_expr_reads_view\ visibility squish\ hprcArrV1 Rearrangements bigBed 9 + Rearrangements including indels, inversions, and duplications 0 100 0 0 0 100 50 0 0 0 0\ This track shows various rearrangements in the HPRC assemblies with respect to hg38. The types include indels, duplications, inversions, and other more complicated \ rearrangements. There are five tracks in the Rearrangement composite track:\ \
\ All items are labeled by the number of HPRC assemblies that have the rearrangement. The indel tracks have one or \ two additional fields that specify how large the indel is in base pairs. \ For the Insertions and Deletions track there's only one number with "bp" after it. \ For insertions, it is the size of the insertion in hg38. \ For deletions, it is the size of the sequence deleted in hg38. \ For the Other Rearrangements track, there are two numbers given: the number of unaligned \ bases in hg38 and the number of unaligned bases in the HPRC assemblies.\
\ All these tracks are built from the HPRC chains and nets. \ The actual instructions used to create these tracks are in the files hprcRearrange.txt and hprcInDel.txt.\ The first step for all the tracks is to find the orthologous sequences in each HPRC assembly for each chromosome in hg38. \ These sequences are called the query sequences. For each query sequence, we select the \ longest chain to the hg38 sequence. This is called the orthologous chain. \ Following are the specific methods for each track.\
\ Wen-Wei Liao, Mobin Asri, Jana Ebler, ...et al, Heng Lin,\ Benedict Paten\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ PMID: 37165242;\ PMC: PMC1017212;\ DOI: 10.1038/s41586-023-05896-x\
\ \\ Glenn Hickey, Jean Monlong, Jana Ebler, Adam M Novak, Jordan M Eizenga,\ Yan Gao; Human Pangenome Reference Consortium; Tobias Marschall, Heng Li,\ Benedict Paten\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nature Biotechnology. 2023 May 10. doi: 10.1038/s41587-023-01793-w.\ PMID: 37165083;\ DOI: 10.1038/s41587-023-01793-w\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q,\ Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ PMID: 33177663;\ PMC: PMC7673649;\ DOI: 10.1038/s41586-020-2871-y\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ PMID: 21665927;\ PMC: PMC3166836;\ DOI: 10.1101/gr.123356.111\
\ \ hprc 1 altColor 100,50,0\ color 0,0,0\ compositeTrack on\ filter.score 1\ filterLabel.score Minimum number of assemblies with arrangement\ group hprc\ longLabel Rearrangements including indels, inversions, and duplications\ priority 100\ shortLabel Rearrangements\ track hprcArrV1\ type bigBed 9 +\ visibility hide\ recombRate2 Recomb Rate bed Recombination rate: Genetic maps from deCODE and 1000 Genomes 0 100 0 130 0 127 192 127 0 0 0\ The recombination rate track represents calculated rates of recombination based\ on the genetic maps from deCODE (Halldorsson et al., 2019) and 1000 Genomes\ (2013 Phase 3 release, lifted from hg19). The deCODE map is more recent, has a higher \ resolution and was natively created on hg38 and therefore recommended. \ For the Recomb. deCODE average track, the recombination rates for chrX represent the female rate.\
\ \This track also includes a subtrack with all the\ individual deCODE recombination events and another subtrack with several thousand\ de-novo mutations found in the deCODE sequencing data. These two tracks are hidden by\ default and have to be switched on explicitly on the configuration page.\
\ \\ This is a super track that contains different subtracks, three with the deCODE\ recombination rates (paternal, maternal and average) and one with the 1000\ Genomes recombination rate (average). These tracks are in \ signal graph\ (wiggle) format. By default, to show most recombination hotspots, their maximum\ value is set to 100 cM, even though many regions have values higher than 100.\ The maximum value can be changed on the configuration pages of the tracks.\
\ \\ There are two more tracks that show additional details provided by deCODE: one\ subtrack with the raw data of all cross-overs tagged with their proband ID and\ another one with around 8000 human de-novo mutation variants that are linked to\ cross-over changes.\
\ \\ The deCODE genetic map was created at \ deCODE Genetics. It is based \ on microarrays assaying 626,828 SNP markers that allowed to identify 1,476,140 crossovers in\ 56,321 paternal meioses and 3,055,395 crossovers in 70,086 maternal meioses.\ In total, the data is based on 4,531,535 crossovers in 126,427 meioses. By\ using WGS data with 9,305,070 SNPs, the boundaries for 761,981 crossovers were\ refined: 247,942 crossovers in 9423 paternal meioses and 514,039 crossovers in\ 11,750 maternal meioses. The average resolution of the genetic map is 682 base\ pairs (bp): 655 and 708 bp for the paternal and maternal maps, respectively.\
\ \The 1000 Genomes genetic map is based on the IMPUTE genetic map based on 1000 Genomes Phase 3, on hg19 coordinates. It\ was converted to hg38 by Po-Ru Loh at the Broad Institute. After a run of \ liftOver, he post-processed the data to deal with situations in which\ consecutive map locations became much closer/farther after lifting. The\ heuristic used is sufficient for statistical phasing but may not be optimal for\ other analyses. For this reason, and because of its higher resolution, the DeCODE\ map is therefore recommended for hg38.\
\ \As with all other tracks, the data conversion commands and pointers to the\ original data files are documented in the \ makeDoc file of this track.
\ \\ The raw data can be explored interactively with the Table Browser, or\ the Data Integrator. For automated access, this track, like all\ others, is available via our API. However, for bulk\ processing, it is recommended to download the dataset.\
\ \\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig and bigBed\
files that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools bigWigToWig\
or bigBedToBed which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to a given range, e.g.,\
\
bigWigToBedGraph -chrom=chr17 -start=45941345 -end=45942345 http://hgdownload.soe.ucsc.edu/gbdb/hg38/recombRate/recombAvg.bw stdout\
\
\ Please refer to our\ Data Access FAQ\ for more information.\
\ \\ This track was produced at UCSC using data that are freely available for\ the deCODE\ and 1000 Genomes genetic maps. Thanks to Po-Ru Loh at the\ Broad Institute for providing the code to lift the hg19 1000 Genomes map data to hg38.\
\ \\ 1000 Genomes Project Consortium., Abecasis GR, Altshuler D, Auton A, Brooks LD, Durbin RM, Gibbs RA,\ Hurles ME, McVean GA.\ \ A map of human genome variation from population-scale sequencing.\ Nature. 2010 Oct 28;467(7319):1061-73.\ PMID: 20981092; PMC: PMC3042601\
\ \\ Halldorsson BV, Palsson G, Stefansson OA, Jonsson H, Hardarson MT, Eggertsson HP, Gunnarsson B,\ Oddsson A, Halldorsson GH, Zink F et al.\ \ Characterizing mutagenic effects of recombination through a sequence-level genetic map.\ Science. 2019 Jan 25;363(6425).\ PMID: 30679340\
\ map 1 color 0,130,0\ group map\ longLabel Recombination rate: Genetic maps from deCODE and 1000 Genomes\ shortLabel Recomb Rate\ superTrack on hide\ track recombRate2\ type bed\ visibility hide\ recount3 recount3 bigBed 9 + recount3 introns 0 100 0 0 0 127 127 127 0 0 0\ Recount3 is a comprehensive resource for re-analyzing RNA-seq data. It provides uniformly processed\ RNA-seq data and associated metadata from a wide range of studies, enabling researchers to access\ and analyze gene expression data in a consistent manner. Recount3 aggregates data from multiple\ sources, including the\ Sequence Read Archive (SRA)\ and the\ Genotype-Tissue Expression (GTEx) project,\ and reprocesses it using a standardized pipeline. This allows for cross-study comparisons and\ meta-analyses, facilitating discoveries in genomics and transcriptomics. Processed recount3 data\ were integrated into the\ Snaptron system\ for indexing and querying data summaries. Recount3 is available\ at: http://rna.recount.bio.\
\\ These tracks display the recount3 intron data, including split read counts and splice junction\ motifs. For hg38, tracks are available for GTEx, TCGA, SRA, and CCLE data sources, while mm10\ includes the SRA track only.\
\ \\ Intron items are colored based on splice junction motifs and read support. Darker colors indicate\ higher read coverage. Split read counts and splice motifs are shown on mouseover.\ By default, only introns with a minimum read count of 10,000 are shown. This threshold can be\ changed on the track configuration page.\
\\ The intron items are color-coded (darker colors indicate higher coverage):\
\\ Introns can be filtered by:\
\\ A distributed processing system for RNA-seq data called Monorail was developed. Using Monorail,\ recount3 processed and summarized 316,443 human and 416,803 mouse RNA-seq run accessions collected\ from the Sequence Read Archive (SRA), with the human runs including large-scale consortia such as\ GTEx v8 and The Cancer Genome Atlas (TCGA).\
\\ Junction files were converted to BED format. For grayscaling total read count was log10\ transformed and multiplied by 10 to get a score between 0 and 225, which can be found\ in the BED score field.\
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ For automated analysis, the data may be queried from our\ REST API.
\\ Please refer to our\ mailing list archives\ for questions or our\ Data Access FAQ\ for more information.\
\\ The original junction files for human can be found at:\
\\ The mouse junction file is available at:\
\ \ \\ Wilks C, Zheng SC, Chen FY, Charles R, Solomon B, Ling JP, Imada EL, Zhang D, Joseph L, Leek JT\ et al.\ \ recount3: summaries and queries for large-scale RNA-seq expression and splicing.\ Genome Biol. 2021 Nov 29;22(1):323.\ PMID: 34844637; PMC: PMC8628444\
\ rna 1 compositeTrack on\ group rna\ html recount3\ longLabel recount3 introns\ noParentConfig on\ shortLabel recount3\ track recount3\ type bigBed 9 +\ visibility hide\ rectumWangCellType Rectum Cells bigBarChart Rectum cells binned by cell type from Wang et al 2020 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-intestine+rectum&gene=$$\ This track shows data from Single-cell transcriptome analysis reveals differential\ nutrient absorption functions in human intestine. Droplet-based\ single-cell RNA sequencing (scRNA-seq) was used to survey gene expression\ profiles of the epithelium in the human ileum, colon, and rectum. A total of 7\ cell clusters were identified: enterocytes (EC), goblet cells (G), paneth-like\ cells (PLC), enteroendocrine cells (EEC), progenitor cells (PRO),\ transient-amplifying cells (TA) and stem cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in rectum\ cells where cells are grouped by cell type\ (Rectum Cells) or donor\ (Rectum Donor). The default track\ displayed is Rectum Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. Note that the Rectum Donor track\ is colored by donor for improved clarity.
\ \\ Using scRNA-seq, RNA profiles of intestinal epithelial cells were obtained for\ 3,898 cells from two human rectum samples. Tissue samples belonged to two\ female donors diagnosed with Adenocarcinoma age 66 (Rectum-1) and age 50\ (Rectum-2). The healthy intestinal mucous membranes used for each sample were\ cut away from the tumor border in surgically removed rectal tissue.\ Additionally, the intestinal tissues were washed in Hank's balanced salt\ solution (HBSS) to remove mucus, blood cells, and muscle tissue. The sample was\ enriched for epithelial cells through centrifugation before being dissociated\ with Tryple to obtain single-cell suspensions. RNA-seq libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina Hiseq X Ten\ PE150.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720\
\ singleCell 1 barChartBars enteroendocrine_cell enterocyte goblet_cell paneth-like_cell progenitor_cell stem_cell transit-amplifying_cell\ barChartColors #c7d2e5 #0198c0 #0251fc #7197d7 #4d689b #9e9fa2 #949dae\ barChartLimit 1.6\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/rectumWang/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/rectumWang/cell_type.bb\ defaultLabelFields name\ html rectumWang\ labelFields name,name2\ longLabel Rectum cells binned by cell type from Wang et al 2020\ parent rectumWang\ shortLabel Rectum Cells\ track rectumWangCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-intestine+rectum&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ rectumWangDonor Rectum Donor bigBarChart Rectum cells binned by organ donor from Wang et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=human-intestine+rectum&gene=$$\ This track shows data from Single-cell transcriptome analysis reveals differential\ nutrient absorption functions in human intestine. Droplet-based\ single-cell RNA sequencing (scRNA-seq) was used to survey gene expression\ profiles of the epithelium in the human ileum, colon, and rectum. A total of 7\ cell clusters were identified: enterocytes (EC), goblet cells (G), paneth-like\ cells (PLC), enteroendocrine cells (EEC), progenitor cells (PRO),\ transient-amplifying cells (TA) and stem cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in rectum\ cells where cells are grouped by cell type\ (Rectum Cells) or donor\ (Rectum Donor). The default track\ displayed is Rectum Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. Note that the Rectum Donor track\ is colored by donor for improved clarity.
\ \\ Using scRNA-seq, RNA profiles of intestinal epithelial cells were obtained for\ 3,898 cells from two human rectum samples. Tissue samples belonged to two\ female donors diagnosed with Adenocarcinoma age 66 (Rectum-1) and age 50\ (Rectum-2). The healthy intestinal mucous membranes used for each sample were\ cut away from the tumor border in surgically removed rectal tissue.\ Additionally, the intestinal tissues were washed in Hank's balanced salt\ solution (HBSS) to remove mucus, blood cells, and muscle tissue. The sample was\ enriched for epithelial cells through centrifugation before being dissociated\ with Tryple to obtain single-cell suspensions. RNA-seq libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina Hiseq X Ten\ PE150.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720\
\ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/rectumWang/donor.colors\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/rectumWang/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/rectumWang/donor.bb\ defaultLabelFields name\ html rectumWang\ labelFields name,name2\ longLabel Rectum cells binned by organ donor from Wang et al 2020\ parent rectumWang\ shortLabel Rectum Donor\ track rectumWangDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=human-intestine+rectum&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ rectumWang Rectum Wang Rectum single cell sequencing from Wang et al 2020 0 100 0 0 0 127 127 127 0 0 0\ This track shows data from Single-cell transcriptome analysis reveals differential\ nutrient absorption functions in human intestine. Droplet-based\ single-cell RNA sequencing (scRNA-seq) was used to survey gene expression\ profiles of the epithelium in the human ileum, colon, and rectum. A total of 7\ cell clusters were identified: enterocytes (EC), goblet cells (G), paneth-like\ cells (PLC), enteroendocrine cells (EEC), progenitor cells (PRO),\ transient-amplifying cells (TA) and stem cells (SC).
\ \\ This track collection contains two bar chart tracks of RNA expression in rectum\ cells where cells are grouped by cell type\ (Rectum Cells) or donor\ (Rectum Donor). The default track\ displayed is Rectum Cells.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| epithelial | |
| secretory | |
| stem cell |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. Note that the Rectum Donor track\ is colored by donor for improved clarity.
\ \\ Using scRNA-seq, RNA profiles of intestinal epithelial cells were obtained for\ 3,898 cells from two human rectum samples. Tissue samples belonged to two\ female donors diagnosed with Adenocarcinoma age 66 (Rectum-1) and age 50\ (Rectum-2). The healthy intestinal mucous membranes used for each sample were\ cut away from the tumor border in surgically removed rectal tissue.\ Additionally, the intestinal tissues were washed in Hank's balanced salt\ solution (HBSS) to remove mucus, blood cells, and muscle tissue. The sample was\ enriched for epithelial cells through centrifugation before being dissociated\ with Tryple to obtain single-cell suspensions. RNA-seq libraries were prepared\ using 10x Genomics 3' v2 kit and sequenced on an Illumina Hiseq X Ten\ PE150.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Yalong Wang, Wanlu Song, and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Luis Nassar. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720\
\ singleCell 0 group singleCell\ longLabel Rectum single cell sequencing from Wang et al 2020\ shortLabel Rectum Wang\ superTrack on\ track rectumWang\ visibility hide\ ucscToRefSeq RefSeq Acc bed 4 RefSeq Accession 0 100 0 0 0 127 127 127 0 0 0 https://www.ncbi.nlm.nih.gov/nuccore/$$\ This track associates UCSC Genome Browser chromosome names to accession\ identifiers from the NCBI Reference Sequence Database (RefSeq).\
\ \\ The data were downloaded from the NCBI assembly database.\
\ \The data for this track was prepared by\ Hiram Clawson.\ map 1 group map\ longLabel RefSeq Accession\ shortLabel RefSeq Acc\ track ucscToRefSeq\ type bed 4\ url https://www.ncbi.nlm.nih.gov/nuccore/$$\ urlLabel RefSeq accession:\ visibility hide\ refSeqFuncElems RefSeq Func Elems bigBed 9 + NCBI RefSeq Functional Elements 0 100 0 0 0 127 127 127 0 0 0
\ NCBI recently announced a new release of\ functional regulatory elements.\ \ NCBI is now providing \ RefSeq and \ Gene\ records for non-genic functional elements that have been described in the literature and are \ experimentally validated. Elements in scope include experimentally-verified gene regulatory \ regions (e.g., enhancers, silencers, locus control regions), known structural elements\ (e.g., insulators, DNase I hypersensitive sites, matrix/scaffold-associated regions), \ well-characterized DNA replication origins, and clinically-significant sites of DNA recombination\ and genomic instability. Priority is given to genomic regions that are implicated in human disease \ or are otherwise of significant interest to the research community. Currently, the scope of this \ project is restricted to human and mouse. The current scope does not include functional elements\ predicted from large-scale epigenomic mapping studies, nor elements based on disease-associated \ variation.
\ \\ Functional elements are colored by Sequence Ontology (SO) term\ using the same scheme as NCBI's Genome Data Viewer:\
\ NCBI manually curated features in accordance with International Nucleotide \ Sequence Database Collaboration (INSDC) standards. Features that are supported by direct \ experimental evidence include at least one experiment qualifier with an evidence code (ECO ID) \ from the Evidence and Conclusion Ontology, and at least one citation from PubMed. Currently\ 971 distinct PubMed citations are included in this track. \
\ \\ This track was made with assistance from\ Terence Murphy at NCBI.
\ \\ The raw data can be explored interactively with the Table Browser, or the Data Integrator. For automated analysis, the data may be \ queried from our REST API,\ and the genome annotations are stored in files that can be downloaded from our \ download server, with more information available on\ our blog.
\ \\ Several new enhancements to the RefSeq Functional Elements dataset are available as a Public Hub.\ The hub can be found on the Public Hub page.\ The track hub was prepared by Dr. Catherine M. Farrell, NCBI/NLM/NIH with further insights discussed\ in a related NCBI blog post.
\ \\ Pruitt KD, Brown GR, Hiatt SM, Thibaud-Nissen F, Astashyn A, Ermolaeva O, Farrell CM, Hart J,\ Landrum MJ, McGarvey KM et al.\ RefSeq: an update on mammalian reference sequences.\ Nucleic Acids Res. 2014 Jan;42(Database issue):D756-63.\ PMID: 24259432; PMC: PMC3965018\
\ \\ Pruitt KD, Tatusova T, Maglott DR.\ NCBI Reference Sequence (RefSeq): a curated non-redundant\ sequence database of genomes, transcripts and proteins.\ Nucleic Acids Res. 2005 Jan 1;33(Database issue):D501-4.\ PMID: 15608248; PMC: PMC539979\
\ regulation 1 bigDataUrl /gbdb/hg38/ncbiRefSeq/refSeqFuncElems.bb\ group regulation\ itemRgb on\ longLabel NCBI RefSeq Functional Elements\ mouseOverField _mouseOver\ noScoreFilter .\ shortLabel RefSeq Func Elems\ track refSeqFuncElems\ type bigBed 9 +\ urls geneIds=https://www.ncbi.nlm.nih.gov/gene?cmd=Retrieve&dopt=full_report&list_uids=$$ pubMedIds=https://www.ncbi.nlm.nih.gov/pubmed/$$ soTerm=http://www.sequenceontology.org/browser/obob.cgi?rm=term_list&release=current_svn&obo_query=$$\ ghGeneHancer Reg Elem bigBed 9 + GeneHancer Regulatory Elements and Gene Interactions 1 100 0 0 0 127 127 127 0 0 0 http://www.genecards.org/Search/Keyword?queryString=$$ regulation 1 exonArrows off\ itemRgb on\ longLabel GeneHancer Regulatory Elements and Gene Interactions\ mouseOverField elementType\ parent geneHancer\ searchIndex name\ shortLabel Reg Elem\ track ghGeneHancer\ type bigBed 9 +\ url http://www.genecards.org/Search/Keyword?queryString=$$\ urlLabel In GeneCards:\ view a_GH\ visibility dense\ ReMap ReMap ChIP-seq bigBed 9 + ReMap Atlas of Regulatory Regions 0 100 0 0 0 127 127 127 0 0 0\ This track represents the ReMap Atlas of regulatory regions, which consists of a\ large-scale integrative analysis of all Public ChIP-seq data for transcriptional\ regulators from GEO, ArrayExpress, and ENCODE. \
\ \\ Below is a schematic diagram of the types of regulatory regions: \
\
\
\ This 4th release of ReMap (2022) presents the analysis of a total of 8,103 \ quality controlled ChIP-seq (n=7,895) and ChIP-exo (n=208) data sets from public\ sources (GEO, ArrayExpress, ENCODE). The ChIP-seq/exo data sets have been mapped\ to the GRCh38/hg38 human assembly. The data set is defined as a ChIP-seq \ experiment in a given series (e.g. GSE46237), for a given TF (e.g. NR2C2), in a\ particular biological condition (i.e. cell line, tissue type, disease state, or\ experimental conditions; e.g. HELA). Data sets were labeled by concatenating\ these three pieces of information, such as GSE46237.NR2C2.HELA. \ \
\Those merged analyses cover a total of 1,211 DNA-binding proteins\ (transcriptional regulators) such as a variety of transcription factors (TFs),\ transcription co-activators (TCFs), and chromatin-remodeling factors (CRFs) for\ 182 million peaks. \
\ \
\
\
\
Public ChIP-seq data sets were extracted from Gene Expression Omnibus (GEO) and\
ArrayExpress (AE) databases. For GEO, the query\
\
'('chip seq' OR 'chipseq' OR\
'chip sequencing') AND 'Genome binding/occupancy profiling by high throughput\
sequencing' AND 'homo sapiens'[organism] AND NOT 'ENCODE'[project]'\
\
was used to return a list of all potential data sets to analyze, which were then manually \
assessed for further analyses. Data sets involving polymerases (i.e. Pol2 and\
Pol3), and some mutated or fused TFs (e.g. KAP1 N/C terminal mutation, GSE27929)\
were excluded.\
\ Available ENCODE ChIP-seq data sets for transcriptional regulators from the\ ENCODE portal were processed with the\ standardized ReMap pipeline. The list of ENCODE data was retrieved as FASTQ files from the\ ENCODE portal\ using the following filters:\
\ Both Public and ENCODE data were processed similarly. Bowtie 2 (PMC3322381) (version 2.2.9) with options -end-to-end -sensitive was used to align all\ reads on the genome. Biological and technical\ replicates for each unique combination of GSE/TF/Cell type or Biological condition\ were used for peak calling. TFBS were identified using MACS2 peak-calling tool\ (PMC3120977) (version 2.1.1.2) in order to follow ENCODE ChIP-seq guidelines,\ with stringent thresholds (MACS2 default thresholds, p-value: 1e-5). An input data\ set was used when available.\
\ \ \\ To assess the quality of public data sets, a score was computed based on the\ cross-correlation and the FRiP (fraction of reads in peaks) metrics developed by\ the ENCODE Consortium (https://genome.ucsc.edu/ENCODE/qualityMetrics.html). Two\ thresholds were defined for each of the two cross-correlation ratios (NSC,\ normalized strand coefficient: 1.05 and 1.10; RSC, relative strand coefficient:\ 0.8 and 1.0). Detailed descriptions of the ENCODE quality coefficients can be\ found at https://genome.ucsc.edu/ENCODE/qualityMetrics.html. The\ phantompeak tools suite was used\ (https://code.google.com/p/phantompeakqualtools/) to compute\ RSC and NSC.\
\\ Please refer to the ReMap 2022, 2020, and 2018 publications for more details\ (citation below).\
\ \ \ \\ ReMap Atlas of regulatory regions data can be explored interactively with the\ Table Browser and cross-referenced with the \ Data Integrator. For programmatic access,\ the track can be accessed using the Genome Browser's\ REST API.\ ReMap annotations can be downloaded from the\ Genome Browser's download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \\ Individual BED files for specific TFs, cells/biotypes, or data sets can be\ found and downloaded on the ReMap website.\
\ \\ Chèneby J, Gheorghe M, Artufel M, Mathelier A, Ballester B.\ \ ReMap 2018: an updated atlas of regulatory regions from an integrative analysis of DNA-binding ChIP-\ seq experiments.\ Nucleic Acids Res. 2018 Jan 4;46(D1):D267-D275.\ PMID: 29126285; PMC: PMC5753247\
\\ Chèneby J, Ménétrier Z, Mestdagh M, Rosnet T, Douida A, Rhalloussi W, Bergon A, Lopez\ F, Ballester B.\ \ ReMap 2020: a database of regulatory regions from an integrative analysis of Human and Arabidopsis\ DNA-binding sequencing experiments.\ Nucleic Acids Res. 2020 Jan 8;48(D1):D180-D188.\ PMID: 31665499; PMC: PMC7145625\
\\ Griffon A, Barbier Q, Dalino J, van Helden J, Spicuglia S, Ballester B.\ \ Integrative analysis of public ChIP-seq experiments reveals a complex multi-cell regulatory\ landscape.\ Nucleic Acids Res. 2015 Feb 27;43(4):e27.\ PMID: 25477382; PMC: PMC4344487\
\\ Hammal F, de Langen P, Bergon A, Lopez F, Ballester B.\ \ ReMap 2022: a database of Human, Mouse, Drosophila and Arabidopsis regulatory regions from an\ integrative analysis of DNA-binding sequencing experiments.\ Nucleic Acids Res. 2022 Jan 7;50(D1):D316-D325.\ PMID: 34751401; PMC: PMC8728178\
\ \ regulation 1 compositeTrack on\ group regulation\ html ../reMap\ longLabel ReMap Atlas of Regulatory Regions\ noParentConfig on\ noScoreFilter on\ shortLabel ReMap ChIP-seq\ track ReMap\ type bigBed 9 +\ visibility hide\ ucscRetroAli9 RetroGenes V9 psl Retroposed Genes V9, Including Pseudogenes 0 100 20 0 250 137 127 252 0 0 0\ Retrotransposition is a process involving the copying of DNA by a group of\ enzymes that have the ability to reverse transcribe spliced mRNAs, and the \ insertion of these processed mRNAs back into the genome resulting\ in single-exon copies of genes and sometime chimeric genes. Retrogenes are \ mostly non-functional pseudogenes but some are functional genes that have \ acquired a promoter from a neighboring gene, or transcribed pseudogenes, and \ some are anti-sense transcripts that may impede mRNA translation.\
\ \\ All mRNAs of a species from GenBank were aligned to the genome using\ lastz\ (Miller lab, Pennsylvania State University). mRNAs that aligned twice in the genome\ (once with introns and once without introns) were initially screened. Next, a series\ of features were scored to determine candidates for retrotransposition events. \ These features included position and length of the polyA tail, percent coverage of the \ retrogene alignment to the parent, degree of synteny with mouse, coverage of repetitive \ elements, number of exons that can still be aligned to the retrogene, number of putative \ introns removed at the retrogene locus and degree of divergence from the parent gene.\ Retrogenes were classified using a threshold score function that is a linear combination \ of this set of features.\ Retrogenes in the final set were selected using a score threshold based on a ROC plot\ against the Vega annotated\ pseudogenes.\
\ \\ Retrogenes inserted into the genome since the mouse/human divergence show a break\ in the human genome syntenic net alignments to the mouse genome. A break in orthology score is \ calculated and weighted before contributing to the final retrogene score. The break in orthology score\ ranges from 0-130 and it represents the portion of the genome that is missing in each species relative\ to the reference genome (human hg38) at the retrogene locus as defined by syntenic\ alignment nets. If the score is 0, there is orthologous DNA and no break in ortholog with the other species; this \ could be an ancient retrogene; duplicated pseudogenes may also score low because they are often generated \ via large segmental duplication events so the size of the pseudogene is small relative to the size of the \ inserted duplicated sequence. Scores greater than 100 represent cases where the retrogene alignment has no \ flanking alignment resulting from an ancient insertion or other complex rearrangement.\
\\ Breaks in orthology with human and dog tend to be due to genomic\ insertions in the rodent lineage so sequence gaps are not treated as orthology breaks. \ Relative orthology of human/mouse and dog/mouse nets are used to avoid false positives due to deletions \ in the human genome. Since older retrogenes will not show a break in orthology, this feature is \ weighted lower than other features when scoring putative retrogenes.\
\ \\ The RetroFinder program and browser track were developed by\ Robert Baertsch at UCSC.\
\ \ \\ Baertsch R, Diekhans M, Kent WJ, Haussler D, Brosius J.\ \ Retrocopy contributions to the evolution of the human genome.\ BMC Genomics. 2008 Oct 8;9:466.\ PMID: 18842134; PMC: PMC2584115\
\ \\ Kent WJ, Baertsch R, Hinrichs A, Miller W, Haussler D.\ \ Evolution's cauldron: duplication, deletion, and rearrangement in the mouse and human genomes.\ Proc Natl Acad Sci U S A. 2003 Sep 30;100(20):11484-9.\ PMID: 14500911; PMC: PMC208784\
\ \\ Pei B, Sisu C, Frankish A, Howald C, Habegger L, Mu XJ, Harte R, Balasubramanian S, Tanzer A,\ Diekhans M et al.\ \ The GENCODE pseudogene resource.\ Genome Biol. 2012 Sep 26;13(9):R51.\ PMID: 22951037; PMC: PMC3491395\
\ \\ Schwartz S, Kent WJ, Smit A, Zhang Z, Baertsch R, Hardison RC, Haussler D, Miller W.\ \ Human-mouse alignments with BLASTZ.\ Genome Res. 2003 Jan;13(1):103-7.\ PMID: 12529312; PMC: PMC430961\
\ \\ Zheng D, Frankish A, Baertsch R, Kapranov P, Reymond A, Choo SW, Lu Y, Denoeud F, Antonarakis SE,\ Snyder M et al.\ \ Pseudogenes in the ENCODE regions: consensus annotation, analysis of transcription, and\ evolution.\ Genome Res. 2007 Jun;17(6):839-51.\ PMID: 17568002; PMC: PMC1891343\
\ genes 1 baseColorDefault diffCodons\ baseColorUseCds table ucscRetroCds9\ baseColorUseSequence extFile ucscRetroSeq9 ucscRetroExtFile9\ color 20,0,250\ dataVersion Jan. 2015\ exonNumbers off\ group genes\ indelDoubleInsert on\ indelQueryInsert on\ longLabel Retroposed Genes V9, Including Pseudogenes\ shortLabel RetroGenes V9\ showCdsAllScales .\ showCdsMaxZoom 10000.0\ showDiffBasesAllScales .\ showDiffBasesMaxZoom 10000.0\ track ucscRetroAli9\ type psl\ ucscRetroInfo ucscRetroInfo9\ visibility hide\ revel REVEL Scores bigWig REVEL Pathogenicity Score for single-base coding mutations (zoom for exact score) 0 100 150 80 200 202 167 227 0 0 0This track collection shows Rare Exome Variant Ensemble Learner (REVEL) scores that can be\ used as evidence for pathogenicity classifications.\
\ \\ REVEL is an ensemble method for predicting a score for missense variants \ based on a combination of scores from 13 individual tools: MutPred, FATHMM v2.3, \ VEST 3.0, PolyPhen-2, SIFT, PROVEAN, MutationAssessor, MutationTaster, LRT, GERP++, \ SiPhy, phyloP, and phastCons. REVEL was trained using recently discovered pathogenic \ and rare neutral missense variants, excluding those previously used to train its \ constituent tools. The REVEL score for an individual missense variant can range \ from 0 to 1, with higher scores reflecting greater likelihood that the variant is \ damaging.\
\ \Most authors of deleteriousness scores argue against using fixed cutoffs in\ diagnostics. But to give an idea of the meaning of the score value, the REVEL\ authors note: "For example, 75.4% of disease mutations but only 10.9% of\ neutral variants (and 12.4% of all ESVs) have a REVEL score above 0.5,\ corresponding to a sensitivity of 0.754 and specificity of 0.891. Selecting a\ more stringent REVEL score threshold of 0.75 would result in higher specificity\ but lower sensitivity, with 52.1% of disease mutations, 3.3% of neutral\ variants, and 4.1% of all ESVs being classified as pathogenic". (Figure S1 of\ the reference below)\
\ \\ There are five subtracks for this track:\
Four lettered subtracks, one for every nucleotide, showing\ scores for the variant from the reference to that\ nucleotide. All subtracks show the REVEL ensemble score on mouseover. Across the exome, \ there are three values per position, one for every possible\ nucleotide variant. The fourth value, "no variant", representing\ the reference allele, e.g. A to A, is always set to zero, "0.0". REVEL only\ takes into account amino acid changes, so a nucleotide variant that predicts no\ amino acid change (synonymous) also receives the score "0.0". \
\ In rare cases, two scores are output for the same variant at a \ genome position. This happens when there are two transcripts with\ distinct splicing patterns and since some input scores for REVEL take into account\ the sequence context, the same variant can get two different scores. In these cases,\ only the maximum score is shown in the four per-nucleotide subtracks. The complete set of \ scores are shown in the Overlaps track.\
\ \One subtrack, Overlaps, shows alternate REVEL scores when applicable. \ In rare cases (0.05% of genome positions), multiple scores exist with a single variant, \ due to multiple, overlapping transcripts. For example, if there are \ two transcripts and one covers only half of an exon, then the amino acids\ that overlap both transcripts will get two distinct REVEL scores, since some of the underlying\ scores (polyPhen for example) take into account the amino acid sequence context and \ this context is different depending on the transcript.\ For these cases, this subtrack contains at least two\ graphical features, for each affected genome position. Each feature is labeled\ with the reference or variant (A, C, T, or G). The transcript IDs and resulting score is\ shown when hovering over the feature or clicking\ it. For the large majority of the genome, this subtrack has no features.\ This is because REVEL usually outputs only a single score per nucleotide and \ most transcript-derived amino acid sequence contexts are identical.\
\\ Note that in most diagnostic testing scenarios, variants are called using WGS\ pipelines, not RNA-seq. As a result, variants are originally located on the\ genome, not on transcripts, and the choice of transcript is made by\ a variant calling software using a heuristic. In addition, clinically, in the\ field, some transcripts have been agreed-on as more relevant for a disease, e.g.\ because only certain transcripts may be expressed in the relevant tissue. So\ the choice of the most relevant transcript, and as such the REVEL score, may be\ a question of manual curation standards rather than a result of the variant itself.\
\\ Note further that these thresholds represent the recommended score\ cutoffs for genes with no Variant Curation Expert Panel (VCEP) rules.\ For genes with published VCEP rules, the VCEP might\ select different thresholds, which are adjusted for the frequency of the\ relevant disorders. These are available in the ClinGen Criteria\ Specification.
\\ When using this track, zoom in until you can see every basepair at the\ top of the display. Otherwise, there are several nucleotides per pixel under \ your mouse cursor and no score will be shown on the mouseover tooltip.\
\ \Track colors
\\ This track is colored according to Table 2 in Pejaver et al. The colors represent the recommended\ ClinGen score cutoffs.\ \
| Range | \Classification | \
|---|---|
| ≥ 0.644 | \Pathogenic supporting | \
| 0.643 - 0.291 | \Neutral | \
| ≤ 0.290 | \Benign supporting | \
\
More details on these scoring ranges can be found in Bergquist et al. Genet Med 2025, Table 2:
\
\
\
For hg38, note that the data were converted from the hg19 data using the UCSC\
liftOver program, by the REVEL authors. This can lead to missing values or\
duplicated values. When a hg38 position is annotated with two scores due to the\
lifting, the authors removed all the scores for this position. They did the same when\
the reference nucleotide has changed from hg19 to hg38. Also, on hg38, the track has\
the
"lifted" icon to indicate\
this. You can double-check if a nucleotide\
position is possibly affected by the lifting procedure by activating the track\
"Hg19 Mapping" under "Mapping and Sequencing".\
\ REVEL scores are available at the \ \ REVEL website. \ The site provides precomputed REVEL scores for all possible human missense variants \ to facilitate the identification of pathogenic variants among the large number of \ rare variants discovered in sequencing studies.\ \
\ \\
The REVEL data on the UCSC Genome Browser can be explored interactively with the\
Table Browser or the\
Data Integrator. The previous overlap bigBed version file is\
available in the\
archives of our downloads server.\
For automated download and analysis, the genome annotation is stored at UCSC in bigWig\
files that can be downloaded from\
our download server.\
The files for this track are called a.bw, c.bw, g.bw, t.bw. Individual\
regions or the genome annotation can be obtained using our tool bigWigToWig,\
which can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.\
The tools can also be used to obtain features confined to given range, e.g.\
\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500 http://hgdownload.soe.ucsc.edu/gbdb/hg38/revel/a.bw stdout\
\
\
\ Data were converted from the files provided on\ the REVEL Downloads website. As with all other tracks,\ a full log of all commands used for the conversion is available in our \ source repository, for hg19 and hg38. The release used for each assembly is shown on the track description page.\
\ \\ Thanks to the REVEL development team for providing precomputed data and fixing duplicated values in the hg38 files.\
\ \\ Ioannidis NM, Rothstein JH, Pejaver V, Middha S, McDonnell SK, Baheti S, \ Musolf A, Li Q, Holzinger E, Karyadi D, et al.\ \ REVEL: An Ensemble Method for Predicting the Pathogenicity of Rare Missense Variants\ Am J Hum Genet. 2016 Oct 6;99(4):877-885.\ PMID: 27666373;\ PMC: PMC5065685\
\ \\ Bergquist T, Stenton SL, Nadeau EAW, Byrne AB, Greenblatt MS, Harrison SM, Tavtigian SV,\ O'Donnell-Luria A, Biesecker LG, Radivojac P et al.\ \ Calibration of additional computational tools expands ClinGen recommendation options for variant\ classification with PP3/BP4 criteria.\ Genet Med. 2025 Mar 10;27(6):101402.\ PMID: 40084623\
\ \ phenDis 0 color 150,80,200\ compositeTrack on\ dataVersion /gbdb/$D/revel/version.txt\ group phenDis\ longLabel REVEL Pathogenicity Score for single-base coding mutations (zoom for exact score)\ origAssembly hg19\ pennantIcon 19.jpg ../goldenPath/help/liftOver.html "lifted from hg19"\ shortLabel REVEL Scores\ track revel\ type bigWig\ visibility hide\ sample_models_view Sample models bigBed Capture long-seq long-read lncRNAs 4 100 0 0 0 127 127 127 0 0 0 rna 1 longLabel Capture long-seq long-read lncRNAs\ noScoreFilter on\ parent clsLongReadRnaTrack\ shortLabel Sample models\ track sample_models_view\ type bigBed\ view sample_models_view\ visibility squish\ scaffolds Scaffolds bed 4 . GRCh38 Defined Scaffold Identifiers 3 100 0 0 0 127 127 127 0 0 0\ This track shows the Genome Reference Consortium (GRC) names for the \ scaffolds in the GRCh38 (hg38) assembly, downloaded from the GRCh38\ acc2name file in GenBank. \
\ map 1 color 0,0,0\ longLabel GRCh38 Defined Scaffold Identifiers\ shortLabel Scaffolds\ superTrack assemblyContainer pack\ track scaffolds\ type bed 4 .\ sgpGene SGP Genes genePred sgpPep SGP Gene Predictions Using Mouse/Human Homology 0 100 0 90 100 127 172 177 0 0 0\ This track shows short nucleotide variants of a few base pairs when aligning\ HPRC genomes to the hg38 reference assembly. The alignment was made with the\ Minigraph-cactus approach described in the references below.\
\ \There are three subtracks in this superTrack:\
\ VCF Decomposition from\ HPRC Pangenome Resources Github:\ "The Raw VCF files contain a site for each bubble in the graph. Nested bubbles will result in\ overlapping sites. The nesting relationships are denoted with the PS (parent snarl), LV (level) and\ AT (allele traversal) tags and need to be taken into account when interpreting the VCF.\ Alternatively, you can use the 'Decomposed VCFs' which have been normalized by using\ vcfbub to 'pop'\ bubbles with alleles larger than 100k and\ vcfwave\ to realign each alt\ (script). Note that in order to reproduce the PanGenie analyses from the papers, you should instead\ use the\ PanGenie HPRC Workflow. This workflow has a\ CHM13 branch to use when working with that reference.\
\ The exact tools and commands used to produce the VCFs are given\ here."
\ \\ The Name of the items are the pair of node labels that denote the site's location\ in the graph, with the '>' and '<' denoting the forward and reverse\ orientation of the node. Mouseover on items in "squish" and "pack" modes shows the items Name and\ Genotypes. Mouseover on items in "full" mode shows Alleles.\ \
\ The Minigraph-Cactus HPRC v1.0 graph was converted to VCF using vg deconstruct.\ This result was further postprocessed using vcfbub to flatten nested sites then\ vcfwave to normalize by realigning alt alleles to the reference. All steps are\ described in Hickey et al 2023. The postprocessing command lines and data can be found on\ Github.\ Finally, the resulting VCF was filtered by length and split into two VCFs using a cutoff of 3bp.\
\ \\ Thanks to Glenn Hickey for providing the HAL file from the HPRC project and for making these VCFs from them.\
\ \\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q,\ Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ PMID: 33177663;\ PMC: PMC7673649;\ DOI: 10.1038/s41586-020-2871-y\
\ \\ Glenn Hickey, Jean Monlong, Jana Ebler, Adam M Novak, Jordan M Eizenga,\ Yan Gao; Human Pangenome Reference Consortium; Tobias Marschall, Heng Li,\ Benedict Paten\ \ Pangenome graph construction from genome alignments with Minigraph-Cactus.\ Nature Biotechnology. 2023 May 10. doi: 10.1038/s41587-023-01793-w.\ PMID: 37165083;\ DOI: 10.1038/s41587-023-01793-w\
\ \\ Paten B, Earl D, Nguyen N, Diekhans M, Zerbino D, Haussler D.\ \ Cactus: Algorithms for genome multiple sequence alignment.\ Genome Res. 2011 Sep;21(9):1512-28.\ PMID: 21665927;\ PMC: PMC3166836;\ DOI: 10.1101/gr.123356.111\
\ \\ Wen-Wei Liao, Mobin Asri, Jana Ebler, ...et al, Heng Lin,\ Benedict Paten\ \ A draft human pangenome reference.\ Nature. 2023 May;617(7960):312-324.\ PMID: 37165242;\ PMC: PMC1017212;\ DOI: 10.1038/s41586-023-05896-x\
\ hprc 0 group hprc\ html hprcVCF\ longLabel Short Variants\ shortLabel Short Variants\ superTrack on\ track hprcVCF\ sibTxGraph SIB Alt-Splicing altGraphX Alternative Splicing Graph from Swiss Institute of Bioinformatics 0 100 0 0 0 127 127 127 0 0 0 http://ccg.vital-it.ch/cgi-bin/tromer/tromergraph2draw.pl?db=hg38&species=H.+sapiens&tromer=$$\ This track shows the graphs constructed by analyzing experimental RNA\ transcripts and serves as basis for the predicted alternative splicing\ transcripts shown in the SIB Genes track. The blocks represent exons; lines\ indicate introns. The graphical display is drawn such that no exons\ overlap, making alternative events easier to view when the track is in full\ display mode and the resolution is set to approximately gene-level.
\Further information on the graphs can be found on the\ Transcriptome \ Web interface.
\ \\ The splicing graphs were generated using a multi-step pipeline: \
\ The SIB Alternative Splicing Graphs track was produced on the Vital-IT high-performance \ computing platform\ using a computational pipeline developed by Christian Iseli with help from\ colleagues at the Ludwig \ Institute for Cancer\ Research and the Swiss \ Institute of Bioinformatics. It is based on data from NCBI RefSeq and GenBank/EMBL. Our\ thanks to the people running these databases and to the scientists worldwide\ who have made contributions to them.
\ rna 1 group rna\ idInUrlSql select name from sibTxGraph where id=%s\ longLabel Alternative Splicing Graph from Swiss Institute of Bioinformatics\ shortLabel SIB Alt-Splicing\ track sibTxGraph\ type altGraphX\ url http://ccg.vital-it.ch/cgi-bin/tromer/tromergraph2draw.pl?db=hg38&species=H.+sapiens&tromer=$$\ urlLabel SIB link:\ visibility hide\ sibGene SIB Genes genePred Swiss Institute of Bioinformatics Gene Predictions from mRNA and ESTs 0 100 195 90 0 225 172 127 0 0 0 http://ccg.vital-it.ch/cgi-bin/tromer/tromer_quick_search_internal.pl?db=hg38&query_str=$$\ The SIB Genes track is a transcript-based set of gene predictions based\ on data from RefSeq and EMBL/GenBank. Genes all have the support of at\ least one GenBank full length RNA sequence, one RefSeq RNA, or one spliced\ EST. The track includes both protein-coding and non-coding transcripts.\ The coding regions are predicted using\ ESTScan.
\ \\ This track in general follows the display conventions for\ gene prediction\ tracks. The exons for putative non-coding genes and untranslated regions \ are represented by relatively thin blocks while those for coding open \ reading frames are thicker.
\\ This track contains an optional codon coloring\ feature that allows users to quickly validate and compare gene predictions.\ To display codon colors, select the genomic codons option from the\ Color track by codons pull-down menu. Go to the\ Coloring Gene Predictions and\ Annotations by Codon page for more information about this feature.
\Further information on the predicted transcripts can be found on the\ Transcriptome Web\ interface.
\ \ \\ The SIB Genes are built using a multi-step pipeline: \
\ The SIB Genes track was produced on the Vital-IT high-performance \ computing platform\ using a computational pipeline developed by Christian Iseli with help from\ colleagues at the Ludwig Institute\ for Cancer\ Research and the Swiss Institute \ of Bioinformatics. It is based on data from NCBI RefSeq and GenBank/EMBL. Our\ thanks to the people running these databases and to the scientists worldwide\ who have made contributions to them.
\ \\ Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Wheeler DL.\ GenBank: update.\ Nucleic Acids Res. 2004 Jan 1;32(Database issue):D23-6.\ PMID: 14681350; PMC: PMC308779\
\ genes 1 color 195,90,0\ group genes\ html ../../sibGene\ longLabel Swiss Institute of Bioinformatics Gene Predictions from mRNA and ESTs\ parent genePredArchive\ shortLabel SIB Genes\ track sibGene\ type genePred\ url http://ccg.vital-it.ch/cgi-bin/tromer/tromer_quick_search_internal.pl?db=hg38&query_str=$$\ urlLabel SIB link:\ visibility hide\ singleCellMerged Single Cell Expression bigBarChart Single cell RNA expression levels cell types from many organs 0 100 0 0 0 127 127 127 0 0 0\
This track displays single-cell data from 12 papers covering 14 organs. Cells are grouped \
together by organ and cell type. The cell types are based on annotations published alongside\
the papers. These were curated at UCSC as much as possible to use the same cell type \
terminologies across papers and organs. In some cases, we merged together small populations\
of cells annotated as distinct and related types into a single type so as to have enough cells \
to call gene expression levels accurate.\
\
The gene expression levels are normalized so that the total level of expression for all genes in a\
single cell or cell type adds up to one million. \
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of \
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ The cell types are colored by which class they belong to according to the following table.\
\\ Please note, the coloring algorithm allows cells that show some mixed characteristics to =\ show blended colors so there will be some color variation within a class. In addition,\ cells with less than 100 transcripts will be a lighter shade and less \ concentrated in color to represent a low number of transcripts. \ \
\
| Color | \Cell classification | \
|---|---|
| neural | |
| adipose | |
| fibroblast | |
| immune | |
| muscle | |
| hepatocyte | |
| trophoblast | |
| secretory | |
| ciliated | |
| epithelial | |
| endothelial | |
| glia | |
| stem cell or progenitor cell |
\ Each organ or tissue was integrated and curated into the Genome Browser indiviually. \ \
\ The raw barChart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array, \ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \\ Many thanks to the data contributing labs for sharing their high quality research. \ Thanks to the Cell Browser team including Matt Speir and Max Haeussler, for their work\ in integratinging these datasets into the Cell Browser. In most cases, their efforts were\ ahead of our own and we could leverage their work making the job much easier. Within the\ Genome Browser group, Jim Kent did the initial wrangling, and Brittney Wick did substantial data\ cleanup and coordination with the labs.\
\ \\ Baron M, Veres A, Wolock SL, Faust AL, Gaujoux R, Vetere A, Ryu JH, Wagner BK, Shen-Orr SS, Klein AM\ et al.\ \ A Single-Cell Transcriptomic Map of the Human and Mouse Pancreas Reveals Inter- and Intra-cell\ Population Structure.\ Cell Syst. 2016 Oct 26;3(4):346-360.e4.\ PMID: 27667365; PMC: PMC5228327
\ \ \ \\ Cao J, O'Day DR, Pliner HA, Kingsley PD, Deng M, Daza RM, Zager MA, Aldinger KA, Blecher-Gonen R,\ Zhang F et al.\ \ A human cell atlas of fetal gene expression.\ Science. 2020 Nov 13;370(6518).\ PMID: 33184181; PMC: PMC7780123\
\ \\ Cao J, Spielmann M, Qiu X, Huang X, Ibrahim DM, Hill AJ, Zhang F, Mundlos S, Christiansen L,\ Steemers FJ et al.\ \ The single-cell transcriptional landscape of mammalian organogenesis.\ Nature. 2019 Feb;566(7745):496-502.\ PMID: 30787437; PMC: PMC6434952\
\ \ \ \\ De Micheli AJ, Spector JA, Elemento O, Cosgrove BD.\ \ A reference single-cell transcriptomic atlas of human skeletal muscle tissue reveals bifurcated\ muscle stem cell populations.\ Skelet Muscle. 2020 Jul 6;10(1):19.\ PMID: 32624006; PMC: PMC7336639
\ \ \ \\ Hao Y, Hao S, Andersen-Nissen E, Mauck WM 3rd, Zheng S, Butler A, Lee MJ, Wilk AJ, Darby C, Zager M\ et al.\ \ Integrated analysis of multimodal single-cell data.\ Cell. 2021 Jun 24;184(13):3573-3587.e29.\ PMID: 34062119; PMC: PMC8238499\
\ \ \ \\ Litviňuková M, Talavera-López C, Maatz H, Reichart D, Worth CL, Lindberg EL, Kanda M,\ Polanski K, Heinig M, Lee M et al.\ \ Cells of the adult human heart.\ Nature. 2020 Dec;588(7838):466-472.\ PMID: 32971526; PMC: PMC7681775\
\ \ \ \ \\ MacParland SA, Liu JC, Ma XZ, Innes BT, Bartczak AM, Gage BK, Manuel J, Khuu N, Echeverri J, Linares\ I et al.\ \ Single cell RNA sequencing of human liver reveals distinct intrahepatic macrophage populations.\ Nat Commun. 2018 Oct 22;9(1):4383.\ PMID: 30348985; PMC: PMC6197289
\ \ \ \ \\ Solé-Boldo L, Raddatz G, Schütz S, Mallm JP, Rippe K, Lonsdorf AS, Rodríguez-Paredes\ M, Lyko F.\ \ Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.\ Commun Biol. 2020 Apr 23;3(1):188.\ PMID: 32327715; PMC: PMC7181753\
\ \ \ \\ Stewart BJ, Ferdinand JR, Young MD, Mitchell TJ, Loudon KW, Riding AM, Richoz N, Frazer GL,\ Staniforth JUL, Vieira Braga FA et al.\ \ Spatiotemporal immune zonation of the human kidney.\ Science. 2019 Sep 27;365(6460):1461-1466.\ PMID: 31604275; PMC: PMC7343525\
\ \ \ \\ Travaglini KJ, Nabhan AN, Penland L, Sinha R, Gillich A, Sit RV, Chang S, Conley SD, Mori Y, Seita J\ et al.\ \ A molecular cell atlas of the human lung from single-cell RNA sequencing.\ Nature. 2020 Nov;587(7835):619-625.\ PMID: 33208946; PMC: PMC7704697\
\ \ \ \\ Velmeshev D, Schirmer L, Jung D, Haeussler M, Perez Y, Mayer S, Bhaduri A, Goyal N, Rowitch DH,\ Kriegstein AR.\ \ Single-cell genomics identifies cell type-specific molecular changes in autism.\ Science. 2019 May 17;364(6441):685-689.\ PMID: 31097668; PMC: PMC7678724\
\ \ \ \\ Vento-Tormo R, Efremova M, Botting RA, Turco MY, Vento-Tormo M, Meyer KB, Park JE, Stephenson E,\ Polański K, Goncalves A et al.\ \ Single-cell reconstruction of the early maternal-fetal interface in humans.\ Nature. 2018 Nov;563(7731):347-353.\ PMID: 30429548\
\ \ \\ Wang Y, Song W, Wang J, Wang T, Xiong X, Qi Z, Fu W, Yang X, Chen YG.\ \ Single-cell transcriptome analysis reveals differential nutrient absorption functions in human\ intestine.\ J Exp Med. 2020 Feb 3;217(2).\ PMID: 31753849; PMC: PMC7041720
\ expression 1 barChartBars Blood_B Blood_CD4_T Blood_CD8_T Blood_DC Blood_Mono Blood_NK Blood_other Blood_other_T Brain_AST-FB Brain_AST-PP Brain_Endothelial Brain_IN-PV Brain_IN-SST Brain_IN-SV2C Brain_IN-VIP Brain_L2/3 Brain_L4 Brain_L5/6 Brain_L5/6-CC Brain_Microglia Brain_Neu-NRGN-I Brain_Neu-NRGN-II Brain_Neu-mat Brain_OPC Brain_Oligodendrocytes Colon_Enteriendocrine Colon_Enterocyte Colon_Goblet Colon_Paneth-like Colon_Progenitor Colon_Stem_Cell Colon_TA Fetal_AFP_ALB_positive_cells Fetal_Acinar_cells Fetal_Adrenocortical_cells Fetal_Amacrine_cells Fetal_Antigen_presenting_cells Fetal_Astrocytes Fetal_Bipolar_cells Fetal_Bronchiolar_and_alveolar_epithelial_cells Fetal_CCL19_CCL21_positive_cells Fetal_CLC_IL5RA_positive_cells Fetal_CSH1_CSH2_positive_cells Fetal_Cardiomyocytes Fetal_Chromaffin_cells Fetal_Ciliated_epithelial_cells Fetal_Corneal_and_conjunctival_epithelial_cells Fetal_Ductal_cells Fetal_ELF3_AGBL2_positive_cells Fetal_ENS_glia Fetal_ENS_neurons Fetal_Endocardial_cells Fetal_Epicardial_fat_cells Fetal_Erythroblasts Fetal_Excitatory_neurons Fetal_Extravillous_trophoblasts Fetal_Ganglion_cells Fetal_Goblet_cells Fetal_Granule_neurons Fetal_Hematopoietic_stem_cells Fetal_Hepatoblasts Fetal_Horizontal_cells Fetal_IGFBP1_DKK1_positive_cells Fetal_Inhibitory_interneurons Fetal_Inhibitory_neurons Fetal_Intestinal_epithelial_cells Fetal_Islet_endocrine_cells Fetal_Lens_fibre_cells Fetal_Limbic_system_neurons Fetal_Lymphatic_endothelial_cells Fetal_Lymphoid_cells Fetal_MUC13_DMBT1_positive_cells Fetal_Megakaryocytes Fetal_Mesangial_cells Fetal_Mesothelial_cells Fetal_Metanephric_cells Fetal_Microglia Fetal_Myeloid_cells Fetal_Neuroendocrine_cells Fetal_Oligodendrocytes Fetal_PAEP_MECOM_positive_cells Fetal_PDE11A_FAM19A2_positive_cells Fetal_PDE1C_ACSM3_positive_cells Fetal_Parietal_and_chief_cells Fetal_Photoreceptor_cells Fetal_Purkinje_neurons Fetal_Retinal_pigment_cells Fetal_Retinal_progenitors_and_Muller_glia Fetal_SATB2_LRRC7_positive_cells Fetal_SKOR2_NPSR1_positive_cells Fetal_SLC24A4_PEX5L_positive_cells Fetal_SLC26A4_PAEP_positive_cells Fetal_STC2_TLX1_positive_cells Fetal_Satellite_cells Fetal_Schwann_cells Fetal_Skeletal_muscle_cells Fetal_Smooth_muscle_cells Fetal_Squamous_epithelial_cells Fetal_Stellate_cells Fetal_Stromal_cells Fetal_Sympathoblasts Fetal_Syncytiotrophoblasts_and_villous_cytotrophoblasts Fetal_Thymic_epithelial_cells Fetal_Thymocytes Fetal_Trophoblast_giant_cells Fetal_Unipolar_brush_cells Fetal_Ureteric_bud_cells Fetal_Vascular_endothelial_cells Fetal_Visceral_neurons Heart_Adipocytes Heart_Atrial_Cardiomyocyte Heart_Endothelial Heart_Fibroblast Heart_Lymphoid Heart_Mesothelial Heart_Myeloid Heart_Neuronal Heart_NotAssigned Heart_Pericytes Heart_Smooth_muscle_cells Heart_Ventricular_Cardiomyocyte Heart_doublets Ileum_Enteriendocrine Ileum_Enterocyte Ileum_Goblet Ileum_Paneth-like Ileum_Progenitor Ileum_Stem_Cell Ileum_TA Kidney_Ascending_vasa_recta_endothelium Kidney_B_cell Kidney_CD4_T_cell Kidney_CD8_T_cell Kidney_Connecting_tubule Kidney_Descending_vasa_recta_endothelium Kidney_Epithelial_progenitor_cell Kidney_Fibroblast Kidney_Glomerular_endothelium Kidney_Intercalated_cell Kidney_MNP Kidney_NK_cell Kidney_Other_immune Kidney_Pelvic_epithelium Kidney_Peritubular_capillary_endothelium Kidney_Podocyte Kidney_Principal_cell Kidney_Proximal_tubule Kidney_Thick_ascending_limb_of_Loop_of_Henle Kidney_Transitional_urothelium Liver_B_cell Liver_Cholangiocyte Liver_Erythroid Liver_Hepatocyte Liver_Inflammatory_Macs Liver_LSEC_1 Liver_LSEC_2,3 Liver_NK-like Liver_Non-inflammatory_Macs Liver_Plasma Liver_Portal_endothelial Liver_Stellate Liver_abT_cell Liver_gdT_cell_1 Liver_gdT_cell_2 Lung_Airway_Smooth_Muscle Lung_Alveolar_Epithelial_Type_1 Lung_Alveolar_Epithelial_Type_2 Lung_Artery_or_Vein Lung_Basal Lung_Basophil/Mast Lung_Bronchial_Vessel Lung_Capillary Lung_Ciliated Lung_Club Lung_Dendretic Lung_Fibroblast Lung_Goblet Lung_Lymphatic Lung_Lymphocyte Lung_Macrophage_or_Monocyte Lung_Mucous Lung_Other/Rare Lung_Pericyte Lung_Vascular_Smooth_Muscle Muscle_ACTA1+_Mature_skeletal_muscle Muscle_ACTA2+_MYH11+_MYL9+_Smooth_muscle_cells Muscle_APOD+_CFD+_PLAC9+_Adipocytes Muscle_C1QA+_CD74+_Macrophages Muscle_CD36+_VWF+_Platelets Muscle_CLDN5+_PECAM1+_Endothelial Muscle_COL1A1+_Fibroblasts Muscle_DCN+_GSN+_MYOC+_Fibroblasts Muscle_FBN1+_MFAP5+_CD55+_Fibroblasts Muscle_HBA1+_Erythroblasts Muscle_ICAM1+_SELE+_VCAM1+_Endothelial Muscle_IL7R+_PTPRC+_NKG7+_B/T/NK_cells Muscle_PAX7+_DLK1+_MuSCs_and_progenitors Muscle_PAX7low_MYF5+_MuSCs_and_progenitors Muscle_RGS5+_MYL9+_Pericytes Muscle_S100A9+_LYZ+_Inflammatory_macrophages Pancreas_acinar Pancreas_activated_stellate Pancreas_alpha Pancreas_beta Pancreas_delta Pancreas_ductal Pancreas_endothelial Pancreas_epsilon Pancreas_gamma Pancreas_other Pancreas_quiescent_stellate Placenta_CD4+_T Placenta_CD8+_T Placenta_EVT Placenta_Endo Placenta_MAIT Placenta_Myeloid Placenta_NK Placenta_Other_immune Placenta_SCT Placenta_VCT Placenta_dP Placenta_dS Placenta_fFB Rectum_Enteriendocrine Rectum_Enterocyte Rectum_Goblet Rectum_Paneth-like Rectum_Progenitor Rectum_Stem_Cell Rectum_TA Skin_Diff._Keratinocytes Skin_EpSC_and_undiff._progenitors Skin_Erythrocytes Skin_Lymphatic_EC Skin_Macrophages+DC Skin_Melanocytes Skin_Mesenchymal Skin_Pericytes Skin_Pro-inflammatory Skin_Secretory-papilliary Skin_Secretory-reticular Skin_T_cells Skin_Vascular_EC\ barChartColors #fe3247 #fe3248 #fe3248 #e92812 #e02900 #fb2e3e #f01111 #fe3247 #81ce00 #81cd00 #01c000 #ebbf00 #ebbf00 #eabe00 #ebbf00 #ecbf00 #ecbf00 #ecbf00 #edbf00 #ef1211 #c8b701 #c5b701 #ebbf00 #c5be01 #86c601 #c7d2e5 #0198c0 #0251fc #7197d7 #4d689b #9e9fa2 #949dae #c75cc6 #3259c7 #7d8952 #d3ac19 #de201f #adb119 #be9c2d #577881 #a4a096 #b787ac #9275da #af1ea8 #aa973d #477f92 #65b5cb #2f5cc6 #c471c0 #80c709 #cba81f #489338 #fe8839 #8a7352 #e1b60c #5f37bb #ddb311 #305cc5 #deb410 #ad4e3b #b001af #b99b2f #7c7062 #deb40f #e7ba08 #536a95 #3f61b4 #ad9f9a #e1b60d #0aba08 #d02b29 #4766a4 #8b6651 #82953b #d07f49 #8c9840 #d92422 #e31b1b #6c7676 #bca424 #756d72 #b39635 #999eaa #2b59cd #ae9537 #dcb212 #88775c #b09f2b #dbc46b #dcb212 #dab014 #c9c6b4 #618237 #8d656b #80c60a #b80db6 #8d5675 #2889a7 #838546 #809836 #958951 #79785f #87a9b4 #b5443b #5425d7 #d7b015 #507093 #12b50d #c9a721 #f1803d #c1229a #07bc02 #b5562a #eb1613 #1494b3 #de2b02 #e6af0e #c12792 #c15f4e #b06a5a #c1229b #d69f85 #bcd0f3 #0198c0 #568bfd #629be4 #436ca1 #9ea0a1 #919eb1 #5bd05a #ec374a #f7354b #f7354b #5f66ed #5fcd5b #60afce #c98b6b #0ab707 #181dda #de2a02 #f1374b #e7a69c #5cb6cf #05bb04 #9f968b #6496d4 #0e0ceb #181cd9 #bfd7e4 #f1798a #908ffd #d3c4db #af01af #d42c0d #5e97d5 #5d8fe8 #f0798a #e3725c #c27d9a #58d05c #e7cbbe #e93650 #e87a8c #cc7d95 #be04bb #905d31 #0695bc #339a1b #4a4eb4 #c82c38 #c74050 #04bd03 #0371d4 #1451e7 #e41819 #af5022 #0950f5 #ab435d #fb344b #df2901 #2652d0 #3b4ebb #a05331 #bd05b9 #d55acd #bb1b98 #fd8738 #da2f08 #b6513e #11b606 #b65928 #b35024 #b25023 #cf8b7e #419916 #fc344a #d33e3f #98672c #1dad0c #dc2c04 #0d55e6 #c68c6e #2a58bc #1754d9 #2457c4 #0298be #57d457 #c2cfe7 #7290d0 #f9b9b9 #c58c6e #f63247 #fa3248 #6026c2 #06bb03 #f73247 #de2903 #f03142 #ee1313 #5823d1 #5923cf #a1288a #be03bb #af4f22 #c7d2e5 #0198c0 #0251fc #7197d7 #4d689b #9e9fa2 #949dae #0298be #1293ac #b1987c #4b9021 #df2a01 #62b7c6 #9e5d22 #3d9c12 #aa5421 #ac5321 #ad5221 #fa3549 #05bd02\ barChartFacets organ,cell_class,stage,cell_type\ barChartLimit 100\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/singleCellMerged/singleCellMerged.stats\ barChartStretchToItem on\ barChartUnit ppm/cell\ bigDataUrl /gbdb/hg38/bbi/singleCellMerged/singleCellMerged.bb\ configureByPopup off\ defaultLabelFields name\ group expression\ labelFields name,name2\ longLabel Single cell RNA expression levels cell types from many organs\ maxItems 200\ shortLabel Single Cell Expression\ track singleCellMerged\ transformFunc NONE\ type bigBarChart\ visibility hide\ bismapBigBed Single-read mappability bigBed 6 Single-read and multi-read mappability after bisulfite conversion 1 100 0 0 0 127 127 127 0 0 0 map 1 longLabel Single-read and multi-read mappability after bisulfite conversion\ parent bismap\ shortLabel Single-read mappability\ track bismapBigBed\ type bigBed 6\ view SR\ visibility dense\ skinSoleBoldoAge Skin Age bigBarChart Skin single cell RNA binned by skin donor's age from Sole-Boldo et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ This track displays data from Single-cell transcriptomes of the human skin reveal\ age-related loss of fibroblast priming. Single cell RNA sequencing (scRNA-seq) \ was performed on sun-protected skin samples prepared using droplet-sequencing \ (drop-seq). RNA profiles were generated for 15,457 cells after quality control \ and subsequent clustering identified 17 clusters with distinct expression profiles\ as found in Solé-Boldo et al., 2020. \
\ \\ This track collection contains four bar chart tracks of RNA expression in the\ human skin where cells are grouped by cell type \ (Skin Cell), age \ (Skin Age),\ donor \ (Skin Donor), and cell type and donor's age \ (Skin Cell+Age). The default\ track displayed is Skin Cell.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Skin Cell subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy skin samples were obtained from whole-skin specimens belonging to 5\ male donors (ages 25-70) with fair skin. Donors underwent full body skin\ examinations by a dermatologist and medical records were checked for skin\ diseases and/or comorbidities that affect the skin. 4-mm punch biopsies were\ taken from surgically removed skin belonging to the inguinal region of the body\ also known as the groin. Skin samples were kept in MACS Tissue Storage Solution\ for less than 1 hour to avoid necrosis and apoptosis. Enzymatical and\ mechanical dissociation was done using the Miltenyi Biotec Whole Skin\ Dissociation kit for human material and the Miltenyi Biotec Gentle MACS\ dissociator. Drop-seq libraries were prepared using a 10x Genomics 3' v2 kit\ and sequenced on an Illumina HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Llorenç Solé-Boldo and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Solé-Boldo L, Raddatz G, Schütz S, Mallm JP, Rippe K, Lonsdorf AS, Rodríguez-Paredes\ M, Lyko F.\ \ Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.\ Commun Biol. 2020 Apr 23;3(1):188.\ PMID: 32327715; PMC: PMC7181753\
\ \ \ singleCell 1 barChartBars OLD YOUNG\ barChartColors #4c8c2c #877227\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/skinSoleBoldo/age.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/skinSoleBoldo/age.bb\ defaultLabelFields name\ html skinSoleBoldo\ labelFields name,name2\ longLabel Skin single cell RNA binned by skin donor's age from Sole-Boldo et al 2020\ parent skinSoleBoldo\ shortLabel Skin Age\ track skinSoleBoldoAge\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ skinSoleBoldoCellType Skin Cell bigBarChart Skin single cell RNA binned by cell type from Sole-Boldo et al 2020 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ This track displays data from Single-cell transcriptomes of the human skin reveal\ age-related loss of fibroblast priming. Single cell RNA sequencing (scRNA-seq) \ was performed on sun-protected skin samples prepared using droplet-sequencing \ (drop-seq). RNA profiles were generated for 15,457 cells after quality control \ and subsequent clustering identified 17 clusters with distinct expression profiles\ as found in Solé-Boldo et al., 2020. \
\ \\ This track collection contains four bar chart tracks of RNA expression in the\ human skin where cells are grouped by cell type \ (Skin Cell), age \ (Skin Age),\ donor \ (Skin Donor), and cell type and donor's age \ (Skin Cell+Age). The default\ track displayed is Skin Cell.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Skin Cell subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy skin samples were obtained from whole-skin specimens belonging to 5\ male donors (ages 25-70) with fair skin. Donors underwent full body skin\ examinations by a dermatologist and medical records were checked for skin\ diseases and/or comorbidities that affect the skin. 4-mm punch biopsies were\ taken from surgically removed skin belonging to the inguinal region of the body\ also known as the groin. Skin samples were kept in MACS Tissue Storage Solution\ for less than 1 hour to avoid necrosis and apoptosis. Enzymatical and\ mechanical dissociation was done using the Miltenyi Biotec Whole Skin\ Dissociation kit for human material and the Miltenyi Biotec Gentle MACS\ dissociator. Drop-seq libraries were prepared using a 10x Genomics 3' v2 kit\ and sequenced on an Illumina HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Llorenç Solé-Boldo and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Solé-Boldo L, Raddatz G, Schütz S, Mallm JP, Rippe K, Lonsdorf AS, Rodríguez-Paredes\ M, Lyko F.\ \ Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.\ Commun Biol. 2020 Apr 23;3(1):188.\ PMID: 32327715; PMC: PMC7181753\
\ \ \ singleCell 1 barChartBars keratinocyte epidermal_stem_(EpSC)_and__progenitor_cell erythrocyte endothelial_lymphatic_cell macrophage/dendritic_cell melanocyte fibroblast_(mesenchymal) pericyte fibroblast_(pro-inflammatory) fibroblast_(secretory-papilliary) fibroblast_(secretory-reticular) T_cell endothelial_vascular_cell\ barChartColors #0298be #1293ac #b1987c #4b9021 #df2a01 #62b7c6 #9e5d22 #3d9c12 #aa5421 #ac5321 #ad5221 #fa3549 #05bd02\ barChartLimit 4\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/skinSoleBoldo/cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/skinSoleBoldo/cell_type.bb\ defaultLabelFields name\ html skinSoleBoldo\ labelFields name,name2\ longLabel Skin single cell RNA binned by cell type from Sole-Boldo et al 2020\ parent skinSoleBoldo\ shortLabel Skin Cell\ track skinSoleBoldoCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ skinSoleBoldoAgeCellType Skin Cell+Age bigBarChart Skin single cell RNA binned by cell type and donor's age from Sole-Boldo et all 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ This track displays data from Single-cell transcriptomes of the human skin reveal\ age-related loss of fibroblast priming. Single cell RNA sequencing (scRNA-seq) \ was performed on sun-protected skin samples prepared using droplet-sequencing \ (drop-seq). RNA profiles were generated for 15,457 cells after quality control \ and subsequent clustering identified 17 clusters with distinct expression profiles\ as found in Solé-Boldo et al., 2020. \
\ \\ This track collection contains four bar chart tracks of RNA expression in the\ human skin where cells are grouped by cell type \ (Skin Cell), age \ (Skin Age),\ donor \ (Skin Donor), and cell type and donor's age \ (Skin Cell+Age). The default\ track displayed is Skin Cell.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Skin Cell subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy skin samples were obtained from whole-skin specimens belonging to 5\ male donors (ages 25-70) with fair skin. Donors underwent full body skin\ examinations by a dermatologist and medical records were checked for skin\ diseases and/or comorbidities that affect the skin. 4-mm punch biopsies were\ taken from surgically removed skin belonging to the inguinal region of the body\ also known as the groin. Skin samples were kept in MACS Tissue Storage Solution\ for less than 1 hour to avoid necrosis and apoptosis. Enzymatical and\ mechanical dissociation was done using the Miltenyi Biotec Whole Skin\ Dissociation kit for human material and the Miltenyi Biotec Gentle MACS\ dissociator. Drop-seq libraries were prepared using a 10x Genomics 3' v2 kit\ and sequenced on an Illumina HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Llorenç Solé-Boldo and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Solé-Boldo L, Raddatz G, Schütz S, Mallm JP, Rippe K, Lonsdorf AS, Rodríguez-Paredes\ M, Lyko F.\ \ Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.\ Commun Biol. 2020 Apr 23;3(1):188.\ PMID: 32327715; PMC: PMC7181753\
\ \ \ singleCell 1 barChartBars Diff_Keratinocytes_OLD Diff_Keratinocytes_YOUNG EpSC_and_undiff_progenitors_OLD EpSC_and_undiff_progenitors_YOUNG Erythrocytes_OLD Erythrocytes_YOUNG Lymphatic_EC_OLD Lymphatic_EC_YOUNG Macrophages+DC_OLD Macrophages+DC_YOUNG Melanocytes_OLD Melanocytes_YOUNG Mesenchymal_OLD Mesenchymal_YOUNG Pericytes_OLD Pericytes_YOUNG Pro-inflammatory_OLD Pro-inflammatory_YOUNG Secretory-papilliary_OLD Secretory-papilliary_YOUNG Secretory-reticular_OLD Secretory-reticular_YOUNG T_cells_OLD T_cells_YOUNG Vascular_EC_OLD Vascular_EC_YOUNG\ barChartColors #0298be #0597bb #0f94ae #1c90a0 #c8bca7 #b1987c #499026 #b8ca9b #dd2b01 #dd2b02 #60b8c8 #9ccdd1 #bf916d #976222 #23ab0b #519018 #a95422 #a75622 #ac5221 #a55822 #ad5221 #ab5322 #ec8181 #fa3649 #09ba03 #0eb705\ barChartLimit 4\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/skinSoleBoldo/age_cell_type.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/skinSoleBoldo/age_cell_type.bb\ defaultLabelFields name\ html skinSoleBoldo\ labelFields name,name2\ longLabel Skin single cell RNA binned by cell type and donor's age from Sole-Boldo et all 2020\ parent skinSoleBoldo\ shortLabel Skin Cell+Age\ track skinSoleBoldoAgeCellType\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ skinSoleBoldoDonor Skin Donor bigBarChart Skin single cell RNA binned by skin donor from Sole-Boldo et al 2020 0 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ This track displays data from Single-cell transcriptomes of the human skin reveal\ age-related loss of fibroblast priming. Single cell RNA sequencing (scRNA-seq) \ was performed on sun-protected skin samples prepared using droplet-sequencing \ (drop-seq). RNA profiles were generated for 15,457 cells after quality control \ and subsequent clustering identified 17 clusters with distinct expression profiles\ as found in Solé-Boldo et al., 2020. \
\ \\ This track collection contains four bar chart tracks of RNA expression in the\ human skin where cells are grouped by cell type \ (Skin Cell), age \ (Skin Age),\ donor \ (Skin Donor), and cell type and donor's age \ (Skin Cell+Age). The default\ track displayed is Skin Cell.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Skin Cell subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy skin samples were obtained from whole-skin specimens belonging to 5\ male donors (ages 25-70) with fair skin. Donors underwent full body skin\ examinations by a dermatologist and medical records were checked for skin\ diseases and/or comorbidities that affect the skin. 4-mm punch biopsies were\ taken from surgically removed skin belonging to the inguinal region of the body\ also known as the groin. Skin samples were kept in MACS Tissue Storage Solution\ for less than 1 hour to avoid necrosis and apoptosis. Enzymatical and\ mechanical dissociation was done using the Miltenyi Biotec Whole Skin\ Dissociation kit for human material and the Miltenyi Biotec Gentle MACS\ dissociator. Drop-seq libraries were prepared using a 10x Genomics 3' v2 kit\ and sequenced on an Illumina HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Llorenç Solé-Boldo and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Solé-Boldo L, Raddatz G, Schütz S, Mallm JP, Rippe K, Lonsdorf AS, Rodríguez-Paredes\ M, Lyko F.\ \ Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.\ Commun Biol. 2020 Apr 23;3(1):188.\ PMID: 32327715; PMC: PMC7181753\
\ \ \ singleCell 1 barChartBars S1 S2 S3 S4 S5\ barChartColors #6d8120 #916a2a #479220 #1294aa #8f6622\ barChartLimit 2\ barChartMetric mean\ barChartStatsUrl /gbdb/hg38/bbi/skinSoleBoldo/donor.stats\ barChartUnit UMI/cell\ bigDataUrl /gbdb/hg38/bbi/skinSoleBoldo/donor.bb\ defaultLabelFields name\ html skinSoleBoldo\ labelFields name,name2\ longLabel Skin single cell RNA binned by skin donor from Sole-Boldo et al 2020\ parent skinSoleBoldo\ shortLabel Skin Donor\ track skinSoleBoldoDonor\ transformFunc NONE\ type bigBarChart\ url https://cells.ucsc.edu/?ds=aging-human-skin&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ skinSoleBoldo Skin Sole-Boldo Skin single cell data from Sole-Boldo et al 2020 0 100 0 0 0 127 127 127 0 0 0\ This track displays data from Single-cell transcriptomes of the human skin reveal\ age-related loss of fibroblast priming. Single cell RNA sequencing (scRNA-seq) \ was performed on sun-protected skin samples prepared using droplet-sequencing \ (drop-seq). RNA profiles were generated for 15,457 cells after quality control \ and subsequent clustering identified 17 clusters with distinct expression profiles\ as found in Solé-Boldo et al., 2020. \
\ \\ This track collection contains four bar chart tracks of RNA expression in the\ human skin where cells are grouped by cell type \ (Skin Cell), age \ (Skin Age),\ donor \ (Skin Donor), and cell type and donor's age \ (Skin Cell+Age). The default\ track displayed is Skin Cell.
\ \\ The cell types are colored by which class they belong to according to the following table.
\ \\
| Color | \Cell classification | \
|---|---|
| fibroblast | |
| immune | |
| epithelial | |
| endothelial |
\ Cells that fall into multiple classes will be colored by blending the colors associated\ with those classes. The colors will be purest in the\ Skin Cell subtrack, where\ the bars represent relatively pure cell types. They can give an overview of the\ cell composition within other categories in other subtracks as well.
\ \\ Healthy skin samples were obtained from whole-skin specimens belonging to 5\ male donors (ages 25-70) with fair skin. Donors underwent full body skin\ examinations by a dermatologist and medical records were checked for skin\ diseases and/or comorbidities that affect the skin. 4-mm punch biopsies were\ taken from surgically removed skin belonging to the inguinal region of the body\ also known as the groin. Skin samples were kept in MACS Tissue Storage Solution\ for less than 1 hour to avoid necrosis and apoptosis. Enzymatical and\ mechanical dissociation was done using the Miltenyi Biotec Whole Skin\ Dissociation kit for human material and the Miltenyi Biotec Gentle MACS\ dissociator. Drop-seq libraries were prepared using a 10x Genomics 3' v2 kit\ and sequenced on an Illumina HiSeq4000.
\ \The cell/gene matrix and cell-level metadata was downloaded from the \ UCSC Cell Browser.\ The UCSC command line utility matrixClusterColumns, matrixToBarChart, and bedToBigBed were used\ to transform these into a bar chart format bigBed file that can be visualized. The coloring \ was done by defining colors for the broad level cell classes and then using another UCSC utility,\ hcaColorCells, to interpolate the colors across all cell types. The UCSC utilities can be found on\ our download server.
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\ \\ Thanks to Llorenç Solé-Boldo and to the many authors who worked on\ producing and publishing this data set. The data were integrated into the UCSC\ Genome Browser by Jim Kent and Brittney Wick then reviewed by Gerardo Perez. The \ UCSC work was paid for by the Chan Zuckerberg Initiative.
\ \\ Solé-Boldo L, Raddatz G, Schütz S, Mallm JP, Rippe K, Lonsdorf AS, Rodríguez-Paredes\ M, Lyko F.\ \ Single-cell transcriptomes of the human skin reveal age-related loss of fibroblast priming.\ Commun Biol. 2020 Apr 23;3(1):188.\ PMID: 32327715; PMC: PMC7181753\
\ \ \ singleCell 0 group singleCell\ longLabel Skin single cell data from Sole-Boldo et al 2020\ pennantIcon 19.jpg liftover.html "lifted from hg19"\ shortLabel Skin Sole-Boldo\ superTrack on\ track skinSoleBoldo\ visibility hide\ gnomADPextSkin_NotSunExposed_Suprapubic Skin-Not Sun Exposed (Suprapubic) bigWig 0 1 gnomAD pext Skin-Not Sun Exposed (Suprapubic) 0 100 0 0 255 127 127 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Skin_NotSunExposed_Suprapubic.bw\ color 0,0,255\ longLabel gnomAD pext Skin-Not Sun Exposed (Suprapubic)\ parent gnomadPext off\ shortLabel Skin-Not Sun Exposed (Suprapubic)\ track gnomADPextSkin_NotSunExposed_Suprapubic\ visibility hide\ gnomADPextSkin_SunExposed_Lowerleg Skin-Sun Exposed (Lowerleg) bigWig 0 1 gnomAD pext Skin-Sun Exposed (Lowerleg) 0 100 119 119 255 187 187 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Skin_SunExposed_Lowerleg.bw\ color 119,119,255\ longLabel gnomAD pext Skin-Sun Exposed (Lowerleg)\ parent gnomadPext off\ shortLabel Skin-Sun Exposed (Lowerleg)\ track gnomADPextSkin_SunExposed_Lowerleg\ visibility hide\ gnomADPextSmallIntestine_TerminalIleum Small Intestine-Terminal Ileum bigWig 0 1 gnomAD pext Small Intestine-Terminal Ileum 0 100 85 85 34 170 170 144 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/SmallIntestine_TerminalIleum.bw\ color 85,85,34\ longLabel gnomAD pext Small Intestine-Terminal Ileum\ parent gnomadPext off\ shortLabel Small Intestine-Terminal Ileum\ track gnomADPextSmallIntestine_TerminalIleum\ visibility hide\ wgRna sno/miRNA bed 8 + C/D and H/ACA Box snoRNAs, scaRNAs, and microRNAs from snoRNABase and miRBase 0 100 200 80 0 227 167 127 0 0 0 http://www-snorna.biotoul.fr/plus.php?id=$$\ This track displays positions of four different types of RNA in the human \ genome: \
\ C/D box and H/ACA box snoRNAs are guides for the 2'O-ribose methylation and \ the pseudouridilation, respectively, of rRNAs and snRNAs, although many of \ them have no documented target RNA. The scaRNAs guide modifications of the\ spliceosomal snRNAs transcribed by RNA polymerase II, and often contain both \ C/D and H/ACA domains.
\ \\ This track follows the general display conventions for \ gene prediction \ tracks.
\\ The miRNA precursor forms (pre-miRNA) are represented by red blocks.
\\ C/D box snoRNAs, H/ACA box snoRNAs and scaRNAs are represented by blue, \ green and magenta blocks, respectively. At a zoomed-in resolution, arrows \ superimposed on the blocks indicate the sense orientation of the snoRNAs.
\ \\ Precursor miRNA genomic locations from\ \ miRBase\ were calculated using wublastn for sequence alignment with the requirement of\ 100% identity. \ The extents of the precursor sequences were not generally known and were\ predicted based on base-paired hairpin structure. miRBase is\ described in Griffiths-Jones, S. (2004) and Weber, M.J. (2005) in the \ References section below.
\\ The snoRNAs and scaRNAs from the snoRNABase were aligned against the \ human genome using blat. \
\ \\ \
\ When making use of these data, please cite the folowing articles in addition to\ the primary sources of the miRNA sequences:
\\ Griffiths-Jones S, Saini HK, van Dongen S, Enright AJ.\ miRBase: tools for microRNA genomics.\ Nucleic Acids Res. 2008 Jan 1;36(Database issue):D154-8.
\\ Griffiths-Jones S, Grocock RJ, van Dongen S, Bateman A, Enright AJ.\ miRBase: microRNA sequences, targets and gene nomenclature.\ Nucleic Acids Res. 2006 Jan 1;34(Database issue):D140-4.
\\ Griffiths-Jones S.\ The microRNA Registry.\ Nucleic Acids Res. 2004 Jan 1;32(Database issue):D109-11.
\\ Weber MJ.\ New human and mouse microRNA genes found by homology search.\
\ You may also want to cite The Wellcome Trust Sanger Institute \ miRBase and The Laboratoire de Biologie Moleculaire \ Eucaryote snoRNABase.
\\ The following publication provides guidelines on miRNA annotation:\ Ambros V. et al., \ A uniform system for microRNA annotation. \ RNA. 2003;9(3):277-9.
\\ genes 1 color 200,80,0\ dataVersion miRBase Release 22 (March 2018) and snoRNABase Version 3 (lifted from hg19)\ group genes\ longLabel C/D and H/ACA Box snoRNAs, scaRNAs, and microRNAs from snoRNABase and miRBase\ noScoreFilter .\ shortLabel sno/miRNA\ superTrack nonCodingRNAs pack\ track wgRna\ type bed 8 +\ url http://www-snorna.biotoul.fr/plus.php?id=$$\ url2 http://www.mirbase.org/cgi-bin/query.pl?terms=$$\ url2Label miRBase:\ urlLabel Laboratoire de Biologie Moleculaire Eucaryote:\ visibility hide\ snpedia SNPedia bed 4 SNPedia 0 100 50 0 100 152 127 177 0 0 0
\ SNPedia is a wiki investigating human\ genetics with information about the effects of variations in DNA, citing\ peer-reviewed scientific publications.\ \
\ The track "SNPedia all" shows all SNPs that exist as a page in \ SNPedia.com. As SNPedia's user collaboration grows, more \ detail will be added to SNPedia.com pages. For now, most of the pages are auto-generated by bots \ and have empty pages. According to Mike Carioso (SNPedia.com founder), SNPedia entries are mostly \ ClinVar entries marked as pathogenic with at least 4 stars as defined by the\ \ ClinVar review status. \
\ \\ The track "SNPedia with text" is a subset of the "SNPedia all" track. This track \ displays only SNPedia entries with a text page that was created manually by a user who typed in \ some text (approximately 5,000 entries). In the browser, click on the "configure" button\ and select "next/previous item navigation" to show clickable arrows in the browser which\ will jump to the next or previous item.\
\\ Clicks on the features show the text from the SNPedia.com page and a link to the original page.\
\ \\ Genomic locations of SNPedia entries are labeled with the dbSNP ID.\
\ \\ In the track "SNPedia all SNPs", the features are colored based on the SNPedia microarray \ annotation: grey for SNPs that are on no microarray, dark blue for Affymetrix, dark purple for \ Illumina and black for features on both arrays.\
\ \\ The mappings displayed in this track were used as provided in the SNPedia GFF file.\ For the "SNPedia with text" track, all SNPedia pages were downloaded and their content \ checked with a script that tries to remove pages that were auto-generated and not created manually \ by a user.\
\ \\ Thanks to Mike Cariaso for help with the GFF download and Max Haeussler at UCSC for building this \ track.\
\ \Cariaso Michael; Lennon Greg. \ \ SNPedia: a wiki supporting personal genome annotation, interpretation and analysis. \ Nucleic acids research. 2012 40Database issue:D1308-12.\ PMID: 22140107; \ PMC: \ PMC3245045
\ \ phenDis 1 color 50,0,100\ compositeTrack on\ group phenDis\ longLabel SNPedia\ shortLabel SNPedia\ track snpedia\ type bed 4\ visibility hide\ varFreqs SNV Frequencies bed 12 SNV Frequencies from various cohorts or national projects 0 100 0 0 0 127 127 127 0 0 0\ This track collection gathers variant allele frequencies from population-scale sequencing\ and genotyping projects worldwide, from a total of ~1.7 million genomes/exomes/arrays.\ Unlike gnomAD, the data was not reprocessed in a harmonized way; the variant VCFs were collected from the\ projects as-is. The goal is a single place to compare how common a variant is across\ different populations, ancestries, and cohorts, for projects that gnomAD is unlikely to\ reprocess soon. Three combined tracks aggregate the source data along different lines, and\ there is also one subtrack per project with the original VCF data and all the annotations\ that the project provides. The different projects use different pipelines and sequencing\ technologies. Click any of the projects above or below for a summary of their sample\ selection, sequencing assay and software pipeline. Many projects do not allow us to\ distribute the data, but we document how to request it and provide all converters, see Data Download below.\
\ \\ The browser has other tracks with variant frequencies. We have of course the data \ from gnomAD in separate tracks. Two projects that\ provide haplotype-phased genotypes can also be found in their own tracks:\ 1000 Genomes is a separate track, and the phased\ genotypes HGDP, SGDP, HGDP+1000 Genomes and Mexico Biobank are in the\ Phased Variants track. Their VCF versions below show\ only the allele frequency per variant, not the phased genotypes.\
\ \Please contact us (genome@soe.ucsc.edu) if you know of a project that we should add. So far,\ we have requested data from Regeneron's Million Exomes and the Mexico City studies (both requests rejected);\ Taiwan Biobank and the full UK Biobank WGS data requests are pending.
\ \\ Three combined tracks merge variants from the individual subtracks into single bigBed files\ with predicted protein consequences and cross-database filtering. All three use the same\ filter conventions (variant type, consequence, source database, allele frequency, allele\ count, and per-database AF/AC).\
\\ On the Disease and Population reference tracks, Affected AF and Background AF\ are pooled across contributing cohort arms (sum of allele counts divided by sum of allele\ numbers), not the maximum across arms, so the displayed frequency matches the carrier-count\ scale and a small cohort with a high local frequency does not dominate the value. See the\ "Pooled allele frequency" section on each combined track's description page for\ which cohorts contribute to the pool numerator and denominator.\
\ \\
All three combined tracks share the same Consequence filter (Missense, Synonymous, Stop\
Gained, Frameshift, Splice Donor, Splice Acceptor, Intron, 3' UTR, 5' UTR, Non-coding,\
Intergenic, Other). The filter uses OR logic across the comma-separated consequence terms\
on each variant: a variant tagged stop_gained,frameshift is selected by either\
the "Stop Gained" or the "Frameshift" filter. The "Other"\
bucket catches the less common\
Sequence Ontology consequence\
that don't fit the named buckets above. Examples\
include splice_region (variant near a splice site but outside the canonical\
donor/acceptor), start_lost / stop_lost (variant disrupts the\
start codon or replaces the stop codon with a coding amino acid),\
stop_retained (variant changes the stop codon but keeps it a stop),\
inframe_insertion / inframe_deletion (in-frame indel that adds or\
removes whole codons), and coding_sequence (CDS variant where the precise\
impact is undetermined). If you include "Other" in the filter selection, no\
records will be hidden by the consequence filter.\
| Combined tracks | ||||||
| Database | \Region | \N | \Data Type | \Cohort | \Sub-populations | \Downloadable from UCSC | \
|---|---|---|---|---|---|---|
| Disease cohorts | \Sequencing-based disease cohorts | \~130k | \WGS/WES/long-read | \Affected/case arms of SFARI SPARK WES/WGS, SCHEMA, GREGoR, GA4K | \Affected/case AF and AC; background AF for contrast | \No | \
| Population reference | \Sequencing-based, population + unaffected | \~1.5mil | \WGS/WES/long-read | \Population cohorts + unaffected/control arms | \Background AF and AC; per-cohort and ancestry breakdowns | \No | \
| Genotyping Array Databases Combined | \TPMI, MexBB, UKBB | \~530k | \Array / imputed | \14.7M variants | \— | \No | \
| Individual project datasets | ||||||
| Database | \Region | \N | \Data Type | \Cohort | \Sub-populations | \Downloadable from UCSC | \
| AllOfUs v7 | \USA | \245k | \WGS | \General population, diverse | \African, Indigenous American, East Asian, European, Oceanian, South Asian\ (local ancestry; see Notes below) | \No | \
| TOPMED Freeze 10 | \USA | \151k | \WGS | \Heart, lung, blood, sleep disorder cohorts | \— | \No | \
| SFARI SPARK WES | \USA | \140k | \WES | \Autism families (parents + affected children) | \— | \No | \
| SFARI SPARK WGS | \USA | \12.5k | \WGS | \Autism families (parents + affected children) | \— | \No | \
| NCBI ALFA R4 | \USA | \408k | \WGS/WES/array mix | \Aggregated dbGaP studies, mixed phenotypes | \— | \Yes | \
| FinnGen R12 | \Finland | \500k | \Imputed (8.5k WGS ref panel) | \National biobank, ~10% of population | \— | \No | \
| UK Biobank (Neale Lab v3) | \UK | \361k | \Imputed array (HRC+UK10K+1KGp3 ref panel) | \White British subset of UK Biobank, Neale Lab Round 2 GWAS | \— | \Yes | \
| SweGen | \Sweden | \1k | \WGS | \Cross-section of Swedish population | \— | \No | \
| GoNL | \Netherlands | \498 | \WGS (~13x) | \250 unrelated Dutch trios (parents only) | \— | \Yes | \
| SCHEMA | \Multi-national | \121k | \WES | \Schizophrenia: 24k cases, 97k controls (Singh 2022 primary); VCF aggregates up to ~73k/~182k | \— | \Yes | \
| Japan ToMMO 61k | \Japan | \61k | \WGS | \General population | \— | \Yes | \
| WBBC China | \China | \4.5k | \WGS | \Westlake BioBank for Chinese pilot (now part of China Precision BioBank), autosomes only | \North Han, Central Han, South Han, Lingnan Han (by recruitment region) | \Yes | \
| ChinaMAP phase 1 | \China | \10.5k | \WGS | \China Metabolic Analytics Project, ~40x depth, 27 provinces and 8 ethnic groups, autosomes only | \— | \No | \
| Taiwan TPMI | \Taiwan | \165k | \Axiom SNP array (TPM1) | \Taiwan Precision Medicine Initiative, Han Chinese | \— | \No | \
| Australia MGRB | \Australia | \4k | \WGS | \Healthy elderly (age ≥70) | \— | \No | \
| GenomeAsia Pilot | \Asia (219 groups) | \1.7k | \WGS | \Diverse populations across Asia | \Northeast Asian, Southeast Asian, South Asian, Oceanian, American, African,\ Western European Reference | \Yes | \
| ABraOM Brazil | \Brazil | \1.2k | \WGS | \Elderly admixed individuals (São Paulo) | \— | \Yes | \
| IndiGenomes | \India | \1k | \WGS | \Healthy individuals | \— | \Yes | \
| GenomeIndia 9.7k | \India | \9.8k | \WGS (≥23x) | \83 anthropologically defined endogamous populations across India | \— | \No | \
| KOVA Korea | \Korea | \5.3k | \1.9k WGS + 3.4k WES | \Normal tissue from cancer patients, healthy parents, volunteers | \— | \No | \
| NPM Singapore | \Singapore | \9.8k | \WGS | \Chinese, Indian, Malay ancestry | \— | \No | \
| Saudi Genome | \Saudi Arabia | \302 | \WGS (30x) | \Saudi population | \— | \Yes | \
| HRC | \Multi-national | \~30k | \Low-coverage WGS (7x) | \Imputation reference panel (excl. 1000 Genomes) | \— | \Yes | \
| MXB Mexico Biobank | \Mexico | \6k | \Genotyping array | \Diverse Mexican ancestries, 898 recruitment sites | \By state, by ancestry | \No | \
| SGDP | \Global | \279 | \WGS | \142 diverse populations worldwide | \By population | \Yes | \
| GREGoR R4 | \USA | \3.6k | \WGS | \Rare disease families (10.7k participants, 4.4k families) | \— | \Yes | \
| gnomAD HGDP+1kG | \Global | \4k | \WGS | \80 populations (HGDP + 1000 Genomes reprocessed) | \4k-cohort total AF only; per-population AF columns are full gnomAD v3.1.2\ release values (~76k genomes), see Notes below | \Yes | \
| GA4K | \USA | \552 | \PacBio HiFi long-read WGS | \Genomic Answers for Kids: pediatric rare-disease probands and families (Children's Mercy) | \— | \Yes | \
| CoLoRSdb v1.2.0 | \Multi-national | \1,027 | \PacBio HiFi long-read WGS | \Consortium of Long Read Sequencing: aggregated population-consented samples across multiple research cohorts | \— | \Yes | \
| SVatalog 101 | \Canada (SickKids) | \101 | \10X Genomics linked short-read WGS | \GWAS SVatalog cohort: 101 samples with matched long-read SVs (see chirmade101Sv) | \— | \Yes | \
| Indigenous Africans 180 | \Africa (Ethiopia, Tanzania, Cameroon, Botswana) | \180 | \WGS (>30x) | \12 indigenous populations across all four African language phyla (Khoesan, Niger-Congo, Nilo-Saharan, Afroasiatic) | \— | \No | \
Most tracks only show the variant and allele frequencies on mouseover or clicks.\ When zoomed in, tracks display alleles with base-specific coloring. Homozygote\ data are shown as one letter; heterozygotes are shown with both\ letters. All VCF files are normalized, with one allele per annotation (no multi-allele\ lines).\
\ \\ Each subtrack includes the upstream project's VCF largely as-released,\ sometimes converted from other file formats; per-subtrack pipelines (coordinate\ liftover, format conversion, header normalization) are documented on each\ subtrack's own description page and recorded in the\ build documentation.\ The conversion scripts \ live alongside the makedoc\ in the scripts directory.\
\\
The combined Disease cohorts and Population reference tracks are built by a separate\
pipeline: each per-subtrack VCF is normalized (bcftools norm), all sites are\
merged into a single callset, consequence annotations are recomputed against Ensembl with\
bcftools csq, and the merged callset is split by phenotype. Within each combined\
track, the Affected AF and Background AF columns are\
pooled across contributing cohort arms (sum of allele counts divided by sum of\
allele numbers, with the per-arm AN derived from each cohort's AC and AF), so the displayed\
frequency matches the carrier-count.\
The Genotyping Array Databases Combined track is built the same\
way from the array cohorts only.\
Many of these databases have restrictions on redistribution and download.\ The table above indicates if we are allowed to distribute it in VCF format.\ Click the database link in the table above and see the "Data Access"\ section of the respective track for a description of where to download the\ data. When the data is freely available from our website, the Data Access\ section will also indicate the VCF file location on our download server.\ Because it contains some licensed data, the combined track is not available for\ download, but can be recreated using the conversion scripts in our GitHub repository and the accompanying documentation file.
\ \This track is only possible thanks to the data from millions of volunteers around the world, who donated blood, signed consent forms and provided health information about themselves and sometimes their families. Click any of the tracks in the list above to see the specific credits for each project. Thanks to Alex Ioannidis, UCSC, for the inspiration for this track and to Andreas Lahner, MGZ, for feedback.
\ \\ All of Us Research Program Genomics Investigators.\ \ Genomic data in the All of Us Research Program.\ Nature. 2024 Mar;627(8003):340-346.\ PMID: 38374255; PMC: PMC10937371\
\ \\ Ameur A, Dahlberg J, Olason P, Vezzi F, Karlsson R, Martin M, Viklund J, Kahari AK, Lundin P, Che H\ et al.\ \ SweGen: a whole-genome data resource of genetic variability in a cross-section of the Swedish\ population.\ Eur J Hum Genet. 2017 Nov;25(11):1253-1260.\ PMID: 28832569; PMC: PMC5765326\
\ \\ Bhattacharyya C, Subramanian K, Uppili B, Biswas NK, Ramdas S, Tallapaka KB, Arvind P, Rupanagudi\ KV, Maitra A, Nagabandi T et al.\ \ Mapping genetic diversity with the GenomeIndia project.\ Nat Genet. 2025 Apr;57(4):767-773.\ PMID: 40200122\
\ \\ Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, Motyer A, Vukcevic D, Delaneau O,\ O'Connell J et al.\ \ The UK Biobank resource with deep phenotyping and genomic data.\ Nature. 2018 Oct;562(7726):203-209.\ PMID: 30305743; PMC: PMC6786975\
\ \\ Cao Y, Li L, Xu M, Feng Z, Sun X, Lu J, Xu Y, Du P, Wang T, Hu R et al.\ \ The ChinaMAP analytics of deep whole genome sequences in 10,588 individuals.\ Cell Res. 2020 Sep;30(9):717-731.\ PMID: 32355288; PMC: PMC7609296\
\ \\ Chirmade S, Wang Z, Mastromatteo S, Sanders E, Thiruvahindrapuram B, Nalpathamkalam T, Pellecchia G,\ Lin F, Keenan K, Patel RV et al.\ \ GWAS SVatalog: a visualization tool to aid fine-mapping of GWAS loci with structural variations.\ Heredity (Edinb). 2025 Sep;135(3):199-210.\ PMID: 41203876; PMC: PMC13031531\
\ \\ Cohen ASA, Farrow EG, Abdelmoity AT, Alaimo JT, Amudhavalli SM, Anderson JT, Bansal L, Bartik L,\ Baybayan P, Belden B et al.\ \ Genomic answers for children: Dynamic analyses of >1000 pediatric rare disease genomes.\ Genet Med. 2022 Jun;24(6):1336-1348.\ PMID: 35305867\
\ \\ Cong PK, Bai WY, Li JC, Yang MY, Khederzadeh S, Gai SR, Li N, Liu YH, Yu SH, Zhao WW et al.\ \ Genomic analyses of 10,376 individuals in the Westlake BioBank for Chinese (WBBC) pilot project.\ Nat Commun. 2022 May 26;13(1):2939.\ PMID: 35618720; PMC: PMC9135724\
\ \\ Fan S, Spence JP, Feng Y, Hansen MEB, Terhorst J, Beltrame MH, Ranciaro A, Hirbo J, Beggs W, Thomas\ N et al.\ \ Whole-genome sequencing reveals a complex African population demographic history and signatures of\ local adaptation.\ Cell. 2023 Mar 2;186(5):923-939.e14.\ PMID: 36868214; PMC: PMC10568978\
\ \\ Feliciano P, Daniels AM, Snyder LG, Beaumont A, Camba A, Esler A, Gulsrud AG, Mason A, Nicholson A,\ Paolicelli AM et al; The SPARK Consortium.\ \ SPARK: A US Cohort of 50,000 Families to Accelerate Autism Research.\ Neuron. 2018 Feb 7;97(3):488-493.\ PMID: 29420931; PMC: PMC7444276\
\ \\ Genome of the Netherlands Consortium.\ \ Whole-genome sequence variation, population structure and demographic history of the Dutch\ population.\ Nat Genet. 2014 Aug;46(8):818-25.\ PMID: 24974849\
\ \\ GenomeAsia100K Consortium.\ \ The GenomeAsia 100K Project enables genetic discoveries across Asia.\ Nature. 2019 Dec;576(7785):106-111.\ PMID: 31802016; PMC: PMC7054211\
\ \\ Jain A, Bhoyar RC, Pandhare K, Mishra A, Sharma D, Imran M, Senthivel V, Divakar MK, Rophina M,\ Jolly B et al.\ \ IndiGenomes: a comprehensive resource of genetic variants from over 1000 Indian genomes.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D1225-D1232.\ PMID: 33095885; PMC: PMC7778947\
\ \\ Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alfoldi J, Wang Q, Collins RL, Laricchia KM,\ Ganna A, Birnbaum DP et al.\ \ The mutational constraint spectrum quantified from variation in 141,456 humans.\ Nature. 2020 May;581(7809):434-443.\ PMID: 32461654; PMC: PMC7334197\
\ \\ Koenig Z, Yohannes MT, Nkambule LL, Zhao X, Goodrich JK, Kim HA, Wilson MW, Tiao G, Hao SP, Sahakian\ N et al.\ \ A harmonized public resource of deeply sequenced diverse human genomes.\ Genome Res. 2024 Jun 25;34(5):796-809.\ PMID: 38749656; PMC: PMC11216312\
\ \\ Kurki MI, Karjalainen J, Palta P, Sipila TP, Kristiansson K, Donner KM, Reeve MP, Laivuori H,\ Aavikko M, Kaunisto MA et al.\ \ FinnGen provides genetic insights from a well-phenotyped isolated population.\ Nature. 2023 Jan;613(7944):508-518.\ PMID: 36653562; PMC: PMC9849126\
\ \\ Lacaze P, Pinese M, Kaplan W, Stone A, Brion MJ, Woods RL, McNamara M, McNeil JJ, Dinger ME,\ Thomas DM.\ \ The Medical Genome Reference Bank: a whole-genome data resource of 4000 healthy elderly individuals.\ Rationale and cohort design.\ Eur J Hum Genet. 2019 Feb;27(2):308-316.\ PMID: 30353151; PMC: PMC6336775\
\ \\ Lee S, Seo J, Park J, Nam JY, Choi A, Ignatius JS, Bjornson RD, Chae JH, Jang IJ, Lee S\ et al.\ \ Korean Variant Archive (KOVA): a reference database of genetic variations in the Korean\ population.\ Sci Rep. 2017 Jun 27;7(1):4287.\ PMID: 28655895; PMC: PMC5487339\
\ \\ Mallick S, Li H, Lipson M, Mathieson I, Gymrek M, Racimo F, Zhao M, Chennagiri N, Nordenfelt S,\ Tandon A et al.\ \ The Simons Genome Diversity Project: 300 genomes from 142 diverse populations.\ Nature. 2016 Oct 13;538(7624):201-206.\ PMID: 27654912; PMC: PMC5161557\
\ \\ Malomane DK, Williams MP, Huber CD, Mangul S, Abedalthagafi M, Chiang CWK.\ \ Patterns of population structure and genetic variation within the Saudi Arabian population.\ bioRxiv. 2025 Jan 13;.\ PMID: 39868174; PMC: PMC11761371\
\ \\ McCarthy S, Das S, Kretzschmar W, Delaneau O, Wood AR, Teumer A, Kang HM, Fuchsberger C, Danecek P,\ Sharp K et al.\ \ A reference panel of 64,976 haplotypes for genotype imputation.\ Nat Genet. 2016 Oct;48(10):1279-83.\ PMID: 27548312; PMC: PMC5388176\
\ \\ Naslavsky MS, Scliar MO, Yamamoto GL, Wang JYT, Zverinova S, Karp T, Nunes K, Ceroni JRM,\ de Carvalho DL, da Silva Simões CE et al.\ \ Whole-genome sequencing of 1,171 elderly admixed individuals from São Paulo, Brazil.\ Nat Commun. 2022 Mar 4;13(1):1004.\ PMID: 35246524; PMC: PMC8897431\
\ \\ Singh T, Poterba T, Curtis D, Akil H, Al Eissa M, Barchas JD, Bass N, Bigdeli TB, Breen G,\ Bromet EJ et al.\ \ Rare coding variants in ten genes confer substantial risk for schizophrenia.\ Nature. 2022 Apr;604(7906):509-516.\ PMID: 35396579; PMC: PMC9805802\
\ \\ Sohail M, Palma-Martínez MJ, Chong AY, Quinto-Cortés CD, Barberena-Jonas C,\ Medina-Muñoz SG, Ragsdale A, Delgado-Sánchez G, Cruz-Hervert LP, Ferreyra-Reyes L\ et al.\ \ Mexican Biobank advances population and medical genomics of diverse ancestries.\ Nature. 2023 Oct;622(7984):775-783.\ PMID: 37821706; PMC: PMC10600006\
\ \\ Tadaka S, Kawashima J, Hishinuma E, Saito S, Okamura Y, Otsuki A, Kojima K, Komaki S, Aoki Y,\ Kanno T et al.\ \ jMorp: Japanese Multi-Omics Reference Panel update report 2023.\ Nucleic Acids Res. 2024 Jan 5;52(D1):D622-D632.\ PMID: 37930845; PMC: PMC10767895\
\ \\ Taliun D, Harris DN, Kessler MD, Carlson J, Szpiech ZA, Torres R, Taliun SAG, Corvelo A, Gogarten SM,\ Kang HM et al.\ \ Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program.\ Nature. 2021 Feb;590(7845):290-299.\ PMID: 33568819; PMC: PMC7875770\
\ \\ Wong E, Bertin N, Hebrard M, Tirado-Magallanes R, Bellis C, Lim WK, Chua CY, Tong PML, Chua R, Mak K\ et al.\ \ The Singapore National Precision Medicine Strategy.\ Nat Genet. 2023 Feb;55(2):178-186.\ PMID: 36658435\
\ \\ Wu D, Dou J, Chai X, Bellis C, Wilm A, Shih CC, Soon WWJ, Bertin N, Lin CB, Khor CC et al.\ \ Large-scale whole-genome sequencing of three diverse Asian populations in Singapore.\ Cell. 2019 Oct 17;179(3):736-749.e15.\ PMID: 31626772\
\ \\ Yang HC, Kwok PY, Li LH, Liu YM, Jong YJ, Lee KY, Wang DW, Tsai MF, Yang JH, Chen CH et al.\ \ The Taiwan Precision Medicine Initiative provides a cohort for large-scale studies.\ Nature. 2025 Dec;648(8092):117-127.\ PMID: 41092961; PMC: PMC12675286\
\ varRep 1 group varRep\ longLabel SNV Frequencies from various cohorts or national projects\ pennantIcon New red ../goldenPath/newsarch.html#070126 "Released Jul. 1, 2026"\ shortLabel SNV Frequencies\ superTrack on\ track varFreqs\ type bed 12\ visibility hide\ gnomADPextSpleen Spleen bigWig 0 1 gnomAD pext Spleen 0 100 119 136 85 187 195 170 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Spleen.bw\ color 119,136,85\ longLabel gnomAD pext Spleen\ parent gnomadPext off\ shortLabel Spleen\ track gnomADPextSpleen\ visibility hide\ intronEst Spliced ESTs psl est Human ESTs That Have Been Spliced 0 100 0 0 0 127 127 127 1 0 0\ This track shows alignments between human expressed sequence tags\ (ESTs) in \ GenBank and the genome that show signs of splicing when\ aligned against the genome. ESTs are single-read sequences, typically about\ 500 bases in length, that usually represent fragments of transcribed genes.\
\ \\ To be considered spliced, an EST must show\ evidence of at least one canonical intron (i.e., the genomic\ sequence between EST alignment blocks must be at least 32 bases in\ length and have GT/AG ends). By requiring splicing, the level\ of contamination in the EST databases is drastically reduced\ at the expense of eliminating many genuine 3' ESTs.\ For a display of all ESTs (including unspliced), see the\ human EST track.\
\ \\ This track follows the display conventions for\ \ PSL alignment tracks. In dense display mode, darker shading\ indicates a larger number of aligned ESTs.\
\ \\ The strand information (+/-) indicates the\ direction of the match between the EST and the matching\ genomic sequence. It bears no relationship to the direction\ of transcription of the RNA with which it might be associated.\
\ \\ The description page for this track has a filter that can be used to change\ the display mode, alter the color, and include/exclude a subset of items\ within the track. This may be helpful when many items are shown in the track\ display, especially when only some are relevant to the current task.\
\ \\ To use the filter:\
\ This track may also be configured to display base labeling, a feature that\ allows the user to display all bases in the aligning sequence or only those\ that differ from the genomic sequence. For more information about this option,\ go to the\ \ Base Coloring for Alignment Tracks page.\ Several types of alignment gap may also be colored;\ for more information, go to the\ \ Alignment Insertion/Deletion Display Options page.\
\ \\ To make an EST, RNA is isolated from cells and reverse\ transcribed into cDNA. Typically, the cDNA is cloned\ into a plasmid vector and a read is taken from the 5'\ and/or 3' primer. For most — but not all — ESTs, the\ reverse transcription is primed by an oligo-dT, which\ hybridizes with the poly-A tail of mature mRNA. The\ reverse transcriptase may or may not make it to the 5'\ end of the mRNA, which may or may not be degraded.\
\ \\ In general, the 3' ESTs mark the end of transcription\ reasonably well, but the 5' ESTs may end at any point\ within the transcript. Some of the newer cap-selected\ libraries cover transcription start reasonably well. Before the\ cap-selection techniques\ emerged, some projects used random rather than poly-A\ priming in an attempt to retrieve sequence distant from the\ 3' end. These projects were successful at this, but as\ a side effect also deposited sequences from unprocessed\ mRNA and perhaps even genomic sequences into the EST databases.\ Even outside of the random-primed projects, there is a\ degree of non-mRNA contamination. Because of this, a\ single unspliced EST should be viewed with considerable\ skepticism.\
\ \\ To generate this track, human ESTs from GenBank were aligned\ against the genome using blat. Note that the maximum intron length\ allowed by blat is 750,000 bases, which may eliminate some ESTs with very\ long introns that might otherwise align. When a single\ EST aligned in multiple places, the alignment having the\ highest base identity was identified. Only alignments having\ a base identity level within 0.5% of the best and at least 96% base identity\ with the genomic sequence are displayed in this track.\
\ \\ This track was produced at UCSC from EST sequence data\ submitted to the international public sequence databases by\ scientists worldwide.\
\ \\ Benson DA, Cavanaugh M, Clark K, Karsch-Mizrachi I, Lipman DJ, Ostell J, Sayers EW.\ \ GenBank.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D36-42.\ PMID: 23193287; PMC: PMC3531190\
\ \\ Benson DA, Karsch-Mizrachi I, Lipman DJ, Ostell J, Wheeler DL.\ GenBank: update.\ Nucleic Acids Res. 2004 Jan 1;32(Database issue):D23-6.\ PMID: 14681350; PMC: PMC308779\
\ \\ Kent WJ.\ BLAT - the BLAST-like alignment tool.\ Genome Res. 2002 Apr;12(4):656-64.\ PMID: 11932250; PMC: PMC187518\
\ rna 1 baseColorUseSequence genbank\ group rna\ indelDoubleInsert on\ indelQueryInsert on\ intronGap 30\ longLabel Human ESTs That Have Been Spliced\ maxItems 300\ shortLabel Spliced ESTs\ showDiffBasesAllScales .\ spectrum on\ track intronEst\ type psl est\ visibility hide\ spliceVarDb SpliceVarDB bigLolly SpliceVarDB: Experimentally validated splicing variants 2 100 0 0 0 127 127 127 0 0 0 https://compbio.ccia.org.au/splicevardb/\ The "Splicing Impact" container track contains tracks showing the predicted or validated effect of variants\ close to splice sites.\
\ \AbSplice is a method that predicts aberrant splicing across human tissues, as described in Wagner,\ Çelik et al., 2023. This track displays precomputed AbSplice scores for all possible\ single-nucleotide variants genome-wide. The scores represent the probability that a given variant\ causes aberrant splicing in a given tissue.\ AbSplice scores\ can be computed from VCF files and are based on quantitative tissue-specific splice site annotations\ (SpliceMaps).\ While SpliceMaps can be generated for any tissue of interest from a cohort of RNA-seq samples, this\ track includes 49 tissues available from the\ Genotype-Tissue\ Expression (GTEx) dataset.\
\ \SpliceAI is an open-source deep\ learning splicing prediction algorithm that can predict splicing alterations caused by DNA variations.\ To score variants, the spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added to the sequence and the new sequence is\ scored. Variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a\ lookup tool\ is provided by the Broad institute. \
\ \\ This SpliceAI "Wildtype" container track shows the scores for the genome sequence itself,\ without variants, from predicted splice donor (5' intron boundaries) and splice acceptor\ (3' intron boundaries) sites. Predictions are strand-specific, with separate subtracks for the\ plus and minus strands. These tracks are useful in combination with the variants track for\ evaluating new transcript models. They can be used to assess potential exon boundaries or\ possible splice acceptor sites.
\ \ Why are some variants not scored by SpliceAI?\\ SpliceAI only annotates variants within genes defined by the gene\ annotation file. Additionally, SpliceAI does not annotate variants if they are close to chromosome\ ends (5kb on either side), deletions of length greater than twice the input parameter -D, or\ inconsistent with the reference fasta file.\
\ \ What are the differences between masked and unmasked tracks?\\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \SpliceVarDB is an online database consolidating over 50,000 variants assayed\ for their effects on splicing in over 8,000 human genes. The authors evaluated\ over 500 published data sources and established a spliceogenicity scale to\ standardize, harmonize, and consolidate variant validation data generated by a\ range of experimental protocols. Genes and variant locations were obtained using\ GENCODE v44. Splice regions were calculated as specific distances from the closest\ canonical exon, including 5' and 3' untranslated regions (UTRs). The\ database is available at\ splicevardb.org.
\ \The AbSplice score is a probability estimate of how likely aberrant splicing of some sort takes\ place in a given tissue. The authors suggest three cutoffs which are represented by color in the track.\
\ \\ Mouseover on items shows the gene name, maximum score, and tissues that had this score. Clicking on\ any item brings up a table with scores for all 49 GTEX tissues.\
\ \\ Variants are colored according to Walker et al. 2023 splicing impact:\
\\ The scores range from 0 to 1 and can be interpreted as the\ probability of the variant being splice-altering. In the paper, a detailed characterization is\ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\ These tracks are in bigWig format. The signal height represents the SpliceAI probability score.\ This track may be configured in a variety of ways to highlight different aspects of the displayed\ information. Click the "Graph configuration help" link for an explanation of configuration\ options.
\ \According to the strength of their supporting\ evidence, variants were classified as "splice-altering" (~25%), "not\ splice-altering" (~25%), and "low-frequency splice-altering" (~50%), which\ correspond to weak or indeterminate evidence of spliceogenicity. 55% of the\ splice-altering variants in SpliceVarDB are outside the canonical splice sites\ (5.6% are deep intronic). The data is shown as lollipop plots that can be clicked, \ the details page then shows a link to SpliceVarDB with full details.\
\ \The classification thresholds primarily follow those established by the original study.\ However, most studies only defined criteria for splice-altering variants and did not define\ criteria for variants that resulted in normal splicing. The authors implemented stringent\ thresholds to define the normal category and ensure a high-quality set of control variants.\ Variants that did not meet these criteria were classified as low-frequency splice-altering\ variants with a wide range of sub-optimal scores. Variants that fell between the normal and\ splice-altering classifications were placed into a low-frequency splice-altering category.\ In situations where a variant was validated multiple times, if at least one validation\ returned splice-altering and another returned normal, the "conflicting" category\ was applied.\
\ \\ The lollipop plots are color-coded based on the score value, which corresponds\ to the following classifications:\
Data was converted from the files (AbSplice_DNA_ hg38 _snvs_high_scores.zip) provided by the authors\ at zenodo.org. Files in the\ score_cutoff=0.01 directory were concatenated. To convert the data to bigBed format, scores and\ their tissues were selected from the AbSplice_DNA fields and maximum scores, and then calculated\ using a custom Python script, which can be found in the\ \ makeDoc from our GitHub repository.
\ \\
The data were downloaded from Illumina.\
The spliceAI scores are represented in the VCF INFO field as\
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain\
\ Since most of the values are 0 or almost 0, we selected only those variants\ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \Data was provided by the Michael Hiller lab. SpliceAI was run on the entire genome reference\ chromosomes. Since the algorithm does not know where transcripts start or end, the scores\ can differ from those on other websites, especially for splice sites before the last exon or\ around the first exon.
\ \ \The data was converted by Patricia Sullivan from SpliceVarDB to\ bigLolly format, and the UCSC\ Browser staff downloaded it for display.\
\ \Precomputed AbSplice-DNA scores in all 49 GTEx tissues are available at\ \ Zenodo.
\ \ License\\ The SpliceAI data is not available for download from the Genome Browser.\ The raw data can be found directly on\ Illumina.\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are\ made available by Illumina only for academic or not-for-profit research only.\ By accessing the SpliceAI data, you acknowledge and agree that you may only\ use this data for your own personal academic or not-for-profit research only,\ and not for any other purposes. You may not use this data for any for-profit,\ clinical, or other commercial purpose without obtaining a commercial license\ from Illumina, Inc.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. For automated analysis, the data may\ be queried from our REST API.
\ \\
For automated download and analysis, the genome annotation is stored in a bigBed or a bigWig file\
that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools, e.g.\
\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg19/splicevardb/SVADB.bb\
\ -chrom=chr21 -start=0 -end=100000000 stdout\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500\
\ http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/spliceAi/wildtype/spliceAiAcceptorMinus.bw\
\ stdout\
\
\
These tools can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.
Thanks to Illumina for making SpliceAI available, both the model and the precomputed data files.
\ \Thanks to Francois Lecoquierre from the University of Oxford, Jean-Madeleine de Sainte Agathe\ from Institut Pasteur Paris, and Michael Hiller from the Senckenberg Museum Frankfurt for\ suggesting and then creating the SpliceAI Wildtype annotations.
\ \Thanks to Nils Wagner for helpful comments and suggestions for the AbSplice track.
\ \Thanks to the SpliceVarDB team for converting the data into our data formats.
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Sullivan PJ, Quinn JMW, Wu W, Pinese M, Cowley MJ.\ \ SpliceVarDB: A comprehensive database of experimentally validated human splicing variants.\ Am J Hum Genet. 2024 Oct 3;111(10):2164-2175.\ PMID: 39226898; PMC: PMC11480807\
\ \\ Wagner N, Çelik MH, Hölzlwimmer FR, Mertes C, Prokisch H, Yépez VA, Gagneur J.\ \ Aberrant splicing prediction across human tissues.\ Nat Genet. 2023 May;55(5):861-870.\ PMID: 37142848\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ \ phenDis 1 bigDataUrl /gbdb/hg38/splicevardb/SVDB.bb\ dataVersion Nov 2024\ group phenDis\ html spliceImpactSuper\ itemRgb on\ lollyMaxSize 5\ lollyNoStems on\ lollySizeField lollySize\ longLabel SpliceVarDB: Experimentally validated splicing variants\ parent spliceImpactSuper on\ shortLabel SpliceVarDB\ skipFields lollySize\ track spliceVarDb\ type bigLolly\ url https://compbio.ccia.org.au/splicevardb/\ urlLabel Go to SpliceVarDB\ viewLimits 0:3\ visibility full\ yAxisLabel.0 0 on 140,140,140 Conflicting\ yAxisLabel.1 1 on 140,140,140 Normal\ yAxisLabel.2 2 on 140,140,140 Low\ yAxisLabel.3 3 on 140,140,140 Splice\ yAxisNumLabels off\ spliceImpactSuper Splicing Impact Splicing Impact Prediction Scores and Databases 0 100 0 0 0 127 127 127 0 0 0\ The "Splicing Impact" container track contains tracks showing the predicted or validated effect of variants\ close to splice sites.\
\ \AbSplice is a method that predicts aberrant splicing across human tissues, as described in Wagner,\ Çelik et al., 2023. This track displays precomputed AbSplice scores for all possible\ single-nucleotide variants genome-wide. The scores represent the probability that a given variant\ causes aberrant splicing in a given tissue.\ AbSplice scores\ can be computed from VCF files and are based on quantitative tissue-specific splice site annotations\ (SpliceMaps).\ While SpliceMaps can be generated for any tissue of interest from a cohort of RNA-seq samples, this\ track includes 49 tissues available from the\ Genotype-Tissue\ Expression (GTEx) dataset.\
\ \SpliceAI is an open-source deep\ learning splicing prediction algorithm that can predict splicing alterations caused by DNA variations.\ To score variants, the spliceAI algorithm is run on the genome sequence itself and scores each\ nucleotide for the probability that it is a donor or acceptor site, on both the\ forward and the reverse strand. Then variants are added to the sequence and the new sequence is\ scored. Variants may activate nearby cryptic splice sites, leading to abnormal transcript isoforms.\ SpliceAI was developed at Illumina; a\ lookup tool\ is provided by the Broad institute. \
\ \\ This SpliceAI "Wildtype" container track shows the scores for the genome sequence itself,\ without variants, from predicted splice donor (5' intron boundaries) and splice acceptor\ (3' intron boundaries) sites. Predictions are strand-specific, with separate subtracks for the\ plus and minus strands. These tracks are useful in combination with the variants track for\ evaluating new transcript models. They can be used to assess potential exon boundaries or\ possible splice acceptor sites.
\ \ Why are some variants not scored by SpliceAI?\\ SpliceAI only annotates variants within genes defined by the gene\ annotation file. Additionally, SpliceAI does not annotate variants if they are close to chromosome\ ends (5kb on either side), deletions of length greater than twice the input parameter -D, or\ inconsistent with the reference fasta file.\
\ \ What are the differences between masked and unmasked tracks?\\ The unmasked tracks include splicing changes corresponding to strengthening annotated splice sites\ and weakening unannotated splice sites, which are typically much less pathogenic than weakening\ annotated splice sites and strengthening unannotated splice sites. The delta scores of such splicing\ changes are set to 0 in the masked files. We recommend using the unmasked tracks for alternative\ splicing analysis and masked tracks for variant interpretation.\
\ \SpliceVarDB is an online database consolidating over 50,000 variants assayed\ for their effects on splicing in over 8,000 human genes. The authors evaluated\ over 500 published data sources and established a spliceogenicity scale to\ standardize, harmonize, and consolidate variant validation data generated by a\ range of experimental protocols. Genes and variant locations were obtained using\ GENCODE v44. Splice regions were calculated as specific distances from the closest\ canonical exon, including 5' and 3' untranslated regions (UTRs). The\ database is available at\ splicevardb.org.
\ \The AbSplice score is a probability estimate of how likely aberrant splicing of some sort takes\ place in a given tissue. The authors suggest three cutoffs which are represented by color in the track.\
\ \\ Mouseover on items shows the gene name, maximum score, and tissues that had this score. Clicking on\ any item brings up a table with scores for all 49 GTEX tissues.\
\ \\ Variants are colored according to Walker et al. 2023 splicing impact:\
\\ The scores range from 0 to 1 and can be interpreted as the\ probability of the variant being splice-altering. In the paper, a detailed characterization is\ provided for 0.2 (high recall), 0.5 (recommended), and 0.8 (high precision) cutoffs.
\ \\ These tracks are in bigWig format. The signal height represents the SpliceAI probability score.\ This track may be configured in a variety of ways to highlight different aspects of the displayed\ information. Click the "Graph configuration help" link for an explanation of configuration\ options.
\ \According to the strength of their supporting\ evidence, variants were classified as "splice-altering" (~25%), "not\ splice-altering" (~25%), and "low-frequency splice-altering" (~50%), which\ correspond to weak or indeterminate evidence of spliceogenicity. 55% of the\ splice-altering variants in SpliceVarDB are outside the canonical splice sites\ (5.6% are deep intronic). The data is shown as lollipop plots that can be clicked, \ the details page then shows a link to SpliceVarDB with full details.\
\ \The classification thresholds primarily follow those established by the original study.\ However, most studies only defined criteria for splice-altering variants and did not define\ criteria for variants that resulted in normal splicing. The authors implemented stringent\ thresholds to define the normal category and ensure a high-quality set of control variants.\ Variants that did not meet these criteria were classified as low-frequency splice-altering\ variants with a wide range of sub-optimal scores. Variants that fell between the normal and\ splice-altering classifications were placed into a low-frequency splice-altering category.\ In situations where a variant was validated multiple times, if at least one validation\ returned splice-altering and another returned normal, the "conflicting" category\ was applied.\
\ \\ The lollipop plots are color-coded based on the score value, which corresponds\ to the following classifications:\
Data was converted from the files (AbSplice_DNA_ hg38 _snvs_high_scores.zip) provided by the authors\ at zenodo.org. Files in the\ score_cutoff=0.01 directory were concatenated. To convert the data to bigBed format, scores and\ their tissues were selected from the AbSplice_DNA fields and maximum scores, and then calculated\ using a custom Python script, which can be found in the\ \ makeDoc from our GitHub repository.
\ \\
The data were downloaded from Illumina.\
The spliceAI scores are represented in the VCF INFO field as\
SpliceAI=G|OR4F5|0.01|0.00|0.00|0.00|-32|49|-40|-31
\
Here, the pipe-separated fields contain\
\ Since most of the values are 0 or almost 0, we selected only those variants\ with a score equal to or greater than 0.02.\
\\ The complete processing of this track can be found in the \ makedoc.\
\ \Data was provided by the Michael Hiller lab. SpliceAI was run on the entire genome reference\ chromosomes. Since the algorithm does not know where transcripts start or end, the scores\ can differ from those on other websites, especially for splice sites before the last exon or\ around the first exon.
\ \ \The data was converted by Patricia Sullivan from SpliceVarDB to\ bigLolly format, and the UCSC\ Browser staff downloaded it for display.\
\ \Precomputed AbSplice-DNA scores in all 49 GTEx tissues are available at\ \ Zenodo.
\ \ License\\ The SpliceAI data is not available for download from the Genome Browser.\ The raw data can be found directly on\ Illumina.\ FOR ACADEMIC AND NOT-FOR-PROFIT RESEARCH USE ONLY. The SpliceAI scores are\ made available by Illumina only for academic or not-for-profit research only.\ By accessing the SpliceAI data, you acknowledge and agree that you may only\ use this data for your own personal academic or not-for-profit research only,\ and not for any other purposes. You may not use this data for any for-profit,\ clinical, or other commercial purpose without obtaining a commercial license\ from Illumina, Inc.\
\ \\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. For automated analysis, the data may\ be queried from our REST API.
\ \\
For automated download and analysis, the genome annotation is stored in a bigBed or a bigWig file\
that can be downloaded from\
our download server.\
Individual regions or the whole genome annotation can be obtained using our tools, e.g.\
\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg19/splicevardb/SVADB.bb\
\ -chrom=chr21 -start=0 -end=100000000 stdout\
\
bigWigToBedGraph -chrom=chr1 -start=100000 -end=100500\
\ http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/spliceAi/wildtype/spliceAiAcceptorMinus.bw\
\ stdout\
\
\
These tools can be compiled from the source code or downloaded as a precompiled\
binary for your system. Instructions for downloading source code and binaries can be found\
here.
Thanks to Illumina for making SpliceAI available, both the model and the precomputed data files.
\ \Thanks to Francois Lecoquierre from the University of Oxford, Jean-Madeleine de Sainte Agathe\ from Institut Pasteur Paris, and Michael Hiller from the Senckenberg Museum Frankfurt for\ suggesting and then creating the SpliceAI Wildtype annotations.
\ \Thanks to Nils Wagner for helpful comments and suggestions for the AbSplice track.
\ \Thanks to the SpliceVarDB team for converting the data into our data formats.
\ \\ Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, Darbandi SF, Knowles D, Li YI, Kosmicki JA,\ Arbelaez J, Cui W, Schwartz GB et al.\ \ Predicting Splicing from Primary Sequence with Deep Learning.\ Cell. 2019 Jan 24;176(3):535-548.e24.\ PMID: 30661751\
\ \\ Sullivan PJ, Quinn JMW, Wu W, Pinese M, Cowley MJ.\ \ SpliceVarDB: A comprehensive database of experimentally validated human splicing variants.\ Am J Hum Genet. 2024 Oct 3;111(10):2164-2175.\ PMID: 39226898; PMC: PMC11480807\
\ \\ Wagner N, Çelik MH, Hölzlwimmer FR, Mertes C, Prokisch H, Yépez VA, Gagneur J.\ \ Aberrant splicing prediction across human tissues.\ Nat Genet. 2023 May;55(5):861-870.\ PMID: 37142848\
\ \\ Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, Canson DM, Bis-Brewer D, Cass A,\ Tchourbanov A et al.\ \ Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on\ splicing: Recommendations from the ClinGen SVI Splicing Subgroup.\ Am J Hum Genet. 2023 Jul 6;110(7):1046-1067.\ PMID: 37352859; PMC: PMC10357475\
\ \ phenDis 0 cartVersion 8\ group phenDis\ longLabel Splicing Impact Prediction Scores and Databases\ shortLabel Splicing Impact\ superTrack on hide\ track spliceImpactSuper\ gnomADPextStomach Stomach bigWig 0 1 gnomAD pext Stomach 0 100 255 221 153 255 238 204 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Stomach.bw\ color 255,221,153\ longLabel gnomAD pext Stomach\ parent gnomadPext off\ shortLabel Stomach\ track gnomADPextStomach\ visibility hide\ strchive STRchive bigBed 9 + STRchive Disease-Associated Short Tandem Repeat Loci 3 100 0 0 0 127 127 127 0 0 0 https://strchive.org/loci/$$\ The STRchive track displays 75 disease-associated short tandem repeat (STR) loci\ curated by the STRchive project.\ STRchive is a dynamic, community-driven resource that compiles population-level and\ locus-specific data for tandem repeat loci implicated in human genetic diseases.
\ \\ Tandem repeat expansion disorders are caused by the expansion of short repetitive DNA\ sequences beyond a pathogenic threshold. These expansions can cause a wide range of\ neurological, neuromuscular, and developmental disorders, including Huntington disease,\ fragile X syndrome, Friedreich ataxia, and many forms of spinocerebellar ataxia.
\ \\ This track shows the genomic positions of disease-associated STR loci from the STRchive\ catalog, along with the reference and pathogenic repeat motifs, minimum pathogenic repeat\ count thresholds, mode of inheritance, and associated diseases. The data are based on\ the GRCh38/hg38 reference assembly.
\ \\ Items are colored by mode of inheritance:
\\ Each item is labeled by its STRchive locus ID, which combines the disease abbreviation\ and gene symbol (e.g., "HD_HTT" for Huntington disease at the HTT\ gene). Hovering over an item shows the repeat motif, gene, pathogenic threshold,\ and inheritance mode. Clicking an item links to the corresponding\ STRchive locus page with detailed\ clinical and population-level information.
\ \\
The STRchive disease locus catalog was downloaded from the\
STRchive GitHub\
repository (file STRchive-disease-loci.hg38.general.bed). The catalog is\
manually curated by the STRchive team from published literature and contains loci where\
tandem repeat expansions have been reported to cause or be associated with human disease.
\ For each locus, the catalog provides:
\\ The BED file was converted to bigBed format for display in the Genome Browser. Coordinates\ were used as provided (0-based half-open BED format).
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated\ analysis, the data may be queried from our\ REST API. The underlying bigBed\ file can be downloaded from our\ download\ server.
\ \\ The complete STRchive dataset, including additional annotations not shown in this track,\ is available from strchive.org and\ the STRchive GitHub\ repository. The data are released under a\ CC BY 4.0\ license.
\ \\ Thanks to Harriet Dashnow (University of Colorado), Laurel Hiatt (University of Utah),\ Ben Weisburd (Broad Institute), and the STRchive team for creating and maintaining this\ resource.
\ \\ Hiatt L, Weisburd B, Dolzhenko E, Rubinetti V, Avvaru AK,\ VanNoy GE, Kurtas NE, Rehm HL, Quinlan AR, Dashnow H.\ \ STRchive: a dynamic resource detailing population-level and\ locus-specific insights at tandem repeat disease loci.\ Genome Med. 2025 Mar 26;17(1):29.\ PMID: 40140942; PMC: PMC11938676\
\ varRep 1 bigDataUrl /gbdb/hg38/strVar/strchive.bb\ dataVersion /gbdb/hg38/strVar/strchive.version.txt\ itemRgb on\ longLabel STRchive Disease-Associated Short Tandem Repeat Loci\ mouseOver Gene: $geneThis track shows locations of Sequence Tagged Site (STS) markers\ along the draft assembly. These markers have been mapped using either\ genetic mapping (Genethon, Marshfield, and deCODE maps), radiation\ hybridization mapping (Stanford, Whitehead RH, and GeneMap99 maps) or\ YAC mapping (the Whitehead YAC map) techniques. Since August 2001,\ this track no longer displays fluorescent in situ hybridization (FISH)\ clones, which are now displayed in a separate track.
\ \Genetic map markers are shown in blue; radiation hybrid map markers\ are shown in black. When a marker maps to multiple positions in the\ genome, it is shown in a lighter color.
\ \Positions of STS markers are determined using both full sequences\ and primer information. Full sequences are aligned using blat,\ while isPCR (Jim Kent) and ePCR are used to find\ locations using primer information. Both sets of placements are\ combined to give final positions. In nearly all cases, full sequence\ and primer-based locations are in agreement, but in cases of\ disagreement, full sequence positions are used. Sequence and primer\ information for the markers were obtained from the primary sites for\ each of the maps, and from NCBI UniSTS (now part of NCBI\ Probe).\ \
The track filter can be used to change the color or include/exclude\ a set of map data within the track. This is helpful when many items\ are shown in the track display, especially when only some are relevant\ to the current task. To use the filter: \
When you have finished configuring the filter, click the\ Submit button.
\ \This track was designed and implemented by Terry Furey. Many\ thanks to the researchers who worked on these maps, and to Greg\ Schuler, Arek Kasprzyk, Wonhee Jang, and Sanja Rogic for helping\ process the data. Additional data on the individual maps can be found\ at the following links:\
\ This track shows structural variants (SVs) identified by long-read\ whole-genome sequencing of 101 individuals, released together with the\ GWAS SVatalog\ web tool described in Chirmade et al. 2026. GWAS SVatalog computes and\ visualizes linkage disequilibrium between these SVs and GWAS-associated\ SNPs so that investigators can assess whether a SNP association signal\ may be tagging an underlying SV.\
\\ The table contains 87,068 SVs (42,435 deletions, 41,619 insertions,\ 1,394 duplications, 912 inversions, 708 complex events; byte-identical\ duplicate records have been removed). Each SV is\ annotated with gene overlaps, GC content, repeat context, ClinGen\ haploinsufficiency / triplosensitivity scores, gnomAD per-gene constraint\ metrics (pLI, LOEUF, missense O/E), OMIM phenotype associations, ClinVar\ variant IDs, and overlaps with DGV, Decipher and ClinGen regional\ annotations.\
\ \\ Items are colored by SV type:\
\ Filters are available for SV type, SV length and the number of overlapping\ genes. The detail page shows the full annotation row: gene-level constraint\ scores (per overlapping gene), ClinGen / Decipher / ClinVar region matches,\ OMIM phenotype annotations and gnomAD SV frequencies at >=90% reciprocal\ overlap. Because most genomic regions carry no clinical annotation, many\ columns will be blank for an arbitrary SV.\
\ \\ Chirmade et al. 2026 called SVs from 101 whole-genome sequenced individuals\ enrolled in the CF Canada-SickKids Program in Individualized Therapy\ (CFIT), a predominantly-European cohort of people with cystic fibrosis.\ Each sample was sequenced with two long-read / linked-read technologies:\ PacBio continuous long reads on Sequel I (34 samples, 50x) or Sequel II\ (67 samples, 76x), and 10X Genomics linked reads on Illumina HiSeq X at\ ~30x. SVs were called per sample with pbsv v2.2.2 (pbmm2 alignments) and\ Sniffles v1.0.11 (NGMLR alignments) on the PacBio CLR data, and with Long\ Ranger, CNVnator v0.4, ERDS v1.1 and Manta v1.6.0 on the 10XG data.\ Per-platform and cross-platform calls were merged in three steps using a\ 50% reciprocal overlap rule (pbsv anchored, tagged by Sniffles on PacBio;\ Manta anchored, augmented by CNVnator, ERDS and Long Ranger deletions on\ 10XG; then a cross-platform merge with PacBio coordinates preferred), and\ SV records present in fewer than three participants were dropped. The\ released catalog contains 87,183 SVs (42,435 deletions, 41,734 insertions,\ 1,394 duplications, 912 inversions and 708 complex events); the\ pre-computed GWAS SVatalog LD analyses use a common-SV subset of 35,732\ sites against 116,870 GWAS-Catalog SNPs.\
\\ The annotation TSV sv_annotations.tsv was downloaded from the\ Zenodo companion record,\ \ zenodo.org/records/13367574. Coordinates in the TSV are 1-based closed\ and were converted to 0-based half-open BED for this track.\
\\ The step-by-step build commands (download, coordinate shift, format\ conversion, bigBed build) are recorded in the UCSC makeDoc for this track\ container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ The data can be explored interactively in table format with the\ Table Browser or the\ Data Integrator, and accessed\ programmatically through our API,\ track=chirmade101Sv.\
\\ The bigBed is available from\ our\ download server as chirmade101.bb. Example:\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/lrSv/chirmade101.bb -chrom=chr21 -start=0 -end=100000000 stdout.\
\\ The original annotation table is available on Zenodo:\ zenodo.org/records/13367574.\ The GWAS SVatalog web tool itself is at\ svatalog.research.sickkids.ca.\
\ \\ Thanks to Chirmade, Strug and colleagues at The Hospital for Sick Children\ and the University of Toronto for releasing this annotated long-read SV\ callset alongside the GWAS SVatalog tool.\
\ \\ Chirmade S, Wang Z, Mastromatteo S, Sanders E, Thiruvahindrapuram B, Nalpathamkalam T, Pellecchia G,\ Lin F, Keenan K, Patel RV et al.\ \ GWAS SVatalog: a visualization tool to aid fine-mapping of GWAS loci with structural variations.\ Heredity (Edinb). 2026 Mar;135(3):199-210.\ PMID: 41203876; PMC: PMC13031531\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/chirmade101.bb\ filter.geneCount 0:200\ filter.insLen 0:31711\ filter.svLen 0:1321484\ filterByRange.geneCount on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.geneCount Gene Count\ filterLabel.insLen Insertion Length\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterType.svType multipleListOr\ filterValues.svType DEL,INS,DUP,INV,CPX\ itemRgb on\ longLabel Structural Variants from 101 Long-read WGS (GWAS SVatalog, Chirmade 2026)\ mouseOver Var: $name ($svType)\
This track shows data from \
The Tabula Sapiens: a multiple organ single cell\
transcriptomic atlas of humans. The dataset covers ~500,000 cells from\
a total of 24 human tissues and organs from all regions of the body using both \
droplet-based and plate-based single-cell RNA-sequencing (scRNA-seq). \
Samples were taken from the human bladder, blood,\
bone marrow, eye, fat, heart, kidney, large intestine, liver, lung, lymph node,\
mammary, muscle, pancreas, prostate, salivary gland, skin, small intestine,\
spleen, thymus, tongue, trachea, uterus, and vasculature. The dataset includes\
264,009 immune cells, 102,580 epithelial cells, 32,701 endothelial cells, and\
81,529 stromal cells. A total of 475 distinct cell types were identified.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ This track collection contains two bar chart tracks of RNA expression.\ The first track,\ Tabula Tissue Cell\ allows cells to be grouped together and faceted on up to 3 categories: tissue, cell class, and cell\ type. The second track,\ Tabula Details\ allows cells to be grouped together and faceted on up to 7 categories: tissue,\ cell class, cell type, subtissue, sex, donor, and assay.\
\ \\ Please see \ tabula-sapiens-portal.ds.czbiohub.org \ for further interactive displays and additional data.
\ \\ The cell types are colored by which compartment they belong to according to the following table.\ In addition, cells found in the \ Tabula Details\ track with less than 100 transcripts will be a lighter shade and less\ concentrated in color to represent a low number of transcripts.\
\ \\
| Color | \Cell Compartment | \
|---|---|
| epithelial | |
| endothelial | |
| germline | |
| immune | |
| stromal |
\ 36 tissue specimens comprising 24 unique tissues and organs were collected from \ 15 human donors (TSP1-15) with a mean age of 51 years. Tissue specimens were collected at\ various hospital locations in the Northern California region and transported on\ ice in less than one hour to preserve cell viability. Single cell suspensions\ from each organ were prepared in tissue expert laboratories at Stanford and\ UCSF. For each tissue, the dissociated cells were sorted using MACS and FACS to\ balance immune, stromal, epithelial, and endothelial cell types.\
\ \\ Sequencing libraries for all tissues were prepared using 10x 3' v3.1, 10x 5' v2, and\ Smart-seq2 (SS2) protocols for Illumina sequencing. Two 10x reactions per organ were\ loaded with 7,000 cells each with the goal to yield 10,000 QC-passed cells.\ Four 384-well Smartseq2 plates were run per organ. In most organs, one plate\ was used for each compartment (epithelial, endothelial, immune, and stromal),\ however, to capture rare cells, some organ experts allocated cells across the\ four plates differently. \ Sequencing runs for droplet libraries were loaded onto the NovaSeq S4 flow cell in sets\ of 16 to 20 libraries of approximately 5,000 cells per library with the goal of generating\ 50,000 to 75,000 reads per cell. Plate libraries were run in sets of 20 plates on Novaseq\ S4 flow cells to allow generating 1M reads per cell, depending on library quality. 152 10x\ reactions were performed, yielding 454,069 cells passing QC, and 161 smartseq2 plates\ were processed, yielding 27,051 cells passing QC.\
\ \\ Tissues collected from the same donor were used to study the\ clonal distribution of T cells between tissues, to understand the tissue\ specific mutation rate in B cells, and to analyze the cell cycle state and\ proliferative potential of shared cell types across tissues. RNA splicing\ analysis was also used to characterize cell type specific splicing and its\ variation across individuals.\
\ \\ For detailed methods and information on donors for each organ or tissue \ please refer to Quake et al, 2021 or the \ Tabula Sapiens website.\
\ \\ Some cell types, particularly in the intestines, are duplicated due to\ the use of multiple ontologies for the same cell type. In a future version,\ we plan to pool the data from these duplicates.\
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \\
The cell/gene matrix and cell-level metadata was downloaded from the \
UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\
matrixToBarChart, and bedToBigBed were used to transform these into a bar\
chart format bigBed file that can be visualized.\
The UCSC utilities can be found on\
our download server.
Thanks to the Tabula Sapiens Consortium who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent, Brittney\ Wick, and Rachel Schwartz.
\ \\ The Tabula Sapiens Consortium, Quake SR., The Tabula Sapiens: A Multiple Organ Single Cell\ Transcriptomic Atlas of Humans. bioRxiv. 2021 March 4.; doi:\ https://doi.org/10.1101/2021.07.19.452956.\
\ \ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/tabulaSapiens/facet_detailed.categories\ barChartFacets tissue,subtissue,cell_class,cell_type,sex,donor,assay\ barChartMerge on\ barChartMetric gene/genome\ barChartStatsUrl /gbdb/hg38/bbi/tabulaSapiens/facet_detailed.facets\ barChartStretchToItem on\ barChartUnit parts per million\ bigDataUrl /gbdb/hg38/bbi/tabulaSapiens/facet_detailed.bb\ defaultLabelFields name\ html tabulaSapiens\ labelFields name,name2\ longLabel Tabula sapiens full details view\ maxWindowToDraw 10000000\ parent tabulaSapiens\ shortLabel Tabula Details\ track tabulaSapiensFullDetails\ type bigBarChart\ url https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility hide\ tabulaSapiens Tabula Sapiens Tabula Sapiens single cell RNA data from many tissues 0 100 0 0 0 127 127 127 0 0 0\
This track shows data from \
The Tabula Sapiens: a multiple organ single cell\
transcriptomic atlas of humans. The dataset covers ~500,000 cells from\
a total of 24 human tissues and organs from all regions of the body using both \
droplet-based and plate-based single-cell RNA-sequencing (scRNA-seq). \
Samples were taken from the human bladder, blood,\
bone marrow, eye, fat, heart, kidney, large intestine, liver, lung, lymph node,\
mammary, muscle, pancreas, prostate, salivary gland, skin, small intestine,\
spleen, thymus, tongue, trachea, uterus, and vasculature. The dataset includes\
264,009 immune cells, 102,580 epithelial cells, 32,701 endothelial cells, and\
81,529 stromal cells. A total of 475 distinct cell types were identified.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ This track collection contains two bar chart tracks of RNA expression.\ The first track,\ Tabula Tissue Cell\ allows cells to be grouped together and faceted on up to 3 categories: tissue, cell class, and cell\ type. The second track,\ Tabula Details\ allows cells to be grouped together and faceted on up to 7 categories: tissue,\ cell class, cell type, subtissue, sex, donor, and assay.\
\ \\ Please see \ tabula-sapiens-portal.ds.czbiohub.org \ for further interactive displays and additional data.
\ \\ The cell types are colored by which compartment they belong to according to the following table.\ In addition, cells found in the \ Tabula Details\ track with less than 100 transcripts will be a lighter shade and less\ concentrated in color to represent a low number of transcripts.\
\ \\
| Color | \Cell Compartment | \
|---|---|
| epithelial | |
| endothelial | |
| germline | |
| immune | |
| stromal |
\ 36 tissue specimens comprising 24 unique tissues and organs were collected from \ 15 human donors (TSP1-15) with a mean age of 51 years. Tissue specimens were collected at\ various hospital locations in the Northern California region and transported on\ ice in less than one hour to preserve cell viability. Single cell suspensions\ from each organ were prepared in tissue expert laboratories at Stanford and\ UCSF. For each tissue, the dissociated cells were sorted using MACS and FACS to\ balance immune, stromal, epithelial, and endothelial cell types.\
\ \\ Sequencing libraries for all tissues were prepared using 10x 3' v3.1, 10x 5' v2, and\ Smart-seq2 (SS2) protocols for Illumina sequencing. Two 10x reactions per organ were\ loaded with 7,000 cells each with the goal to yield 10,000 QC-passed cells.\ Four 384-well Smartseq2 plates were run per organ. In most organs, one plate\ was used for each compartment (epithelial, endothelial, immune, and stromal),\ however, to capture rare cells, some organ experts allocated cells across the\ four plates differently. \ Sequencing runs for droplet libraries were loaded onto the NovaSeq S4 flow cell in sets\ of 16 to 20 libraries of approximately 5,000 cells per library with the goal of generating\ 50,000 to 75,000 reads per cell. Plate libraries were run in sets of 20 plates on Novaseq\ S4 flow cells to allow generating 1M reads per cell, depending on library quality. 152 10x\ reactions were performed, yielding 454,069 cells passing QC, and 161 smartseq2 plates\ were processed, yielding 27,051 cells passing QC.\
\ \\ Tissues collected from the same donor were used to study the\ clonal distribution of T cells between tissues, to understand the tissue\ specific mutation rate in B cells, and to analyze the cell cycle state and\ proliferative potential of shared cell types across tissues. RNA splicing\ analysis was also used to characterize cell type specific splicing and its\ variation across individuals.\
\ \\ For detailed methods and information on donors for each organ or tissue \ please refer to Quake et al, 2021 or the \ Tabula Sapiens website.\
\ \\ Some cell types, particularly in the intestines, are duplicated due to\ the use of multiple ontologies for the same cell type. In a future version,\ we plan to pool the data from these duplicates.\
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \\
The cell/gene matrix and cell-level metadata was downloaded from the \
UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\
matrixToBarChart, and bedToBigBed were used to transform these into a bar\
chart format bigBed file that can be visualized.\
The UCSC utilities can be found on\
our download server.
Thanks to the Tabula Sapiens Consortium who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent, Brittney\ Wick, and Rachel Schwartz.
\ \\ The Tabula Sapiens Consortium, Quake SR., The Tabula Sapiens: A Multiple Organ Single Cell\ Transcriptomic Atlas of Humans. bioRxiv. 2021 March 4.; doi:\ https://doi.org/10.1101/2021.07.19.452956.\
\ \ singleCell 0 group singleCell\ longLabel Tabula Sapiens single cell RNA data from many tissues\ shortLabel Tabula Sapiens\ superTrack on\ track tabulaSapiens\ visibility hide\ tabulaSapiensTissueCellType Tabula Tissue Cell bigBarChart Tabula sapiens RNA by tissue and cell type 3 100 0 0 0 127 127 127 0 0 0 https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$$\
This track shows data from \
The Tabula Sapiens: a multiple organ single cell\
transcriptomic atlas of humans. The dataset covers ~500,000 cells from\
a total of 24 human tissues and organs from all regions of the body using both \
droplet-based and plate-based single-cell RNA-sequencing (scRNA-seq). \
Samples were taken from the human bladder, blood,\
bone marrow, eye, fat, heart, kidney, large intestine, liver, lung, lymph node,\
mammary, muscle, pancreas, prostate, salivary gland, skin, small intestine,\
spleen, thymus, tongue, trachea, uterus, and vasculature. The dataset includes\
264,009 immune cells, 102,580 epithelial cells, 32,701 endothelial cells, and\
81,529 stromal cells. A total of 475 distinct cell types were identified.\
\
\
The read count is calculated by taking, for this cell type and gene location, the total number of\
transcript reads divided by the number of cells, and is therefore an average or mean value.\
\ This track collection contains two bar chart tracks of RNA expression.\ The first track,\ Tabula Tissue Cell\ allows cells to be grouped together and faceted on up to 3 categories: tissue, cell class, and cell\ type. The second track,\ Tabula Details\ allows cells to be grouped together and faceted on up to 7 categories: tissue,\ cell class, cell type, subtissue, sex, donor, and assay.\
\ \\ Please see \ tabula-sapiens-portal.ds.czbiohub.org \ for further interactive displays and additional data.
\ \\ The cell types are colored by which compartment they belong to according to the following table.\ In addition, cells found in the \ Tabula Details\ track with less than 100 transcripts will be a lighter shade and less\ concentrated in color to represent a low number of transcripts.\
\ \\
| Color | \Cell Compartment | \
|---|---|
| epithelial | |
| endothelial | |
| germline | |
| immune | |
| stromal |
\ 36 tissue specimens comprising 24 unique tissues and organs were collected from \ 15 human donors (TSP1-15) with a mean age of 51 years. Tissue specimens were collected at\ various hospital locations in the Northern California region and transported on\ ice in less than one hour to preserve cell viability. Single cell suspensions\ from each organ were prepared in tissue expert laboratories at Stanford and\ UCSF. For each tissue, the dissociated cells were sorted using MACS and FACS to\ balance immune, stromal, epithelial, and endothelial cell types.\
\ \\ Sequencing libraries for all tissues were prepared using 10x 3' v3.1, 10x 5' v2, and\ Smart-seq2 (SS2) protocols for Illumina sequencing. Two 10x reactions per organ were\ loaded with 7,000 cells each with the goal to yield 10,000 QC-passed cells.\ Four 384-well Smartseq2 plates were run per organ. In most organs, one plate\ was used for each compartment (epithelial, endothelial, immune, and stromal),\ however, to capture rare cells, some organ experts allocated cells across the\ four plates differently. \ Sequencing runs for droplet libraries were loaded onto the NovaSeq S4 flow cell in sets\ of 16 to 20 libraries of approximately 5,000 cells per library with the goal of generating\ 50,000 to 75,000 reads per cell. Plate libraries were run in sets of 20 plates on Novaseq\ S4 flow cells to allow generating 1M reads per cell, depending on library quality. 152 10x\ reactions were performed, yielding 454,069 cells passing QC, and 161 smartseq2 plates\ were processed, yielding 27,051 cells passing QC.\
\ \\ Tissues collected from the same donor were used to study the\ clonal distribution of T cells between tissues, to understand the tissue\ specific mutation rate in B cells, and to analyze the cell cycle state and\ proliferative potential of shared cell types across tissues. RNA splicing\ analysis was also used to characterize cell type specific splicing and its\ variation across individuals.\
\ \\ For detailed methods and information on donors for each organ or tissue \ please refer to Quake et al, 2021 or the \ Tabula Sapiens website.\
\ \\ Some cell types, particularly in the intestines, are duplicated due to\ the use of multiple ontologies for the same cell type. In a future version,\ we plan to pool the data from these duplicates.\
\ \\ The raw bar chart data can be\ explored interactively with the Table\ Browser, or the Data Integrator. For\ automated analysis, the data may be queried from our REST API. Please refer to our mailing\ list archives for questions, or our Data Access FAQ for more\ information.
\\ The expScores field for this track contains a comma-separated list of values for\ each cell type, and the expCount field is the size of the expScores array,\ which is the total number of cell types. The value in the expScores\ field corresponds to the read count for that cell type, and the order of the cell types\ is defined by the barChartBars line in the\ trackDb file for this track.\
\ \\
The cell/gene matrix and cell-level metadata was downloaded from the \
UCSC Cell Browser. The UCSC command line utility matrixClusterColumns,\
matrixToBarChart, and bedToBigBed were used to transform these into a bar\
chart format bigBed file that can be visualized.\
The UCSC utilities can be found on\
our download server.
Thanks to the Tabula Sapiens Consortium who worked on producing and publishing this data set. \ The data were integrated into the UCSC Genome Browser by Jim Kent, Brittney\ Wick, and Rachel Schwartz.
\ \\ The Tabula Sapiens Consortium, Quake SR., The Tabula Sapiens: A Multiple Organ Single Cell\ Transcriptomic Atlas of Humans. bioRxiv. 2021 March 4.; doi:\ https://doi.org/10.1101/2021.07.19.452956.\
\ \ singleCell 1 barChartCategoryUrl /gbdb/hg38/bbi/tabulaSapiens/bw_edit_tissue_cell_type.categories\ barChartFacets tissue,cell_class,cell_type\ barChartMetric gene/genome\ barChartStatsUrl /gbdb/hg38/bbi/tabulaSapiens/bw_edit_tissue_cell_type.facets\ barChartStretchToItem on\ barChartUnit parts per million\ bigDataUrl /gbdb/hg38/bbi/tabulaSapiens/tissue_cell_type.bb\ defaultLabelFields name\ html tabulaSapiens\ labelFields name,name2\ longLabel Tabula sapiens RNA by tissue and cell type\ parent tabulaSapiens\ shortLabel Tabula Tissue Cell\ track tabulaSapiensTissueCellType\ type bigBarChart\ url https://cells.ucsc.edu/?ds=tabula-sapiens+all&gene=$$\ urlLabel View on the UCSC Cell Browser:\ visibility pack\ strVar Tandem Repeat Variation Tandem Repeat Variation 0 100 0 0 0 127 127 127 0 0 0\ Tandem repeats are among the most polymorphic loci in the genome due to high\ rates of repeat unit insertions and deletions caused primarily by polymerase slippage\ during DNA replication.\ The Tandem Repeat Variation track contains a collection of tracks\ displaying population-level genetic variation at tandem repeat loci across\ the human genome. Short tandem repeats (STRs), also known as\ microsatellites, are consecutive repetitions of 1-6 nucleotide motifs.\ Variable Number Tandem Repeats (VNTRs) are tandem repeats of typically\ 7-100 bp.\
\ \\ This super track provides genome-wide tandem repeat annotations, allele frequency data from\ large-scale population cohorts, and curated disease-associated STR loci.
\ \Note that the gnomAD track container also includes an STR variation track, which is not part\ of the container here.
\ \\ Thanks to the data providers of the individual tracks listed above.\ See each track's documentation page for specific credits.
\ varRep 0 group varRep\ longLabel Tandem Repeat Variation\ pennantIcon New red ../goldenPath/newsarch.html#041026 "Released Apr. 10, 2026"\ shortLabel Tandem Repeat Variation\ superTrack on\ track strVar\ visibility hide\ targets_view Targets bigBed Capture long-seq long-read lncRNAs 1 100 0 0 0 127 127 127 0 0 0 rna 1 longLabel Capture long-seq long-read lncRNAs\ noScoreFilter on\ parent clsLongReadRnaTrack\ shortLabel Targets\ track targets_view\ type bigBed\ view targets_view\ visibility dense\ gdcCancer TCGA Pan-Cancer bigLolly 12 + TCGA Pan-Cancer mutations: 33 TCGA Cancer Projects Summary (Pan-Can 33) 0 100 0 0 0 127 127 127 0 0 0\ This track shows the genomic positions of somatic variants found through whole genome sequencing of tumors\ as part of The Cancer Genome Atlas (TCGA) by the National Cancer Institute, made available through\ the Genomic Data Commons Portal. The\ data shown here is sometimes called the "Pan-Cancer dataset", a collection of thirty-three\ TCGA projects processed in a uniform way.
\ \\ Variants can be filtered by project ID and gender from the track details page. Pressing the\ "All" button allows the user to specify whether the checked values all have to be\ true of a particular variant, or if only one of them need be present to satisfy the filter.
\ \\ The vertical viewing range in full mode can also be used to filter what variants are shown. Variants\ that have a sampleCount more or less than the min and max values specificed in the viewing range are\ not displayed.
\ \\ The raw data can be explored interactively with the Table Browser or the Data\ Integrator.\ \
\ For automated download and analysis, the genome annotation for all the thirty-three projects is\ stored in a bigBed file that can be downloaded from\ our\ download server. There are also bigBed files for each of the thirty-three projects in that\ directory. Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here. The tool can also be used to obtain only features within a given range,\ e.g.,
\\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/gdcCancer/gdcCancer.bb -chrom=chr21 -start=0 -end=100000000 stdout\\ \ \
\ All MuTect Variant calls were downloaded from the GDC portal in January 2019 and reformatted at UCSC\ to the bigBed format with a short\ script, cancerMafToBigBed.\
\ \\ Thanks to GDC for making the TCGA data available on their web site.\
\ phenDis 1 compositeTrack on\ genderFilterType multipleListOr\ genderFilterValues male,female\ group phenDis\ longLabel TCGA Pan-Cancer mutations: 33 TCGA Cancer Projects Summary (Pan-Can 33)\ maxItems 500000\ project_idFilterType multipleListOr\ project_idFilterValues TCGA-LAML|Acute Myeloid Leukemia,TCGA-ACC|Adrenocortical carcinoma,TCGA-BLCA|Bladder Urothelial Carcinoma,TCGA-LGG|Brain Lower Grade Glioma,TCGA-BRCA|Breast invasive carcinoma,TCGA-CESC|Cervical squamous cell carcinoma and endocervical adenocarcinoma,TCGA-CHOL|Cholangiocarcinoma,TCGA-LCML|Chronic Myelogenous Leukemia,TCGA-COAD|Colon adenocarcinoma,TCGA-CNTL|Controls,TCGA-ESCA|Esophageal carcinoma,TCGA-FPPP|FFPE Pilot Phase II,TCGA-GBM|Glioblastoma multiforme,TCGA-HNSC|Head and Neck squamous cell carcinoma,TCGA-KICH|Kidney Chromophobe,TCGA-KIRC|Kidney renal clear cell carcinoma,TCGA-KIRP|Kidney renal papillary cell carcinoma,TCGA-LIHC|Liver hepatocellular carcinoma,TCGA-LUAD|Lung adenocarcinoma,TCGA-LUSC|Lung squamous cell carcinoma,TCGA-DLBC|Lymphoid Neoplasm Diffuse Large B-cell Lymphoma,TCGA-MESO|Mesothelioma,TCGA-MISC|Miscellaneous,TCGA-OV|Ovarian serous cystadenocarcinoma,TCGA-PAAD|Pancreatic adenocarcinoma,TCGA-PCPG|Pheochromocytoma and Paraganglioma,TCGA-PRAD|Prostate adenocarcinoma,TCGA-READ|Rectum adenocarcinoma,TCGA-SARC|Sarcoma,TCGA-SKCM|Skin Cutaneous Melanoma,TCGA-STAD|Stomach adenocarcinoma,TCGA-TGCT|Testicular Germ Cell Tumors,TCGA-THYM|Thymoma,TCGA-THCA|Thyroid carcinoma,TCGA-UCS|Uterine Carcinosarcoma,TCGA-UCEC|Uterine Corpus Endometrial Carcinoma,TCGA-UVM|Uveal Melanoma\ shortLabel TCGA Pan-Cancer\ track gdcCancer\ type bigLolly 12 +\ visibility hide\ gnomADPextTestis Testis bigWig 0 1 gnomAD pext Testis 0 100 170 170 170 212 212 212 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Testis.bw\ color 170,170,170\ longLabel gnomAD pext Testis\ parent gnomadPext off\ shortLabel Testis\ track gnomADPextTestis\ visibility hide\ adult_testis_models Testis models bigBed 12 + Adult Testis transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-Testis.bb\ longLabel Adult Testis transcript models\ parent sample_models_view on\ shortLabel Testis models\ subGroups view=sample_models_view sample=adult_testis type=models\ track adult_testis_models\ type bigBed 12 +\ visibility squish\ adult_testis_ont_post_models Testis ONT post models bigBed 12 + Adult Testis ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Testis01Rep1.bb\ itemRgb on\ longLabel Adult Testis ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Testis ONT post models\ subGroups view=per_expr_models_view sample=adult_testis type=post_capture_ont_models\ track adult_testis_ont_post_models\ type bigBed 12 +\ visibility hide\ adult_testis_ont_post_reads Testis ONT post reads bam Adult Testis ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_Testis01Rep1.bam\ longLabel Adult Testis ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Testis ONT post reads\ subGroups view=per_expr_reads_view sample=adult_testis type=post_capture_ont_reads\ track adult_testis_ont_post_reads\ type bam\ visibility hide\ adult_testis_ont_pre_models Testis ONT pre models bigBed 12 + Adult Testis ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Testis01Rep1.bb\ itemRgb on\ longLabel Adult Testis ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Testis ONT pre models\ subGroups view=per_expr_models_view sample=adult_testis type=pre_capture_ont_models\ track adult_testis_ont_pre_models\ type bigBed 12 +\ visibility hide\ adult_testis_ont_pre_reads Testis ONT pre reads bam Adult Testis ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_Testis01Rep1.bam\ longLabel Adult Testis ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Testis ONT pre reads\ subGroups view=per_expr_reads_view sample=adult_testis type=pre_capture_ont_reads\ track adult_testis_ont_pre_reads\ type bam\ visibility hide\ adult_testis_pacbio_post_models Testis PB post models bigBed 12 + Adult Testis PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Testis01Rep1.bb\ itemRgb on\ longLabel Adult Testis PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Testis PB post models\ subGroups view=per_expr_models_view sample=adult_testis type=post_capture_pacbio_models\ track adult_testis_pacbio_post_models\ type bigBed 12 +\ visibility hide\ adult_testis_pacbio_post_reads Testis PB post reads bam Adult Testis PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_Testis01Rep1.bam\ longLabel Adult Testis PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Testis PB post reads\ subGroups view=per_expr_reads_view sample=adult_testis type=post_capture_pacbio_reads\ track adult_testis_pacbio_post_reads\ type bam\ visibility hide\ adult_testis_pacbio_pre_models Testis PB pre models bigBed 12 + Adult Testis PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Testis01Rep1.bb\ itemRgb on\ longLabel Adult Testis PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Testis PB pre models\ subGroups view=per_expr_models_view sample=adult_testis type=pre_capture_pacbio_models\ track adult_testis_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ adult_testis_pacbio_pre_reads Testis PB pre reads bam Adult Testis PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_Testis01Rep1.bam\ longLabel Adult Testis PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Testis PB pre reads\ subGroups view=per_expr_reads_view sample=adult_testis type=pre_capture_pacbio_reads\ track adult_testis_pacbio_pre_reads\ type bam\ visibility hide\ gnomADPextThyroid Thyroid bigWig 0 1 gnomAD pext Thyroid 0 100 0 102 0 127 178 127 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Thyroid.bw\ color 0,102,0\ longLabel gnomAD pext Thyroid\ parent gnomadPext off\ shortLabel Thyroid\ track gnomADPextThyroid\ visibility hide\ adult_tpoola_models Tissue Pool models bigBed 12 + Tissue Pool transcript models 4 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/cls-models-TpoolA.bb\ longLabel Tissue Pool transcript models\ parent sample_models_view on\ shortLabel Tissue Pool models\ subGroups view=sample_models_view sample=adult_tpoola type=models\ track adult_tpoola_models\ type bigBed 12 +\ visibility squish\ adult_tpoola_ont_post_models Tissue Pool ONT post models bigBed 12 + Tissue Pool ONT post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_TpoolA01Rep1.bb\ itemRgb on\ longLabel Tissue Pool ONT post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Tissue Pool ONT post models\ subGroups view=per_expr_models_view sample=adult_tpoola type=post_capture_ont_models\ track adult_tpoola_ont_post_models\ type bigBed 12 +\ visibility hide\ adult_tpoola_ont_post_reads Tissue Pool ONT post reads bam Tissue Pool ONT post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/ont-Crg-CapTrap_Hv3_0+_TpoolA01Rep1.bam\ longLabel Tissue Pool ONT post-capture reads\ parent per_expr_reads_view off\ shortLabel Tissue Pool ONT post reads\ subGroups view=per_expr_reads_view sample=adult_tpoola type=post_capture_ont_reads\ track adult_tpoola_ont_post_reads\ type bam\ visibility hide\ adult_tpoola_ont_pre_models Tissue Pool ONT pre models bigBed 12 + Tissue Pool ONT pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_TpoolA01Rep1.bb\ itemRgb on\ longLabel Tissue Pool ONT pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Tissue Pool ONT pre models\ subGroups view=per_expr_models_view sample=adult_tpoola type=pre_capture_ont_models\ track adult_tpoola_ont_pre_models\ type bigBed 12 +\ visibility hide\ adult_tpoola_ont_pre_reads Tissue Pool ONT pre reads bam Tissue Pool ONT pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/ont-Crg-CapTrap_HpreCap_0+_TpoolA01Rep1.bam\ longLabel Tissue Pool ONT pre-capture reads\ parent per_expr_reads_view off\ shortLabel Tissue Pool ONT pre reads\ subGroups view=per_expr_reads_view sample=adult_tpoola type=pre_capture_ont_reads\ track adult_tpoola_ont_pre_reads\ type bam\ visibility hide\ adult_tpoola_pacbio_post_models Tissue Pool PB post models bigBed 12 + Tissue Pool PacBio post-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_TpoolA01Rep1.bb\ itemRgb on\ longLabel Tissue Pool PacBio post-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Tissue Pool PB post models\ subGroups view=per_expr_models_view sample=adult_tpoola type=post_capture_pacbio_models\ track adult_tpoola_pacbio_post_models\ type bigBed 12 +\ visibility hide\ adult_tpoola_pacbio_post_reads Tissue Pool PB post reads bam Tissue Pool PacBio post-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/post-capture/pacBioSII-Cshl-CapTrap_Hv3_0+_TpoolA01Rep1.bam\ longLabel Tissue Pool PacBio post-capture reads\ parent per_expr_reads_view off\ shortLabel Tissue Pool PB post reads\ subGroups view=per_expr_reads_view sample=adult_tpoola type=post_capture_pacbio_reads\ track adult_tpoola_pacbio_post_reads\ type bam\ visibility hide\ adult_tpoola_pacbio_pre_models Tissue Pool PB pre models bigBed 12 + Tissue Pool PacBio pre-capture transcript models 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_TpoolA01Rep1.bb\ itemRgb on\ longLabel Tissue Pool PacBio pre-capture transcript models\ noScoreFilter on\ parent per_expr_models_view off\ shortLabel Tissue Pool PB pre models\ subGroups view=per_expr_models_view sample=adult_tpoola type=pre_capture_pacbio_models\ track adult_tpoola_pacbio_pre_models\ type bigBed 12 +\ visibility hide\ adult_tpoola_pacbio_pre_reads Tissue Pool PB pre reads bam Tissue Pool PacBio pre-capture reads 0 100 0 0 0 127 127 127 0 0 0 rna 1 bigDataUrl /gbdb/hg38/clsLongReadRna/pre-capture/pacBioSII-Cshl-CapTrap_HpreCap_0+_TpoolA01Rep1.bam\ longLabel Tissue Pool PacBio pre-capture reads\ parent per_expr_reads_view off\ shortLabel Tissue Pool PB pre reads\ subGroups view=per_expr_reads_view sample=adult_tpoola type=pre_capture_pacbio_reads\ track adult_tpoola_pacbio_pre_reads\ type bam\ visibility hide\ tommoJpSv ToMMo 333 SVs bigBed 9 + Structural Variants from 333 Japanese Individuals (ToMMo, 111 Trios) 0 100 0 0 0 127 127 127 0 0 0\ This track shows structural variants (SVs) identified by Oxford Nanopore long-read\ sequencing of 333 Japanese individuals from the Tohoku Medical Megabank (ToMMo)\ project. The 333 individuals form 111 parent-offspring trios, enabling\ Mendelian consistency checks on the SV calls. Activated T lymphocytes were used\ as a source of high-molecular-weight DNA for nanopore sequencing at a median\ coverage of 22.2x with an N50 read length of 25.8 kb.\
\\ The dataset contains 74,201 SVs (37,981 deletions and 36,220 insertions),\ merged across individuals using SURVIVOR v1.0.6. Over 95% of the SVs are\ concordant with Mendelian inheritance in the trio families.\
\ \\ Items are colored by SV type:\
\ Filters are available for SV type, SV length, and allele frequency.\ For insertions, the item is placed at the insertion site with a width of 1 bp;\ for deletions, the item spans the deleted region.\
\\ The detail page for each item shows:\
\ Otsuki et al. 2022 extracted high-molecular-weight genomic DNA from activated\ T lymphocytes of 333 individuals (111 parent-offspring trios) from the Tohoku\ Medical Megabank (ToMMo) BirThree cohort and performed Oxford Nanopore\ whole-genome sequencing on PromethION instruments with R9.4.1 flow cells\ (SQK-LSK109 libraries, Guppy v4.2.2 high-accuracy base-calling). After QC,\ median per-sample sequencing coverage was 22.2x with a read N50 of 25.8 kb.\ Reads were aligned to GRCh38 with LRA, SVs were called per sample with\ CuteSV\ v1.0.9 (-min_sv_length 50), and per-sample calls were merged with\ SURVIVOR\ v1.0.6 (1000 bp distance, type-match, no length-match) into a nonredundant\ panel of 74,201 autosomal SVs (37,981 deletions and 36,220 insertions).\ Over 95% of the SVs were concordant with Mendelian inheritance in the 111\ trio families; allele frequencies in this track are computed from the 222\ unrelated parents to avoid double-counting.\
\\ The site-only VCF tommo-JSV1-20211208-GRCh38-without-genotype-count.vcf.gz\ was downloaded from the jMorp JSV1 download page,\ \ tommo-jsv1-20211208-af.\
\\ The step-by-step build commands (download, format conversion, bigBed build)\ are recorded in the UCSC makeDoc for this track container:\ \ doc/hg38/lrSv.txt. The conversion scripts and autoSql schemas live in\ \ makeDb/scripts/lrSv, and the track configuration is in trackDb/human/lrSv.ra.\
\ \\ Source data is available from the\ tommo-jsv1-20211208-af download page on the jMorp\ portal (ToMMo Japanese Multi Omics Reference Panel).\
\ \\ The information in the ToMMo jMorp database is provided only to persons\ who agree to jMorp's\ \ Conditions of Use. By using these data, you are deemed to have read\ and understood those conditions and to agree to the following obligations:\
\ Thanks to the Tohoku Medical Megabank Organization for making their structural\ variant calls publicly available through the jMorp data portal.\
\ \\ Otsuki A, Okamura Y, Ishida N, Tadaka S, Takayama J, Kumada K, Kawashima J, Taguchi K, Minegishi N,\ Kuriyama S et al.\ \ Construction of a trio-based structural variation panel utilizing activated T lymphocytes and long-\ read sequencing technology.\ Commun Biol. 2022 Sep 20;5(1):991.\ PMID: 36127505; PMC: PMC9489684\
\ \ varRep 1 bigDataUrl /gbdb/hg38/lrSv/tommoJp.bb\ filter.AC 0:444\ filter.alleleFreq 0:1\ filter.insLen 0:30649\ filter.svLen 0:99985\ filterByRange.AC on\ filterByRange.alleleFreq on\ filterByRange.insLen on\ filterByRange.svLen on\ filterLabel.AC Allele Count\ filterLabel.alleleFreq Allele Frequency\ filterLabel.insLen Insertion Length\ filterLabel.svLen SV Length\ filterLabel.svType SV Type\ filterLimits.alleleFreq 0:1\ filterType.svType multipleListOr\ filterValues.svType DEL,INS\ itemRgb on\ longLabel Structural Variants from 333 Japanese Individuals (ToMMo, 111 Trios)\ mouseOver Var: $name ($svType)\ This track shows allele count distributions for 174,300 short tandem repeat (STR)\ loci genotyped across 61,000 Japanese individuals by the\ Tohoku Medical Megabank\ Organization (ToMMo). STR genotyping was performed with\ Expansion Hunter,\ which estimates repeat copy numbers from short-read whole-genome sequencing data.\
\ \\ For each locus, the track provides the repeat motif, the reference copy number, the\ mean and median copy number across the cohort, and a histogram of allele counts\ by repeat size. Click on any locus to see the allele count distribution as a\ bar chart.\
\ \\ Items are colored by expected heterozygosity, computed as\ het = 1 − ∑pi2 from allele counts\ across the 61,000 Japanese individuals:\
\\ The allele count histogram on the detail page shows the number of alleles observed\ at each repeat copy number. The reference allele count is computed as AN minus the\ sum of all alternate allele counts.\
\ \\ Genomic DNA was obtained from peripheral blood, saliva, or cord blood samples\ from participants in the Tohoku Medical Megabank Project. Whole-genome sequencing\ was performed on multiple Illumina and MGI platforms (HiSeq 2500, NovaSeq 6000,\ DNBSeq-T7). STR genotyping was performed with\ Expansion Hunter,\ which uses paired-end reads and read pairs spanning, flanking, and fully contained\ within repeat regions to estimate repeat copy numbers.\
\\ At UCSC, the Expansion Hunter VCF was converted to bigBed format using a\ custom Python script.\ For each STR locus, the <STRn> symbolic alleles in the VCF ALT field encode\ the repeat copy number, and the INFO/AC field provides the allele count for each.\ The reference allele count was computed as AN minus the sum of all alternate AC values.\ These were assembled into a histogram of copies=count pairs for display.\
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator.\ The data can be accessed from scripts through our\ API, the track name is tommoStr.\
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed\ file that can be downloaded from\ our download server.\ The file for this track is called tommoStr.bb.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a\ precompiled binary for your system. Instructions for downloading source code and\ binaries can be found\ here.\ The tool can also be used to obtain features within a given range, e.g.\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/strVar/tommoStr.bb\ -chrom=chr21 -start=0 -end=100000000 stdout\
\ \\ The original data can be downloaded from the\ jMorp 61KJPN-STR Downloads page.\ Use of the data requires agreement to the\ ToMMo conditions of use.\
\ \\ Thanks to the Tohoku Medical Megabank Organization and the participants of the\ ToMMo cohort study for making this data publicly available.\
\ \\ Tadaka S, Hishinuma E, Komaki S, Motoike IN, Kawashima J,\ Saigusa D, Inoue J, Takayama J, Okamura Y, Aoki Y\ et al.\ \ jMorp updates in 2020: large enhancement of multi-omics data\ resources on the general Japanese population.\ Nucleic Acids Res. 2021 Jan 8;49(D1):D536-D544.\ PMID: 33179747; PMC: PMC7779038\
\ \\ Tadaka S, Kawashima J, Hishinuma E, Saito S, Okamura Y,\ Otsuki A, Kojima K, Komaki S, Aoki Y, Kanno T et al.\ \ jMorp: Japanese Multi-Omics Reference Panel update report\ 2023.\ Nucleic Acids Res. 2024 Jan 5;52(D1):D622-D632.\ PMID: 37930845; PMC: PMC10767895\
\ varRep 1 bigDataUrl /gbdb/hg38/strVar/tommoStr.bb\ detailsScript.histogram.alleleHist {"title":"Allele Count Distribution (61K Japanese)","xLabel":"Allele size (repeat copies)"}\ filter.het 0:1\ filterByRange.het on\ filterLimits.het 0:1\ itemRgb on\ longLabel ToMMo 61KJPN Short Tandem Repeat Allele Counts (Expansion Hunter)\ mouseOver Motif: $motif ($period bp)\ These tracks contain cDNA and gene alignments produced by\ the TransMap cross-species alignment algorithm\ from other vertebrate species in the UCSC Genome Browser.\ For closer evolutionary distances, the alignments are created using\ syntenically filtered LASTZ or BLASTZ alignment chains, resulting\ in a prediction of the orthologous genes in human. For more distant\ organisms, reciprocal best alignments are used.\
\ \ TransMap maps genes and related annotations in one species to another\ using synteny-filtered pairwise genome alignments (chains and nets) to\ determine the most likely orthologs. For example, for the mRNA TransMap track\ on the human assembly, more than 400,000 mRNAs from 25 vertebrate species were\ aligned at high stringency to the native assembly using BLAT. The alignments\ were then mapped to the human assembly using the chain and net alignments\ produced using BLASTZ, which has higher sensitivity than BLAT for diverged\ organisms.\\ Compared to translated BLAT, TransMap finds fewer paralogs and aligns more UTR\ bases.\
\ \\ This track follows the display conventions for \ PSL alignment tracks.
\\ This track may also be configured to display codon coloring, a feature that\ allows the user to quickly compare cDNAs against the genomic sequence. For more \ information about this option, click \ here.\ Several types of alignment gap may also be colored; \ for more information, click \ here.\ \
\
\ To ensure unique identifiers for each alignment, cDNA and gene accessions were\ made unique by appending a suffix for each location in the source genome and\ again for each mapped location in the destination genome. The format is:\
\ accession.version-srcUniq.destUniq\\ \ Where srcUniq is a number added to make each source alignment unique, and\ destUniq is added to give the subsequent TransMap alignments unique\ identifiers.\ \
\ For example, in the cow genome, there are two alignments of mRNA BC149621.1.\ These are assigned the identifiers BC149621.1-1 and BC149621.1-2.\ When these are mapped to the human genome, BC149621.1-1 maps to a single\ location and is given the identifier BC149621.1-1.1. However, BC149621.1-2\ maps to two locations, resulting in BC149621.1-2.1 and BC149621.1-2.2. Note\ that multiple TransMap mappings are usually the result of tandem duplications, where both\ chains are identified as syntenic.\
\ \\ The raw data for these tracks can be accessed interactively through the\ Table Browser or the\ Data Integrator.\ For automated analysis, the annotations are stored in\ bigPsl files (containing a\ number of extra columns) and can be downloaded from our\ download server, \ or queried using our API. For more \ information on accessing track data see our \ Track Data Access FAQ.\ The files are associated with these tracks in the following way:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/transMap/V5/hg38.refseq.transMapV5.bigPsl\ -chrom=chr6 -start=0 -end=1000000 stdout\ \ \
\ This track was produced by Mark Diekhans at UCSC from cDNA and EST sequence data\ submitted to the international public sequence databases by \ scientists worldwide and annotations produced by the RefSeq,\ Ensembl, and GENCODE annotations projects.
\ \\ Siepel A, Diekhans M, Brejová B, Langton L, Stevens M, Comstock CL, Davis C, Ewing B, Oommen S,\ Lau C et al.\ \ Targeted discovery of novel human exons by comparative genomics.\ Genome Res. 2007 Dec;17(12):1763-73.\ PMID: 17989246; PMC: PMC2099585\
\ \\ Stanke M, Diekhans M, Baertsch R, Haussler D.\ \ Using native and syntenically mapped cDNA alignments to improve de novo gene finding.\ Bioinformatics. 2008 Mar 1;24(5):637-44.\ PMID: 18218656\
\ \\ Zhu J, Sanborn JZ, Diekhans M, Lowe CB, Pringle TH, Haussler D.\ \ Comparative genomics search for losses of long-established genes on the human lineage.\ PLoS Comput Biol. 2007 Dec;3(12):e247.\ PMID: 18085818; PMC: PMC2134963\
\ \ genes 0 group genes\ html transMapV5\ longLabel TransMap Alignments Version 5\ shortLabel TransMap V5\ superTrack on\ track transMapV5\ trexplorer TRExplorer bigBed 9 + TRExplorer V2 Tandem Repeat Catalog 1 100 0 0 0 127 127 127 0 0 0\ The TRExplorer track displays 5,599,658 tandem repeat (TR) loci from the\ TRExplorer\ catalog. Tandem repeats are adjacent copies of a short DNA sequence motif; they include\ short tandem repeats (STRs, motifs of 1–6 bp) and variable number tandem repeats\ (VNTRs, longer motifs). TRs are among the most polymorphic and mutationally active loci\ in the human genome, contributing to gene expression variation, complex disease risk,\ and over 60 known Mendelian disorders.
\ \\ The catalog integrates loci from multiple sources, including perfect repeats in the\ reference genome, polymorphic TRs discovered in T2T assemblies and the Illumina 174k\ cohort, HipSTR catalog loci, and curated disease-associated repeat expansions. Each\ locus is annotated with repeat purity, gene context, disease associations, and\ population allele frequency data from up to three cohorts.
\ \\ Items are colored by expected heterozygosity, computed as\ het = 1 − ∑pi2 from allele counts\ pooled across the TenK10K and HPRC256 cohorts:
\\ Items are labeled by the repeat motif sequence (truncated with “..” for\ motifs longer than 25 characters). The BED score reflects repeat purity (0–1000).\ Hovering over an item shows the full motif, motif size, number of reference copies,\ repeat purity, gene annotation, and data source.
\ \\ Clicking an item opens the details page, which includes a link to the corresponding\ TRExplorer locus\ page with interactive allele frequency visualizations.
\ \\ Allele frequency histograms are available for two cohorts where genotyping was\ performed:
\\ For each cohort, two parallel fields store allele sizes (in repeat copy numbers) and\ their corresponding counts, preserving the original order for histogram visualization.\ Summary allele counts are also available for the AoU1027\ cohort (1,027 HiFi PacBio samples from the All of Us Research Program\ genotyped using TRGT-LPS).
\ \\ Loci in this catalog were compiled from multiple sources:
\\ The TRExplorer catalog was built by merging tandem repeat annotations from multiple\ reference-based and population-based discovery approaches. For each locus, the repeat\ motif, copy number, and purity were determined from the GRCh38 reference sequence.\ Gene annotations were derived from MANE Select transcripts (with fallback to Gencode).\ Population allele frequencies were obtained by genotyping large cohorts using\ ExpansionHunter and other TR genotyping tools.
\ \\ For the UCSC Genome Browser track, the source catalog (TSV format) was converted to\ bigBed format. Coordinates in the source data are already 0-based half-open (BED\ convention). Allele frequency histograms were split into parallel size and count fields\ to facilitate visualization. Items are colored by expected heterozygosity.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated\ analysis, the data may be queried from our\ REST API. The underlying bigBed\ file can be downloaded from our\ download\ server.
\ \\ The complete TRExplorer dataset and interactive tools are available from the\ TRExplorer web\ portal at the Broad Institute.
\ \Thanks to Ben Weisburd, Egor Dolzhenko, and the TRExplorer team\ for making these data available.
\ \\ Weisburd B, Dolzhenko E, Bennett MF, Danzi MC, Xu IRL,\ Tanudisastro H, Gu B, English A, Hiatt L, Mokveld T\ et al.\ \ TRExplorer: A comprehensive catalog of tandem repeat variation in the human genome.\ bioRxiv. 2024.\ doi: 10.1101/2024.10.04.615514\
\ varRep 1 bigDataUrl /gbdb/hg38/strVar/trexplorer.bb\ detailsScript.histogram.hprcAlleleHist {"title":"HPRC256 Allele Distribution","xLabel":"Allele size (repeat copies)"}\ detailsScript.histogram.tenKAlleleHist {"title":"TenK10K Allele Distribution","xLabel":"Allele size (repeat copies)"}\ filter.het 0:1\ filterByRange.het on\ filterLimits.het 0:1\ itemRgb on\ longLabel TRExplorer V2 Tandem Repeat Catalog\ mouseOver Motif: $referenceMotif ($motifSize bp)\ This track displays tRNA genes predicted by using \ tRNAscan-SE v.1.23. \
\\ tRNAscan-SE is an integrated program that uses tRNAscan (Fichant) and an A/B box motif detection \ algorithm (Pavesi) as pre-filters to obtain an initial list of tRNA candidates. \ The program then filters these candidates with a covariance model-based \ search program \ COVE (Eddy) to obtain a highly specific set of primary sequence \ and secondary structure predictions that represent 99-100% of true tRNAs \ with a false positive rate of fewer than 1 per 15 gigabases.
\\ Detailed tRNA annotations for eukaryotes, bacteria, and archaea are available at\ Genomic tRNA Database (GtRNAdb). \
\\ What does the tRNAscan-SE score mean? Anything with a score above 20 bits is likely to be\ derived from a tRNA, although this does not indicate whether the tRNA gene still encodes a \ functional tRNA molecule (i.e. tRNA-derived SINES probably do not function in the ribosome in translation).\ Vertebrate tRNAs with scores of >60.0 (bits) are likely to encode functional tRNA genes, and \ those with scores below ~45 have sequence or structural features that indicate they probably are\ no longer involved in translation. tRNAs with scores between 45-60 bits are in the "grey" zone, and may\ or may not have all the required features to be functional. In these cases, tRNAs should be inspected\ carefully for loss of specific primary or secondary structure features (usually in alignments with other\ genes of the same isotype), in order to make a better educated guess. These rough score range guides \ are not exact, nor are they based on specific biochemical studies of atypical tRNA features,\ so please treat them accordingly.\
\\ Please note that tRNA genes marked as "Pseudo" are low scoring predictions that are mostly pseudogenes or \ tRNA-derived elements. These genes do not usually fold into a typical cloverleaf tRNA secondary \ structure and the provided images of the predicted secondary structures may appear rotated.\
\ \\ Both tRNAscan-SE and GtRNAdb are maintained by the\ Lowe Lab at UCSC.\
\\ Cove-predicted tRNA secondary structures were rendered by NAVIEW (c) 1988 Robert E. Bruccoleri.\
\ \\ When making use of these data, please cite the following articles:
\\ Chan PP, Lowe TM. \ GtRNAdb: a database of transfer RNA genes detected in genomic sequence.\ Nucleic Acids Res. 2009 Jan;37(Database issue):D93-7.\ PMID: 18984615; PMC: PMC2686519\
\ \\ Eddy SR, Durbin R. \ \ RNA sequence analysis using covariance models.\ Nucleic Acids Res. 1994 Jun 11;22(11):2079-88.\ PMID: 8029015; PMC: PMC308124\
\ \\ Fichant GA, Burks C. \ \ Identifying potential tRNA genes in genomic DNA sequences.\ J Mol Biol. 1991 Aug 5;220(3):659-71.\ PMID: 1870126\
\ \\ Lowe TM, Eddy SR. \ \ tRNAscan-SE: a program for improved detection of transfer RNA genes in genomic sequence.\ Nucleic Acids Res. 1997 Mar 1;25(5):955-64.\ PMID: 9023104; PMC: PMC146525\
\ \\ Pavesi A, Conterio F, Bolchi A, Dieci G, Ottonello S.\ \ Identification of new eukaryotic tRNA genes in genomic DNA databases by a multistep weight matrix\ analysis of transcriptional control regions.\ Nucleic Acids Res. 1994 Apr 11;22(7):1247-56.\ PMID: 8165140; PMC: PMC523650\
\ genes 1 color 0,20,150\ group genes\ longLabel Transfer RNA Genes Identified with tRNAscan-SE\ nextItemButton on\ noScoreFilter .\ shortLabel tRNA Genes\ superTrack nonCodingRNAs pack\ track tRNAs\ type bed 6 +\ visibility hide\ knownAlt UCSC Alt Events bed 6 . Alternative Splicing, Alternative Promoter and Similar Events in UCSC Genes 0 100 90 0 150 172 127 202 0 0 0This track shows various types of alternative splicing and other\ events that result in more than a single transcript from the same\ gene. The label by an item describes the type of event. The events are:
\This track is based on an analysis by the txgAnalyse program of splicing graphs\ produced by the txGraph program. Both of these programs were written by Jim\ Kent at UCSC.
\ genes 1 color 90,0,150\ group genes\ longLabel Alternative Splicing, Alternative Promoter and Similar Events in UCSC Genes\ noScoreFilter .\ shortLabel UCSC Alt Events\ track knownAlt\ type bed 6 .\ visibility hide\ umap Umap bigWig Single-read and multi-read mappability by Umap 2 100 0 0 0 127 127 127 0 0 0\ These tracks indicate regions with uniquely mappable reads of particular lengths before and after\ bisulfite conversion. Both Umap and Bismap tracks contain single-read mappability and multi-read\ mappability tracks for four different read lengths: 24 bp, 36 bp, 50 bp, and 100 bp.
\\ You can use these tracks for many purposes, including filtering unreliable signal from\ sequencing assays. The Bismap track can help filter unreliable signal from sequencing assays\ involving bisulfite conversion, such as whole-genome bisulfite sequencing or reduced representation\ bisulfite sequencing.
\ \ \These tracks mark any region of the bisulfite-converted genome that is uniquely mappable by\ at least one k-mer on the specified strand. Mappability of the forward strand was\ generated by converting all instances of cytosine to thymine. Similarly, mappability of the\ reverse strand was generated by converting all instances of guanine to adenine.
\To calculate the single-read mappability, you must find the overlap of a given region with\ the region that is uniquely mappable on both strands. Regions not uniquely mappable on both\ strands or have a low multi-read mappability might bias the downstream analysis.
These tracks represent the probability that a randomly selected k-mer which overlaps\ with a given position is uniquely mappable. Multi-read mappability track is calculated for\ k-mers that are uniquely mappable on both strands, and thus there is no strand\ specification.
These tracks mark any region of the genome that is uniquely mappable by at least one\ k-mer. To calculate the single-read mappability, you must find the overlap of a given\ region with this track.
These tracks represent the probability that a randomly selected k-mer which overlaps\ with a given position is uniquely mappable.
For greater detail and explanatory diagrams, see the\ preprint, the\ Umap and Bismap project website, or the\ Umap and Bismap software\ documentation.\ \
\ The raw data can be explored interactively with the Table Browser, or the Data Integrator. For automated analysis, genome annotation is stored in a bigBed\ or bigWig file that can be downloaded from the\ download\ server. Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed or bigWigToWig, which can be compiled from the source code or\ downloaded as a precompiled binary for your system. Instructions for downloading source code and\ binaries can be found here.\ The tool can also be used to obtain only features within a given range, for example:
\ bigBedToBed -chrom=chr6 -start=0 -end=1000000\ http://hgdownload.soe.ucsc.edu/gbdb/hg38/hoffmanMappability/k24.Unique.Mappability.bb stdout\\ Please refer to our mailing list archives for questions, or our\ Data Access FAQ for more\ information.
\ \\ Anshul Kundaje (Stanford\ University) created the original Umap software in MATLAB. The original Umap repository is available\ here.\ Mehran Karimzadeh (Michael Hoffman\ lab, Princess Margaret Cancer Centre) implemented the Python version of Umap and added features,\ including Bismap.
\ \\ Karimzadeh M, Ernst C, Kundaje A, Hoffman MM.,\ Umap and Bismap:\ quantifying genome and methylome mappability\ bioRxiv bioRxiv, p. 095463, 2016.; doi: https://doi.org/10.1101/095463.
\ map 0 compositeTrack on\ group map\ html mappability\ longLabel Single-read and multi-read mappability by Umap\ parent mappability\ shortLabel Umap\ subGroup1 view Views SR=Single-read MR=Multi-read\ track umap\ type bigWig\ visibility full\ umapBigBed Umap bigBed 6 Single-read and multi-read mappability by Umap 4 100 0 0 0 127 127 127 0 0 0 map 1 longLabel Single-read and multi-read mappability by Umap\ parent umap on\ shortLabel Umap\ track umapBigBed\ type bigBed 6\ view SR\ visibility squish\ umapBigWig Umap bigWig Single-read and multi-read mappability by Umap 2 100 0 0 0 127 127 127 0 0 0 map 0 longLabel Single-read and multi-read mappability by Umap\ parent umap on\ shortLabel Umap\ track umapBigWig\ type bigWig\ view MR\ viewLimits 0:1\ visibility full\ uniprot UniProt bigBed 12 + UniProt SwissProt/TrEMBL Protein Annotations 0 100 0 0 0 127 127 127 0 0 0\ This track shows protein sequences and annotations on them from the UniProt/SwissProt database,\ mapped to genomic coordinates. \
\\ UniProt/SwissProt data has been curated from scientific publications by the UniProt staff,\ UniProt/TrEMBL data has been predicted by various computational algorithms.\ The annotations are divided into multiple subtracks, based on their "feature type" in UniProt.\ The first two subtracks below - one for SwissProt, one for TrEMBL - show the\ alignments of protein sequences to the genome, all other tracks below are the protein annotations\ mapped through these alignments to the genome.\
\ \| Track Name | \Description | \
|---|---|
| UCSC Alignment, SwissProt = curated protein sequences | \Protein sequences from SwissProt mapped to the genome. All other\ tracks are (start,end) SwissProt annotations on these sequences mapped\ through this alignment. Even protein sequences without a single curated \ annotation (splice isoforms) are visible in this track. Each UniProt protein \ has one main isoform, which is colored in dark. Alternative isoforms are \ sequences that do not have annotations on them and are colored in light-blue. \ They can be hidden with the TrEMBL/Isoform filter (see below). |
| UCSC Alignment, TrEMBL = predicted protein sequences | \Protein sequences from TrEMBL mapped to the genome. All other tracks\ below are (start,end) TrEMBL annotations mapped to the genome using\ this track. This track is hidden by default. To show it, click its\ checkbox on the track configuration page. |
| UniProt Signal Peptides | \Regions found in proteins destined to be secreted, generally cleaved from mature protein. | \
| UniProt Extracellular Domains | \Protein domains with the comment "Extracellular". | \
| UniProt Transmembrane Domains | \Protein domains of the type "Transmembrane". | \
| UniProt Cytoplasmic Domains | \Protein domains with the comment "Cytoplasmic". | \
| UniProt Polypeptide Chains | \Polypeptide chain in mature protein after post-processing. | \
| UniProt Regions of Interest | \Regions that have been experimentally defined, such as the role of a region in mediating protein-protein interactions or some other biological process. | \
| UniProt Domains | \Protein domains, zinc finger regions and topological domains. | \
| UniProt Disulfide Bonds | \Disulfide bonds. | \
| UniProt Amino Acid Modifications | \Glycosylation sites, modified residues and lipid moiety-binding regions. | \
| UniProt Amino Acid Mutations | \Mutagenesis sites and sequence variants. | \
| UniProt Protein Primary/Secondary Structure Annotations | \Beta strands, helices, coiled-coil regions and turns. | \
| UniProt Sequence Conflicts | \Differences between Genbank sequences and the UniProt sequence. | \
| UniProt Repeats | \Regions of repeated sequence motifs or repeated domains. | \
| UniProt Other Annotations | \All other annotations, e.g. compositional bias | \
\ For consistency and convenience for users of mutation-related tracks,\ the subtrack "UniProt/SwissProt Variants" is a copy of the track\ "UniProt Variants" in the track group "Phenotype and Literature", or \ "Variation and Repeats", depending on the assembly.\
\ \\ Genomic locations of UniProt/SwissProt annotations are labeled with a short name for\ the type of annotation (e.g. "glyco", "disulf bond", "Signal peptide"\ etc.). A click on them shows the full annotation and provides a link to the UniProt/SwissProt\ record for more details. TrEMBL annotations are always shown in \ light blue, except in the Signal Peptides,\ Extracellular Domains, Transmembrane Domains, and Cytoplamsic domains subtracks.
\ \\ Mouse over a feature to see the full UniProt annotation comment. For variants, the mouse over will\ show the full name of the UniProt disease acronym.\
\ \\ The subtracks for domains related to subcellular location are sorted from outside to inside of \ the cell: Signal peptide, \ extracellular, \ transmembrane, and cytoplasmic.\
\ \\ Features in the "UniProt Modifications" (modified residues) track are drawn in \ light green. Disulfide bonds are shown in \ dark grey. Topological domains\ in maroon and zinc finger regions in \ olive green.\
\ \\ Duplicate annotations are removed as far as possible: if a TrEMBL annotation\ has the same genome position and same feature type, comment, disease and\ mutated amino acids as a SwissProt annotation, it is not shown again. Two\ annotations mapped through different protein sequence alignments but with the same genome\ coordinates are only shown once.
\ \On the configuration page of this track, you can choose to hide any TrEMBL annotations.\ This filter will also hide the UniProt alternative isoform protein sequences because\ both types of information are less relevant to most users. Please contact us if you\ want more detailed filtering features.
\ \Note that for the human hg38 assembly and SwissProt annotations, there\ also is a public\ track hub prepared by UniProt itself, with \ genome annotations maintained by UniProt using their own mapping\ method based on those Gencode/Ensembl gene models that are annotated in UniProt\ for a given protein. For proteins that differ from the genome, UniProt's mapping method\ will, in most cases, map a protein and its annotations to an unexpected location\ (see below for details on UCSC's mapping method).
\ \\ Briefly, UniProt protein sequences were aligned to the transcripts associated\ with the protein, the top-scoring alignments were retained, and the result was\ projected to the genome through a transcript-to-genome alignment.\ Depending on the genome, the transcript-genome alignments was either\ provided by the source database (NBCI RefSeq), created at UCSC (UCSC RefSeq) or\ derived from the transcripts (Ensembl/Augustus). The transcript set is NCBI\ RefSeq for hg38, UCSC RefSeq for hg19 (due to alt/fix haplotype misplacements \ in the NCBI RefSeq set on hg19). For other genomes, RefSeq, Ensembl and Augustus \ are tried, in this order. The resulting protein-genome alignments of this process \ are available in the file formats for liftOver or pslMap from our data archive\ (see "Data Access" section below).\
\ \An important step of the mapping process protein -> transcript ->\ genome is filtering the alignment from protein to transcript. Due to\ differences between the UniProt proteins and the transcripts (proteins were\ made many years before the transcripts were made, and human genomes have\ variants), the transcript with the highest BLAST score when aligning the\ protein to all transcripts is not always the correct transcript for a protein\ sequence. Therefore, the protein sequence is aligned to only a very short list\ of one or sometimes more transcripts, selected by a three-step procedure:\
\ For strategy 2 and 3, many of the transcripts found do not differ in coding\ sequence, so the resulting alignments on the genome will be identical.\ Therefore, any identical alignments are removed in a final filtering step. The\ details page of these alignments will contain a list of all transcripts that\ result in the same protein-genome alignment. On hg38, only a handful of edge\ cases (pseudogenes, very recently added proteins) remain in 2023 where strategy\ 3 has to be used.
\ \In other words, when an NCBI or UCSC RefSeq track is used for the mapping and to align a\ protein sequence to the correct transcript, we use a three stage process:\
This system was designed to resolve the problem of incorrect mappings of\ proteins, mostly on hg38, due to differences between the SwissProt\ sequences and the genome reference sequence, which has changed since the\ proteins were defined. The problem is most pronounced for gene families\ composed of either very repetitive or very similar proteins. To make sure that\ the alignments always go to the best chromosome location, all _alt and _fix\ reference patch sequences are ignored for the alignment, so the patches are\ entirely free of UniProt annotations. Please contact us if you have feedback on\ this process or example edge cases. We are not aware of a way to evaluate the\ results completely and in an automated manner.
\\ Proteins were aligned to transcripts with TBLASTN, converted to PSL, filtered\ with pslReps (93% query coverage, keep alignments within top 1% score), lifted to genome\ positions with pslMap and filtered again with pslReps. UniProt annotations were\ obtained from the UniProt XML file. The UniProt annotations were then mapped to the\ genome through the alignment described above using the pslMap program. This approach\ draws heavily on the LS-SNP pipeline by Mark Diekhans.\ Like all Genome Browser source code, the main script used to build this track\ can be found on Github.\
\ \\ This track is automatically updated on an ongoing basis, every 2-3 months.\ The current version name is always shown on the track details page, it includes the\ release of UniProt, the version of the transcript set and a unique MD5 that is\ based on the protein sequences, the transcript sequences, the mapping file\ between both and the transcript-genome alignment. The exact transcript\ that was used for the alignment is shown when clicking a protein alignment\ in one of the two alignment tracks.\
\ \\ For reproducibility of older analysis results and for manual inspection, previous versions of this track\ are available for browsing in the form of the UCSC UniProt Archive Track Hub (click this link to connect the hub now). The underlying data of\ all releases of this track (past and current) can be obtained from our downloads server, including the UniProt\ protein-to-genome alignment.
\ \\ The raw data of the current track can be explored interactively with the\ Table Browser, or the\ Data Integrator.\ For automated analysis, the genome annotation is stored in a bigBed file that \ can be downloaded from the\ download server.\ The exact filenames can be found in the \ track configuration file. \ Annotations can be converted to ASCII text by our tool bigBedToBed\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, for example:\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/uniprot/unipStruct.bb -chrom=chr6 -start=0 -end=1000000 stdout \
\ Please refer to our\ mailing list archives\ for questions, or our\ Data Access FAQ\ for more information. \ \ \\ \
To facilitate mapping protein coordinates to the genome, we provide the\ alignment files in formats that are suitable for our command line tools. Our\ command line programs liftOver or pslMap can be used to map\ coordinates on protein sequences to genome coordinates. The filenames are\ unipToGenome.over.chain.gz (liftOver) and unipToGenomeLift.psl.gz (pslMap).
\ \Example commands:\
\ wget -q https://hgdownload.soe.ucsc.edu/goldenPath/archive/hg38/uniprot/2022_03/unipToGenome.over.chain.gz\ wget -q https://hgdownload.soe.ucsc.edu/admin/exe/linux.x86_64/liftOver\ chmod a+x liftOver\ echo 'Q99697 1 10 annotationOnProtein' > prot.bed\ liftOver prot.bed unipToGenome.over.chain.gz genome.bed\ cat genome.bed\\ \ \
\ This track was created by Maximilian Haeussler at UCSC, with a lot of input from Chris\ Lee, Mark Diekhans and Brian Raney, feedback from the UniProt staff, Alejo\ Mujica, Regeneron Pharmaceuticals and Pia Riestra, GeneDx. Thanks to UniProt for making all data\ available for download.\
\ \\ UniProt Consortium.\ \ Reorganizing the protein space at the Universal Protein Resource (UniProt).\ Nucleic Acids Res. 2012 Jan;40(Database issue):D71-5.\ PMID: 22102590; PMC: PMC3245120\
\ \\ Yip YL, Scheib H, Diemand AV, Gattiker A, Famiglietti LM, Gasteiger E, Bairoch A.\ \ The Swiss-Prot variant page and the ModSNP database: a resource for sequence and structure\ information on human protein variants.\ Hum Mutat. 2004 May;23(5):464-70.\ PMID: 15108278\
\ genes 1 allButtonPair on\ compositeTrack on\ dataVersion /gbdb/$D/uniprot/version.txt\ exonNumbers off\ group genes\ hideEmptySubtracks off\ itemRgb on\ longLabel UniProt SwissProt/TrEMBL Protein Annotations\ shortLabel UniProt\ track uniprot\ type bigBed 12 +\ urls uniProtId="http://www.uniprot.org/uniprot/$$#section_features" pmids="https://www.ncbi.nlm.nih.gov/pubmed/$$"\ visibility hide\ spMut UniProt Variants bigBed 12 + UniProt/SwissProt Amino Acid Substitutions 0 100 0 0 0 127 127 127 0 0 0NOTE:
\
This track is intended for use primarily by physicians and other\
professionals concerned with genetic disorders, by genetics researchers, and\
by advanced students in science and medicine. While the genome browser database\
is open to the public, users seeking information about a personal medical or\
genetic condition are urged to consult with a qualified physician for\
diagnosis and for answers to personal questions.
\ This track shows the genomic positions of natural and artifical amino acid variants\ in the UniProt/SwissProt database.\ The data has been curated from scientific publications by the UniProt staff.\
\ \\ Genomic locations of UniProt/SwissProt variants are labeled with the amino acid\ change at a given position and, if known, the abbreviated disease name. A\ "?" is used if there is no disease annotated at this location, but the\ protein is described as being linked to only a single disease in UniProt.\
\ \\ Mouse over a mutation to see the UniProt comments.\
\ \\ Artificially-introduced mutations are colored green and naturally-occurring variants are colored\ red. For full information about a particular variant, click the "UniProt variant" linkout. \ The "UniProt record" linkout lists all variants of a particular protein sequence.\ The "Source articles" linkout lists the articles in PubMed that originally described\ the variant(s) and were used as evidence by the UniProt curators.\
\ \\ UniProt sequences were aligned to RefSeq sequences first with BLAT, then lifted\ to genome positions with pslMap. UniProt variants were parsed from the UniProt\ XML file. The variants were then mapped to the genome through the alignment\ using the pslMap program. This mapping approach\ draws heavily on the LS-SNP pipeline by Mark Diekhans. The complete script is\ part of the kent source tree and is located in src/hg/utils/uniprotMutations. \
\ \\
The raw data can be explored interactively with the\
Table Browser, or the\
Data Integrator.\
For automated analysis, the genome annotation is stored in a bigBed file that\
can be downloaded from the\
download server.\
The underlying data file for this track is called spMut.bb. Individual \
regions or the whole genome annotation can be obtained using our tool bigBedToBed \
which can be compiled from the source code or downloaded as a precompiled binary\
for your system. Instructions for downloading source code and binaries can be found\
here. \
The tool can also be used to obtain only features within a given range, for example:\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/uniprot/spMut.bb -chrom=chr6 -start=0 -end=1000000 stdout \
\
Please refer to our\
mailing list archives\
for questions, or our\
Data Access FAQ\
for more information. \
\ This track was created by Maximilian Haeussler, with advice from Mark Diekhans and Brian Raney.\
\ \\ UniProt Consortium.\ \ Reorganizing the protein space at the Universal Protein Resource (UniProt).\ Nucleic Acids Res. 2012 Jan;40(Database issue):D71-5.\ PMID: 22102590; PMC: PMC3245120\
\ \\ Yip YL, Scheib H, Diemand AV, Gattiker A, Famiglietti LM, Gasteiger E, Bairoch A.\ \ The Swiss-Prot variant page and the ModSNP database: a resource for sequence and structure\ information on human protein variants.\ Hum Mutat. 2004 May;23(5):464-70.\ PMID: 15108278\
\ phenDis 1 bigDataUrl /gbdb/hg38/uniprot/unipMut.bb\ exonNumbers off\ group phenDis\ itemRgb on\ longLabel UniProt/SwissProt Amino Acid Substitutions\ maxWindowCoverage 10000000\ mouseOverField comments\ noScoreFilter on\ shortLabel UniProt Variants\ track spMut\ type bigBed 12 +\ urls variationId="http://www.uniprot.org/uniprot/$$" uniProtId="http://www.uniprot.org/uniprot/$$" pmids="https://www.ncbi.nlm.nih.gov/pubmed/$$"\ visibility hide\ unusualcons Unusually Conserved bed Unusually Conserved Regions - Ultracons, HARs, etc. 0 100 0 0 0 127 127 127 0 0 0These tracks show regions of unusual conservation in human relative to other organisms:
\ \| Track | \Count | \Coverage in bp | \
|---|---|---|
| Ultraconserved | \481 | \126,007 | \
| UCNEBase Chicken | \4351 | \1,415,142 | \
| UCNEBase Paralogs | \987 | \215,800 | \
| UCNEBase UGRBs | \239 | \199,269,634 | \
| Zoo Ultracons. | \4552 | \131,661 | \
| HARs | \2647 | \681,420 | \
| ZooHARs | \312 | \49,173 | \
| HAQERs | \1580 | \1,410,669 | \
| Long hCondels | \583 | \293,809 | \
| Short hCondels | \10,032 | \1,968,123 | \
| Zoo ROCCs | \595,536 | \26,995,284 | \
| Zoo UNICORNs | \423,586 | \16,155,520 | \
All tracks show the locations and name for the features. Only one, Short hCondels, has MPRA test results in the track itself.\
\ \Ultraconserved sequences: downloaded from our hg19 public track hub and lifted to hg38.\
\ \\ HARs: Downloaded from https://docpollard.org/research/.\ Converted from Excel. Lifted 2649 HAR file to hg38.\
\ \\ HAQERs: Converted from supplemental table 1 of Mangan et al, Cell 2023.\
\ \\ Long hCondels from McLean: Excel file converted manually as supplement 2 from\ https://pmc.ncbi.nlm.nih.gov/articles/PMC3071156/ and lifted to hg38 from hg18.\
\ \\ Short hCondels from Xue et al: Excel file converted manually from supplemental file 1.\ From the paper: "We constructed a chimpanzee-anchored multiple sequence alignment across 11\ vertebrate species to detect statistically significant conserved sequences (1,371,766). These\ elements ranged from being deeply conserved throughout vertebrates to being conserved only through\ primates. We then intersected our conserved elements with called deletions (2,042,706) between the\ human (hg38) and chimpanzee (panTro4) genomes to yield 43,588 putative hCONDELs."\
\ \ \\ Zoonomia UNICORNs, ZooUCEs and RoCCs: Downloaded from\ https://cgl.gi.ucsc.edu/data/cactus/zoonomia-2021-track-hub/hg38/\
\ \
\ Thanks to Katie Pollard, Hiram Clawson, James Xue, Matt Christmas\ (matthew.christmas@imbim.uu.se),\ and Mark Diekhans for providing the data.
\ \\ Bejerano G, Pheasant M, Makunin I, Stephen S, Kent WJ, Mattick JS, Haussler D.\ \ Ultraconserved elements in the human genome.\ Science. 2004 May 28;304(5675):1321-5.\ PMID: 15131266\
\ \ \\ Pollard KS, Salama SR, Lambert N, Lambot MA, Coppens S, Pedersen JS, Katzman S, King B, Onodera C,\ Siepel A et al.\ \ An RNA gene expressed during cortical development evolved rapidly in humans.\ Nature. 2006 Sep 14;443(7108):167-72.\ PMID: 16915236\
\ \ Capra JA, Erwin GD, McKinsey G, Rubenstein JL, Pollard KS.\ \ Many human accelerated regions are developmental enhancers.\ Philos Trans R Soc Lond B Biol Sci. 2013 Dec 19;368(1632):20130025.\ PMID: 24218637; PMC: PMC3826498\ \ \\ Keough KC, Whalen S, Inoue F, Przytycki PF, Fair T, Deng C, Steyert M, Ryu H, Lindblad-Toh K,\ Karlsson E et al.\ \ Three-dimensional genome rewiring in loci with human accelerated regions.\ Science. 2023 Apr 28;380(6643):eabm1696.\ PMID: 37104607; PMC: PMC10999243\
\ \\ Mangan RJ, Alsina FC, Mosti F, Sotelo-Fonseca JE, Snellings DA, Au EH, Carvalho J, Sathyan L,\ Johnson GD, Reddy TE et al.\ \ Adaptive sequence divergence forged new neurodevelopmental enhancers in humans.\ Cell. 2022 Nov 23;185(24):4587-4603.e23.\ PMID: 36423581; PMC: PMC10013929\
\ \\ McLean CY, Reno PL, Pollen AA, Bassan AI, Capellini TD, Guenther C, Indjeian VB, Lim X, Menke DB,\ Schaar BT et al.\ \ Human-specific loss of regulatory DNA and the evolution of human-specific traits.\ Nature. 2011 Mar 10;471(7337):216-9.\ PMID: 21390129; PMC: PMC3071156\
\ \\ Xue JR, Mackay-Smith A, Mouri K, Garcia MF, Dong MX, Akers JF, Noble M, Li X, Zoonomia Consortium,\ Lindblad-Toh K et al.\ \ The functional and evolutionary impacts of human-specific deletions in conserved elements.\ Science. 2023 Apr 28;380(6643):eabn2253.\ PMID: 37104592; PMC: PMC10202372\
\ \\ Dimitrieva S, Bucher P.\ \ UCNEbase--a database of ultraconserved non-coding elements and genomic regulatory blocks.\ Nucleic Acids Res. 2013 Jan;41(Database issue):D101-9.\ PMID: 23193254; PMC: PMC3531063\
\ \ compGeno 1 compositeTrack on\ group compGeno\ longLabel Unusually Conserved Regions - Ultracons, HARs, etc.\ shortLabel Unusually Conserved\ track unusualcons\ type bed\ gnomADPextUterus Uterus bigWig 0 1 gnomAD pext Uterus 0 100 255 102 255 255 178 255 0 0 0 varRep 0 bigDataUrl /gbdb/hg38/gnomAD/pext/Uterus.bw\ color 255,102,255\ longLabel gnomAD pext Uterus\ parent gnomadPext off\ shortLabel Uterus\ track gnomADPextUterus\ visibility hide\ utrAnnotUorfs UTRannotator uORFs bigGenePred ncORFs: Upstream Open Reading Frames (uORFs) from UTRannotator 3 100 0 0 0 127 127 127 0 0 0\ This track shows 44k upstream open reading frames (uORFs) in 5' UTRs of human genes,\ curated from ribosome profiling data by the\ UTRannotator\ project, annotated by UCSC with the Kozak strength and translational efficiency.\
\ \\ uORFs are small open reading frames located in the 5' UTR of mRNAs, upstream of the main\ protein-coding sequence. They play an important role in translational regulation: ribosomes\ scanning from the 5' cap may translate a uORF first, which can reduce translation of the\ downstream main ORF. Genetic variants that create or disrupt uORFs can therefore alter\ protein expression and contribute to disease.\
\ \\ UTRannotator is a plugin for the\ Ensembl\ Variant Effect Predictor (VEP) that annotates 5' UTR variants with respect to uORFs.\ It detects five types of uORF-perturbing events (AUG gained/lost, stop lost/gained, frameshift).\ This plugin needs a database of uORFs to annotate, so the authors compiled a\ curated reference set of translated small ORFs in human 5' UTRs, derived from\ ribosome profiling data in the\ sorfs.org database. This reference set\ is what is displayed in this track. Almost all of these ORFs are annotated as 5' uORFs, only \ a tiny fraction, 270 of them, are annotated as 5'UTR+3'UTR uORF, when transcripts overlap.\
\ \\ Items are displayed in bigGenePred format. Each item is labeled with the gene symbol of\ the host transcript. Color reflects the categorical Kozak consensus strength:\
\\
Strong – A/G at position −3 and G at position +4
\
Moderate – only one of those positions matches
\
Weak – neither position matches
\
non-ATG – near-cognate start codon; the Kozak rule does not apply
\
no context – chromosome edge or context unavailable\
\ The UTRannotator source data has no exon/intron structure, so each uORF is projected\ onto a same-strand host transcript whose coordinates overlap the uORF range. The host's\ exons are clipped to the uORF range, so any host intron inside the overlap becomes an\ intron of the displayed feature; a uORF that extends past either end of the host gets a\ single bridging block for the orphan portion. The primary donor pool is the\ MANE Select / MANE Plus Clinical set;\ if every MANE candidate is rejected (e.g. the original UTRannotator transcript had a\ different UTR exon boundary), the full GENCODE comprehensive set is consulted as a\ fallback. The chosen donor transcript ID is stored in intronsSource\ (none if no host was found in either pool).\
\ \\ Mouseover shows the gene symbol, uORF type, start codon, Kozak strength and\ translational efficiency, and the host transcript whose exons supplied the intron\ structure.\
\ \\ The track offers the following filters: start codon, Kozak strength, Kozak TE (range),\ uORF type (5'UTR-only vs spans into 3'UTR).\
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. The data can be accessed from\ scripts through our API; the track name is\ "utrAnnotUorfs".\
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ Individual regions or the whole genome annotation can be obtained using our tool\ bigBedToBed, which can be compiled from the source code or downloaded as a precompiled\ binary for your system. Instructions for downloading source code and binaries can be found\ here.\ The tool can also be used to obtain only features within a given range, e.g.\
\ bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/ncOrfs/utrAnnotUorfs.kozak.bb -chrom=chr21 -start=0 -end=100000000 stdout\ \\ The uORF reference data was downloaded from the\ UTRannotator\ GitHub repository (file uORF_5UTR_GRCh38_PUBLIC.txt) and converted to bigBed format\ at UCSC. Coordinates for reverse-strand uORFs were swapped to genomic orientation. Four entries\ with invalid coordinates were excluded. Host transcripts were annotated as described above. \
\ \\ Thanks to Xiaolei Zhang, Nicola Whiffin, and the UTRannotator team at the Imperial College London\ Cardiovascular Genetics group for making this data publicly available.\
\ \\ Whiffin N, Karczewski KJ, Zhang X, Chothani S, Smith MJ, Evans DG, Roberts AM, Quaife NM, Schafer S,\ Rackham O et al.\ \ Characterising the loss-of-function impact of 5' untranslated region variants in 15,708\ individuals.\ Nat Commun. 2020 May 27;11(1):2523.\ PMID: 32461616; PMC: PMC7253449\
\ \\ Zhang X, Wakeling M, Ware J, Whiffin N.\ \ Annotating high-impact 5'untranslated region variants with the UTRannotator.\ Bioinformatics. 2021 May 23;37(8):1171-1173.\ PMID: 32926138; PMC: PMC8150139\
\ genes 1 baseColorDefault genomicCodons\ baseColorUseCds given\ bigDataUrl /gbdb/hg38/ncOrfs/utrAnnotUorfs.kozak.bb\ filter.kozakTE -1:1.5\ filterByRange.kozakTE on\ filterLimits.kozakTE -1:1.5\ filterType.kozakStrength multipleListOr\ filterType.startCodon multipleListOr\ filterType.uorfType multipleListOr\ filterValues.kozakStrength Strong,Moderate,Weak,non-ATG,None\ filterValues.startCodon ATG,CTG,GTG,TTG,ACG,other,none\ filterValues.uorfType 5'UTR uORF|5'UTR-only uORF,5'UTR+3'UTR uORF|Spans into 3'UTR\ itemRgb on\ longLabel ncORFs: Upstream Open Reading Frames (uORFs) from UTRannotator\ mouseOver $name uORF ($uorfType)NOTE:
\
Some rights reserved. This work permits non-commercial use, distribution and reproduction in any\
medium, provided the original author and source are credited.\
\
License and legal information can be found on the Varaico website.
\ Varaico\ (Variation Research Advancing Insight in Complex\ Organisms) was created using\ literature mining, similar to AVADA. Varaico variants are generated by an automated process that\ extracts purely factual information about genes from scientific papers (by matching strings against\ gene names) and HGVS variant descriptions (using regular expressions). Varaico aims to reduce\ false-positive gene and variant mentions and link them together appropriately, but nonetheless, many\ variants displayed are not mapped to the genomic position intended by the authors.\
\ \Varaico Variants (suppl) contains variants extracted from supplementary data files\ using similar methods as in the Varaico track.
\ \\ For data questions, Varaico can be contacted at\ \ jbirgmei@gmail.com\ \
\ \\ Genomic locations of variants are labeled with the HGNC gene symbol and the variant change.\ Mouse over the variants to show the gene, variant, latest author/year/title, number of publications\ mentioning the variant, and variant effect.
\ \\ Clicking on an item will provide a link directly to\ Varaico to view all publications mentioning this variant.
\ \\ The items are colored based on the amount of literature support and are a gradient from the\ colors described on the table below:\
\ \\
| Color | \Level of literature support | \
|---|---|
| \ | ≥20 papers mention the variant | \
| \ | 15 papers mention the variant | \
| \ | 10 papers mention the variant | \
| \ | 5 papers mention the variant | \
| \ | 1 paper mentions the variant | \
\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be accessed from scripts through our\ API, the track name is "varaico".
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ The file for this track is called varaico.bb. Individual\ regions or the whole genome annotation can be obtained using our tool bigBedToBed,\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system.
\\ The previous Varaico Variants version is also available in our\ download archive.
\\
Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g.\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/varaico.bb -chrom=chr21 -start=0 -end=10000000 stdout
NOTE:
\
Some rights reserved. This work permits non-commercial use, distribution and reproduction in any\
medium, provided the original author and source are credited.\
\
License and legal information can be found on the Varaico website.
\ Varaico\ (Variation Research Advancing Insight in Complex\ Organisms) was created using\ literature mining, similar to AVADA. Varaico variants are generated by an automated process that\ extracts purely factual information about genes from scientific papers (by matching strings against\ gene names) and HGVS variant descriptions (using regular expressions). Varaico aims to reduce\ false-positive gene and variant mentions and link them together appropriately, but nonetheless, many\ variants displayed are not mapped to the genomic position intended by the authors.\
\ \Varaico Variants (suppl) contains variants extracted from supplementary data files\ using similar methods as in the Varaico track.
\ \\ For data questions, Varaico can be contacted at\ \ jbirgmei@gmail.com\ \
\ \\ Genomic locations of variants are labeled with the HGNC gene symbol and the variant change.\ Mouse over the variants to show the gene, variant, latest author/year/title, number of publications\ mentioning the variant, and variant effect.
\ \\ Clicking on an item will provide a link directly to\ Varaico to view all publications mentioning this variant.
\ \\ The items are colored based on the amount of literature support and are a gradient from the\ colors described on the table below:\
\ \\
| Color | \Level of literature support | \
|---|---|
| \ | ≥20 papers mention the variant | \
| \ | 15 papers mention the variant | \
| \ | 10 papers mention the variant | \
| \ | 5 papers mention the variant | \
| \ | 1 paper mentions the variant | \
\ The raw data can be explored interactively with the Table Browser\ or the Data Integrator. The data can be accessed from scripts through our\ API, the track name is "varaico".
\ \\ For automated download and analysis, the genome annotation is stored in a bigBed file that\ can be downloaded from\ our download server.\ The file for this track is called varaico.bb. Individual\ regions or the whole genome annotation can be obtained using our tool bigBedToBed,\ which can be compiled from the source code or downloaded as a precompiled\ binary for your system.
\\ The previous Varaico Variants version is also available in our\ download archive.
\\
Instructions for downloading source code and binaries can be found\
here.\
The tool\
can also be used to obtain only features within a given range, e.g.\
\
bigBedToBed http://hgdownload.soe.ucsc.edu/gbdb/hg38/bbi/varaico.bb -chrom=chr21 -start=0 -end=10000000 stdout
The tracks that are listed here contain genetic variants and links to scientific publications that \ mention them.
\\ For additional information please click on the hyperlink of the respective track above.\
\ By default, each variant is labeled with the nucleotide change. Hover over the\ feature to see more information, explained on the track details page of the particular track\ or when clicking onto the feature.
\\ For data provenance, access and descriptions, please click the documentation via the link above.\
\ phenDis 1 group phenDis\ longLabel Genetic Variants mentioned in scientific publications\ pennantIcon Updated red ../goldenPath/newsarch.html#062526 "Updated Varaico Variants and Varaico Variants (suppl) tracks to release 3 Jun. 25, 2026"\ shortLabel Variants in Papers\ superTrack on\ track varsInPubs\ type bed 3\ vistaEnhancersBb VISTA Enhancers bigBed 9 + VISTA Enhancers 0 100 0 0 0 127 127 127 0 0 0 https://enhancer.lbl.gov/vista/element?vistaId=$$This track shows potential enhancers whose activity was experimentally validated in transgenic\ mice. Most of these noncoding elements were selected for testing based on their extreme conservation\ in other vertebrates or epigenomic evidence (ChIP-Seq) of putative enhancer marks. More information\ can be found on the VISTA Enhancer Browser\ page.\
\ \Items appearing in blue (positive) indicate that a\ reproducible pattern was observed in the in vivo enhancer assay under at least one of the\ tested conditions. Items appearing in gray (negative) indicate\ that NO reproducible pattern was observed in the in vivo enhancer assay under any of the tested\ conditions. This does not exclude the possibility that this region is a reproducible enhancer active\ under different conditions, for example at an earlier or later timepoint in development.
\ \Excerpted from the Vista Enhancer Mouse Enhancer Screen Handbook and Methods page at the Lawrence Berkeley\ National Laboratory (LBNL) website:\
Most enhancer candidate sequences are identified by extreme evolutionary sequence conservation or\ by ChIP-seq. Detailed information related to enhancer identification by extreme evolutionary\ conservation can be found in the following publications:\
\Detailed information related to enhancer identification by ChIP-seq can be found in the\ following publications:
\See the Transgenic Mouse Assay section for experimental procedures that were used to perform the\ transgenic assays: Mouse Enhancer Screen Handbook and Methods\ \
UCSC converted the\ vista-data bed files for\ hg38 and mm10 into bigBed format using the bedToBigBed utility. The data for mm39 was lifted over\ from mm10. The data for hg19 was lifted over from hg38.
\ \\ VISTA Enhancers data can be explored interactively with the\ Table Browser and cross-referenced with the\ Data Integrator. For programmatic access, the track can be\ accessed using the Genome Browser's REST API. ReMap\ annotations can be downloaded from the Genome Browser's\ download server\ as a bigBed file. This compressed binary format can be remotely queried through\ command line utilities. Please note that some of the download files can be quite large.
\ \Thanks to the Lawrence Berkeley National Laboratory for providing this data.
\ \ \\ Kosicki M, Baltoumas FA, Kelman G, Boverhof J, Ong Y, Cook LE, Dickel DE, Pavlopoulos GA, Pennacchio\ LA, Visel A.\ \ VISTA Enhancer browser: an updated database of tissue-specific developmental enhancers.\ Nucleic Acids Res. 2025 Jan 6;53(D1):D324-D330.\ PMID: 39470740; PMC: PMC11701537\
\\ Visel A, Minovitsky S, Dubchak I, Pennacchio LA.\ \ VISTA Enhancer Browser--a database of tissue-specific human enhancers.\ Nucleic Acids Res. 2007 Jan;35(Database issue):D88-92.\ PMID: 17130149; PMC: PMC1716724\
\ regulation 1 bigDataUrl /gbdb/hg38/vistaEnhancers/vistaEnhancers.bb\ group regulation\ itemRgb on\ longLabel VISTA Enhancers\ mouseOverField patternExpression\ shortLabel VISTA Enhancers\ track vistaEnhancersBb\ type bigBed 9 +\ url https://enhancer.lbl.gov/vista/element?vistaId=$$\ urlLabel View on the VISTA Enhancer Browser\ webstr WebSTR bigBed 9 + WebSTR Short Tandem Repeat Loci (EnsembleTR Panel, 1000 Genomes) 1 100 0 0 0 127 127 127 0 0 0 https://webstr.ucsd.edu/locus?repeat_id=$\ The WebSTR track displays 1,710,833 short tandem repeat (STR) loci across the\ human genome from the\ WebSTR database.
\ \\ This track is based on the EnsembleTR panel for the GRCh38/hg38 assembly,\ which represents a combined set of tandem repeats genotyped by four separate methods\ (HipSTR, GangSTR, ExpansionHunter, and AdVNTR) on data from the\ 1000 Genomes Project.\ EnsembleTR\ was applied to jointly genotype all 3,550 samples, producing consensus calls at\ over 1.7 million autosomal tandem repeat loci.
\ \\ The track includes allele frequency distributions for five 1000 Genomes continental\ populations:
\\ For each population, allele frequencies are defined as the number of copies of each allele\ divided by the total number of alleles in that population. Alleles are represented as\ the number of repeat unit copies.
\ \\ Items are colored by expected heterozygosity, computed as\ het = 1 − ∑pi2 from allele frequencies\ pooled across all five 1000 Genomes populations weighted by sample count:
\\ Each item is labeled by its repeat motif and copy count. Hovering over an item shows the repeat\ motif, number of reference copies, and heterozygosity. Clicking an item links to the\ corresponding\ WebSTR locus page, which provides\ interactive allele frequency histograms and additional annotations.
\ \\ The EnsembleTR reference panel was constructed as follows:
\\ For the UCSC Genome Browser track, the source data were converted from CSV to bigBed\ format. Per-population allele frequency distributions are stored as extra bigBed fields.
\ \\ The raw data can be explored interactively with the\ Table Browser or the\ Data Integrator. For automated\ analysis, the data may be queried from our\ REST API. The underlying bigBed\ file can be downloaded from our\ download\ server.
\ \\ The complete WebSTR dataset, including additional cohorts and data types not included in\ this track, is available from the\ WebSTR web portal. Programmatic\ access to the full WebSTR database is available through the\ WebSTR REST API.
\ \\ Thanks to Melissa Gymrek (UC San Diego) and the WebSTR team for\ providing the data for this track.
\ \\ Lundström OS, Adriaan Verbiest M, Xia F, Jam HZ, Zlobec I,\ Anisimova M, Gymrek M.\ \ WebSTR: A Population-wide Database of Short Tandem Repeat Variation\ in Humans.\ J Mol Biol. 2023 Oct 15;435(20):168260.\ PMID: 37678708\
\ \\ Ziaei Jam H, Li Y, DeVito R, Mousavi N, Ma N, Lujumba I, Adam Y,\ Maksimov M, Huang B, Dolzhenko E et al.\ \ A deep population reference panel of tandem repeat variation.\ Nat Commun. 2023 Oct 23;14(1):6711.\ PMID: 37872149; PMC: PMC10593948\
\ \ varRep 1 bigDataUrl /gbdb/hg38/strVar/webstr.bb\ detailsScript.histogram.afrHist {"title":"AFR Allele Frequencies","xLabel":"Allele size (repeat copies)"}\ detailsScript.histogram.amrHist {"title":"AMR Allele Frequencies","xLabel":"Allele size (repeat copies)"}\ detailsScript.histogram.easHist {"title":"EAS Allele Frequencies","xLabel":"Allele size (repeat copies)"}\ detailsScript.histogram.eurHist {"title":"EUR Allele Frequencies","xLabel":"Allele size (repeat copies)"}\ detailsScript.histogram.sasHist {"title":"SAS Allele Frequencies","xLabel":"Allele size (repeat copies)"}\ filter.het 0:1\ filterByRange.het on\ filterLimits.het 0:1\ itemRgb on\ longLabel WebSTR Short Tandem Repeat Loci (EnsembleTR Panel, 1000 Genomes)\ mouseOver Repeat motif: $motif ($period bp)\ This track shows a multiple alignment of 447 mammalian genomes made with Cactus and constraint scores derived from it.\ To build this track, the Zoonomia 241 alignment was used as a starting point, all primates and a few outdated\ assemblies were removed and an alignment between 233 newly sequenced primates was added. See the Methods section below for details, and \ also the publications by Kuderna et al. 2023 in the Reference section.\ All alignments and operations on them were performed using the Cactus toolkit.\
\ \\ This track shows four phyloP conservation score subtracks computed from the\ 447-way Cactus alignment (and a primates subset of it):\
\ The SSREV substitution model is strand-symmetric, which avoids\ strand-dependent bias in single-base conservation scores (Pollard\ et al. 2010, supplementary section 2.4) -- relevant when analyzing\ transcript-related nucleotides such as splice sites, miRNA seed regions, or\ other strand-specific sequence features. The REV model is the standard\ phyloP model and is appropriate for general genome-wide conservation\ analysis. The primates subset tracks restrict scoring to the 233 primate\ genomes included in the alignment, useful when conservation across\ non-primate mammals would dilute primate-specific signal.\
\ \\ Downloads for data in this track are available from the directory:\
\ In full and pack display modes, conservation scores are displayed as a\ wiggle track (histogram) in which the height reflects the\ size of the score.\ The conservation wiggles can be configured in a variety of ways to\ highlight different aspects of the displayed information.\ Click the Graph configuration help link for an explanation\ of the configuration options.
\\ Pairwise alignments of each species to the human genome are\ displayed below the conservation histogram as a grayscale density plot (in\ pack mode) or as a wiggle (in full mode) that indicates alignment quality.\ In dense display mode, conservation is shown in grayscale using\ darker values to indicate higher levels of overall conservation\ as scored by phastCons.
\\ Checkboxes on the track configuration page allow selection of the\ species to include in the pairwise display.\ Note that excluding species from the pairwise display does not alter the\ conservation score display.
\\ To view detailed information about the alignments at a specific\ position, zoom the display in to 30,000 or fewer bases, then click on\ the alignment.
\ \\ The Display chains between alignments configuration option\ enables display of gaps between alignment blocks in the pairwise alignments in\ a manner similar to the Chain track display. Missing sequence in any\ assembly is highlighted in the track display by regions of yellow when zoomed\ out and by Ns when displayed at base level. The following conventions are used:\
\ Discontinuities in the genomic context (chromosome, scaffold or region) of the\ aligned DNA in the aligning species are shown as follows:\
\ When zoomed-in to the base-level display, the track shows the base\ composition of each alignment. The numbers and symbols on the Gaps\ line indicate the lengths of gaps in the human sequence at those\ alignment positions relative to the longest non-human sequence.\ If there is sufficient space in the display, the size of the gap is shown.\ If the space is insufficient and the gap size is a multiple of 3, a\ "*" is displayed; other gap sizes are indicated by "+".
\\ Codon translation is available in base-level display mode if the\ displayed region is identified as a coding segment. To display this annotation,\ select the species for translation from the pull-down menu in the Codon\ Translation configuration section at the top of the page. Then, select one of\ the following modes:\
\ Codon translation uses the following gene tracks as the basis for translation:\
\ \\
\ Table 2. Gene tracks used for codon translation.\\ Gene Track Species \ RefSeq Genes Bos mutus, Canis lupus familiaris, Carlito syrichta, Cercocebus atys, Chinchilla lanigera, Colobus angolensis, Condylura cristata, Dipodomys ordii, Elephantulus edwardii, Eptesicus fuscus, Felis catus, Felis catus fca126, Fukomys damarensis, Homo sapiens, Ictidomys tridecemlineatus, Macaca mulatta, Macaca nemestrina, Marmota marmota, Microtus ochrogaster, Miniopterus natalensis, Mus musculus, Mus pahari, Myotis brandtii, Myotis davidii, Myotis lucifugus, Odobenus rosmarus, Orcinus orca, Otolemur garnettii, Peromyscus maniculatus, Piliocolobus tephrosceles, Propithecus coquerelli, Pteropus alecto, Pteropus vampyrus, Rattus norvegicus, Rhinopithecus roxellana, Saimiri boliviensis, Sorex araneus, Sus scrofa, Theropithecus gelada, Tupaia chinensis \ Ensembl Genes Cavia aperea \ Augustus Genes Eidolon helvum, Pteronotus parnellii \ no annotation Acinonyx jubatus, Acomys cahirinus, Ailuropoda melanoleuca, Ailurus fulgens, Allactaga bullata, Allenopithecus nigroviridis, Allochrocebus lhoesti, Allochrocebus preussi, Allochrocebus solatus, Alouatta belzebul, Alouatta caraya, Alouatta discolor, Alouatta juara, Alouatta macconnelli, Alouatta nigerrima, Alouatta palliata, Alouatta puruensis, Alouatta seniculus, Ammotragus lervia, Anoura caudifer, Antilocapra americana, Aotus azarae, Aotus griseimembra, Aotus nancymaae, Aotus trivirgatus, Aotus vociferans, Aplodontia rufa, Arctocebus calabarensis, Artibeus jamaicensis, Ateles geoffroyi_a, Ateles geoffroyi_b, Ateles belzebuth, Ateles chamek, Ateles marginatus, Ateles paniscus, Avahi laniger, Avahi peyrierasi, Balaenoptera acutorostrata, Balaenoptera bonaerensis, Beatragus hunteri, Bison bison, Bos indicus, Bos taurus, Bubalus bubalis, Cacajao ayresi, Cacajao calvus, Cacajao hosomi, Cacajao melanocephalus, Callibella humilis, Callimico goeldii, Callithrix geoffroyi, Callithrix jacchus, Callithrix kuhlii, Camelus bactrianus, Camelus dromedarius, Camelus ferus, Canis lupus VD, Canis lupus dingo, Canis lupus orion, Capra aegagrus, Capra hircus, Capromys pilorides, Carollia perspicillata, Castor canadensis, Catagonus wagneri, Cavia porcellus, Cavia tschudii, Cebuella niveiventris, Cebuella pygmaea, Cebus albifrons, Cebus olivaceus, Cebus unicolor, Cephalopachus bancanus, Ceratotherium simum, Ceratotherium simum cottoni, Cercocebus chrysogaster, Cercocebus lunulatus, Cercocebus torquatus, Cercopithecus ascanius, Cercopithecus cephus, Cercopithecus diana, Cercopithecus hamlyni, Cercopithecus lowei, Cercopithecus albogularis, Cercopithecus mona, Cercopithecus neglectus, Cercopithecus nictitans, Cercopithecus petaurista, Cercopithecus pogonias, Cercopithecus roloway, Chaetophractus vellerosus, Cheirogaleus major, Cheirogaleus medius, Cheracebus lucifer, Cheracebus lugens, Cheracebus regulus, Cheracebus torquatus, Chiropotes albinasus, Chiropotes israelita, Chiropotes sagulatus, Chlorocebus aethiops, Chlorocebus pygerythrus, Chlorocebus sabaeus, Choloepus didactylus, Choloepus hoffmanni, Chrysochloris asiatica, Colobus guereza, Colobus polykomos, Craseonycteris thonglongyai, Cricetomys gambianus, Cricetulus griseus, Crocidura indochinensis, Cryptoprocta ferox, Ctenodactylus gundi, Ctenomys sociabilis, Cuniculus paca, Dasyprocta punctata, Dasypus novemcinctus, Daubentonia madagascariensis, Delphinapterus leucas, Desmodus rotundus, Dicerorhinus sumatrensis, Diceros bicornis, Dinomys branickii, Dipodomys stephensi, Dolichotis patagonum, Echinops telfairi, Elaphurus davidianus, Ellobius lutescens, Ellobius talpinus, Enhydra lutris, Equus asinus, Equus caballus, Equus przewalskii, Erinaceus europaeus, Erythrocebus patas, Eschrichtius robustus, Eubalaena japonica, Eulemur albifrons, Eulemur collaris, Eulemur coronatus, Eulemur flavifrons, Eulemur fulvus, Eulemur macaco, Eulemur mongoz, Eulemur rubriventer, Eulemur rufus, Eulemur sanfordi, Felis nigripes, Galago moholi, Galago senegalensis, Galagoides demidoff, Galeopterus variegatus, Giraffa tippelskirchi, Glis glis, Gorilla beringei, Gorilla gorilla, Graphiurus murinus, Hapalemur alaotrensis, Hapalemur gilberti, Hapalemur griseus, Hapalemur meridionalis, Hapalemur occidentalis, Helogale parvula, Hemitragus hylocrius, Heterocephalus glaber, Heterohyrax brucei, Hippopotamus amphibius, Hipposideros armiger, Hipposideros galeritus, Hoolock leuconedys, Hyaena hyaena, Hydrochoerus hydrochaeris, Hylobates abbotti, Hylobates agilis, Hylobates klossii, Hylobates pileatus, Hylobates muelleri, Hylobates pileatus, Hystrix cristata, Indri indri, Inia geoffrensis, Jaculus jaculus, Kogia breviceps, Lagothrix lagothricha, Lasiurus borealis, Lemur catta, Leontocebus fuscicollis, Leontocebus illigeri, Leontocebus nigricollis, Leontopithecus chrysomelas, Leontopithecus rosalia, Lepilemur ankaranensis, Lepilemur dorsalis, Lepilemur ruficaudatus, Lepilemur septentrionalis, Leptonychotes weddellii, Lepus americanus, Lipotes vexillifer, Lophocebus aterrimus, Loris lydekkerianus, Loris tardigradus, Loxodonta africana, Lycaon pictus, Macaca arctoides, Macaca assamensis, Macaca cyclopis, Macaca fascicularis, Macaca fuscata, Macaca leonina, Macaca maura, Macaca nigra, Macaca radiata, Macaca siberu, Macaca silenus, Macaca thibetana, Macaca tonkeana, Macroglossus sobrinus, Mandrillus leucophaeus, Mandrillus sphinx, Manis javanica, Manis pentadactyla, Megaderma lyra, Mellivora capensis, Meriones unguiculatus, Mesocricetus auratus, Mesoplodon bidens, Mico argentatus, Mico humeralifer, Mico schneideri, Microcebus murinus, Microgale talazaci, Micronycteris hirsuta, Miniopterus schreibersii, Miopithecus ogouensis, Mirounga angustirostris, Mirza zaza, Monodon monoceros, Mormoops blainvillei, Moschus moschiferus, Mungos mungo, Murina feae, Mus caroli, Mus spretus, Muscardinus avellanarius, Mustela putorius, Myocastor coypus, Myotis myotis, Myrmecophaga tridactyla, Nannospalax galili, Nasalis larvatus, Neomonachus schauinslandi, Neophocaena asiaeorientalis, Noctilio leporinus, Nomascus annamensis, Nomascus concolor, Nomascus gabriellae, Nomascus siki_a, Nomascus siki_b, Nyctereutes procyonoides, Nycticebus bengalensis, Nycticebus coucang, Nycticebus pygmaeus, Ochotona princeps, Octodon degus, Odocoileus virginianus, Okapia johnstoni, Ondatra zibethicus, Onychomys torridus, Orycteropus afer, Oryctolagus cuniculus, Otocyon megalotis, Otolemur crassicaudatus, Ovis aries, Ovis canadensis, Pan paniscus, Pan troglodytes, Panthera onca, Panthera pardus, Panthera tigris, Pantholops hodgsonii, Papio anubis, Papio cynocephalus, Papio hamadryas, Papio kindae, Papio papio, Papio ursinus, Paradoxurus hermaphroditus, Perodicticus ibeanus, Perodicticus potto, Perognathus longimembris, Petromus typicus, Phocoena phocoena, Piliocolobus badius, Piliocolobus gordonorum, Piliocolobus kirkii, Pipistrellus pipistrellus, Pithecia albicans, Pithecia chrysocephala, Pithecia hirsuta, Pithecia mittermeieri, Pithecia pissinattii, Pithecia pithecia, Pithecia vanzolinii, Platanista gangetica, Plecturocebus bernhardi, Plecturocebus brunneus, Plecturocebus caligatus, Plecturocebus cinerascens, Plecturocebus cupreus, Plecturocebus dubius, Plecturocebus grovesi, Plecturocebus hoffmannsi, Plecturocebus miltoni, Plecturocebus moloch, Pongo abelii, Pongo pygmaeus, Presbytis comata, Presbytis mitrata, Procavia capensis, Prolemur simus, Propithecus coronatus, Propithecus diadema, Propithecus edwardsi, Propithecus perrieri, Propithecus tattersalli, Propithecus verreauxi, Psammomys obesus, Pteronura brasiliensis, Puma concolor, Pygathrix cinerea, Pygathrix nigripes, Pygathrix nigripes, Rangifer tarandus, Rhinolophus sinicus, Rhinopithecus bieti, Rhinopithecus strykeri, Rousettus aegyptiacus, Saguinus bicolor, Saguinus geoffroyi, Saguinus imperator, Saguinus inustus, Saguinus labiatus, Saguinus midas, Saguinus mystax, Saguinus oedipus, Saiga tatarica, Saimiri cassiquiarensis, Saimiri macrodon, Saimiri oerstedii, Saimiri sciureus, Saimiri ustus, Sapajus apella, Sapajus macrocephalus, Scalopus aquaticus, Semnopithecus entellus, Semnopithecus hypoleucos, Semnopithecus johnii, Semnopithecus priam, Semnopithecus schistaceus, Semnopithecus vetulus, Sigmodon hispidus, Solenodon paradoxus, Spermophilus dauricus, Spilogale gracilis, Suricata suricatta, Symphalangus syndactylus, Tadarida brasiliensis, Tamandua tetradactyla, Tapirus indicus, Tapirus terrestris, Tarsius lariang, Tarsius wallacei, Thryonomys swinderianus, Tolypeutes matacus, Tonatia saurophila, Trachypithecus auratus, Trachypithecus crepusculus, Trachypithecus cristatus, Trachypithecus francoisi, Trachypithecus geei, Trachypithecus germaini, Trachypithecus hatinhensis, Trachypithecus laotum, Trachypithecus leucocephalus, Trachypithecus melamera, Trachypithecus obscurus, Trachypithecus phayrei, Trachypithecus pileatus, Tragulus javanicus, Trichechus manatus, Tupaia tana, Tursiops truncatus, Uropsilus gracilis, Ursus maritimus, Varecia rubra, Varecia variegata, Vicugna pacos, Vulpes lagopus, Xerus inauris, Zalophus californianus, Zapus hudsonius, Ziphius cavirostris\
\ This alignment was created by making three edits (using Cactus) to the\ 241-way mammalian Zoonomia Cactus alignment\ (\ https://cglgenomics.ucsc.edu/data/cactus/).\
\
phyloP scores were computed from the Cactus 447-way alignment using the\
phyloP program from the\
PHAST package.\
Per-base scores were produced with options\
--method LRT --mode CONACC --wig-scores; positive scores\
indicate conservation under purifying selection, negative scores indicate\
acceleration relative to neutral evolution.\
\
For the all-species tracks, base-composition and substitution-rate\
parameters were estimated from 4-fold degenerate sites using\
phyloFit (PHAST, EM algorithm, medium precision) under either the\
REV or strand-symmetric reversible (SSREV) substitution model. Background\
base frequencies were adjusted with modFreqs so that\
complementary bases (A/T and C/G) appear at equal expected frequencies,\
which is required for strand-symmetric scoring.\
\
For the primates-subset tracks, the alignment was restricted to the 233\
primate species and an independent phyloFit / phyloP run was performed on\
that sub-alignment using the SSREV model. All scores were encoded into\
wiggle format and loaded as either bigWig files (REV all-species,\
primates LRT) or wig SQL tables backed by .wib data files\
(SSREV all-species, SSREV primates).\
\ The phylogenic tree was established by the research described\ in A global catalog of whole-genome diversity from 233 primate\ species.\ \
\
\\ \ \\
\ count \common \
nameclade \scientific name \
(link to browser when existing)taxon id \
link to NCBI\ 001 human primates catarrhini Homo sapiens/hg38
reference species9606 \ 002 western gorilla primates catarrhini Gorilla gorilla
GCA_900006655.3_Susie39593 \ 003 Sumatran orangutan primates catarrhini Pongo abelii
GCA_002880775.3_Susie_PABv29601 \ 004 Eastern Gorilla primates catarrhini Gorilla beringei 499232 \ 005 chimpanzee primates catarrhini Pan troglodytes
GCA_002880755.3_Clint_PTRv29598 \ 006 Bornean orangutan primates catarrhini Pongo pygmaeus 9600 \ 007 Rhesus monkey primates catarrhini Macaca mulatta
rheMac109544 \ 008 gelada primates catarrhini Theropithecus gelada
GCF_003255815.1_Tgel_1.09565 \ 009 stump-tailed macaque primates catarrhini Macaca arctoides 9540 \ 010 Northern Talapoin Monkey primates catarrhini Miopithecus ogouensis 100488 \ 011 crab-eating macaque primates catarrhini Macaca fascicularis 9541 \ 012 Allen's swamp monkey primates catarrhini Allenopithecus nigroviridis 54135 \ 013 siamang primates catarrhini Symphalangus syndactylus 9590 \ 014 black crested mangabey primates catarrhini Lophocebus aterrimus 75566 \ 015 drill primates catarrhini Mandrillus leucophaeus 9568 \ 016 Bonnet Macaque primates catarrhini Macaca radiata 9548 \ 017 Red-capped Mangabey primates catarrhini Cercocebus torquatus 9530 \ 018 Golden-bellied Mangabey primates catarrhini Cercocebus chrysogaster 75569 \ 019 Owl-faced Monkey primates catarrhini Cercopithecus hamlyni 9536 \ 020 Siberut Macaque primates catarrhini Macaca siberu 244255 \ 021 pig-tailed macaque primates catarrhini Macaca nemestrina
GCF_000956065.1_Mnem_1.09545 \ 022 White-naped Mangabey primates catarrhini Cercocebus lunulatus (Cercocebus atys lunulatus) 75570 \ 023 Tonkean Macaque primates catarrhini Macaca tonkeana 40843 \ 024 Diana Monkey primates catarrhini Cercopithecus diana 36224 \ 025 red guenon primates catarrhini Erythrocebus patas 9538 \ 026 Northern Pig-tailed Macaque primates catarrhini Macaca leonina 90387 \ 027 Moor Macaque primates catarrhini Macaca maura 90383 \ 028 Guinea Baboon primates catarrhini Papio papio 100937 \ 029 hamadryas baboon primates catarrhini Papio hamadryas 9557 \ 030 liontail macaque primates catarrhini Macaca silenus 54601 \ 031 olive baboon primates catarrhini Papio anubis
GCA_000264685.2_Panu_3.09555 \ 032 Roloway Monkey primates catarrhini Cercopithecus roloway 1137049 \ 033 Kinda Baboon primates catarrhini Papio kindae 208091 \ 034 Chacma Baboon primates catarrhini Papio ursinus 36229 \ 035 Sun-tailed Monkey primates catarrhini Allochrocebus solatus 147650 \ 036 golden snub-nosed monkey primates catarrhini Rhinopithecus roxellana
GCF_007565055.1_ASM756505v161622 \ 037 Vervet Monkey primates catarrhini Chlorocebus pygerythrus 60710 \ 038 sooty mangabey primates catarrhini Cercocebus atys
GCF_000955945.1_Caty_1.09531 \ 039 green monkey primates catarrhini Chlorocebus sabaeus
GCA_000409795.2_Chlorocebus_sabeus_1.160711 \ 040 De Brazza's monkey primates catarrhini Cercopithecus neglectus 36227 \ 041 Yellow Baboon primates catarrhini Papio cynocephalus 9556 \ 042 Celebes crested macaque primates catarrhini Macaca nigra 54600 \ 043 proboscis monkey primates catarrhini Nasalis larvatus 43780 \ 044 Preuss's Monkey primates catarrhini Allochrocebus preussi 147649 \ 045 Putty-nosed Monkey primates catarrhini Cercopithecus nictitans 36228 \ 046 Javan Surili primates catarrhini Presbytis comata 78452 \ 047 Sykes' Monkey primates catarrhini Cercopithecus albogularis 36225 \ 048 LHoests Monkey primates catarrhini Allochrocebus lhoesti 100224 \ 049 Crowned Monkey primates catarrhini Cercopithecus pogonias 102108 \ 050 Southern Mitered Langur primates catarrhini Presbytis mitrata (Presbytis melalophos mitrata) 272115 \ 051 Grey-shanked Douc Langur primates catarrhini Pygathrix cinerea 693712 \ 052 Mona monkey primates catarrhini Cercopithecus mona 36226 \ 053 Spot-nosed Monkey primates catarrhini Cercopithecus petaurista 100487 \ 054 grivet primates catarrhini Chlorocebus aethiops 9534 \ 055 Lowes Monkey primates catarrhini Cercopithecus lowei 304410 \ 056 Northern Yellow-cheeked Crested Gibbon primates catarrhini Nomascus annamensis 1616038 \ 057 Red-cheeked Gibbon primates catarrhini Nomascus gabriellae 61852 \ 058 Japanese macaque primates catarrhini Macaca fuscata 9542 \ 059 Western Red Colobus primates catarrhini Piliocolobus badius 164648 \ 060 southern white-cheeked gibbon primates catarrhini Nomascus siki_a 9586 \ 061 Taiwan macaque primates catarrhini Macaca cyclopis 78449 \ 062 black-shanked douc langur primates catarrhini Pygathrix nigripes 310352 \ 063 King Colobus primates catarrhini Colobus polykomos 9572 \ 064 Black Crested Gibbon primates catarrhini Nomascus concolor 29089 \ 065 Udzungwa Red Colobus primates catarrhini Piliocolobus gordonorum 591933 \ 066 Gee's Golden Langur primates catarrhini Trachypithecus geei 164650 \ 067 Kloss's Gibbon primates catarrhini Hylobates klossii 9587 \ 068 Spectacled Leaf Monkey primates catarrhini Trachypithecus obscurus 54181 \ 069 Zanzibar Red Colobus primates catarrhini Piliocolobus kirkii 591937 \ 070 Indochinese Silvered Langur primates catarrhini Trachypithecus germaini 271260 \ 071 Hatinh Langur primates catarrhini Trachypithecus hatinhensis 867383 \ 072 Moustached Monkey primates catarrhini Cercopithecus cephus 9535 \ 073 Laotian Langur primates catarrhini Trachypithecus laotum 465718 \ 074 Francois's langur primates catarrhini Trachypithecus francoisi 54180 \ 075 Purple-faced Langur primates catarrhini Semnopithecus vetulus (Trachypithecus vetulus) 54137 \ 076 Capped Langur primates catarrhini Trachypithecus pileatus 164651 \ 077 Ugandan red Colobus primates catarrhini Piliocolobus tephrosceles
GCF_002776525.2_ASM277652v2591936 \ 078 Spangled Ebony Langur primates catarrhini Trachypithecus auratus 222416 \ 079 Red-tailed Monkey primates catarrhini Cercopithecus ascanius 36223 \ 080 Silvery Lutung primates catarrhini Trachypithecus cristatus 122765 \ 081 Nilgiri Langur primates catarrhini Semnopithecus johnii (Trachypithecus johnii) 66063 \ 082 Indochinese grey langur primates catarrhini Trachypithecus crepusculus (Trachypithecus phayrei crepuscula) 272121 \ 083 White-headed langur primates catarrhini Trachypithecus leucocephalus (Trachypithecus poliocephalus) 465719 \ 084 pygmy chimpanzee primates catarrhini Pan paniscus
GCA_000258655.2_panpan1.19597 \ 085 northern white-cheeked gibbon primates catarrhini Nomascus siki_b 9586 \ 086 Agile Gibbon primates catarrhini Hylobates agilis 9579 \ 087 Phayre's Leaf-monkey primates catarrhini Trachypithecus melamera n/a \ 088 Nepal Gray Langur primates catarrhini Semnopithecus schistaceus 2804203 \ 089 Abbott's Gray Gibbon primates catarrhini Hylobates abbotti (Hylobates muelleri abbotti) 716694 \ 090 Bornean Gibbon primates catarrhini Hylobates muelleri 9588 \ 091 Tufted Gray Langur primates catarrhini Semnopithecus priam 1208733 \ 092 Black-footed Gray Langur primates catarrhini Semnopithecus hypoleucos 1208734 \ 093 mantled guereza primates catarrhini Colobus guereza 33548 \ 094 Hanuman langur primates catarrhini Semnopithecus entellus 88029 \ 095 pileated gibbon primates catarrhini Hylobates pileatus 9589 \ 096 black snub-nosed monkey primates catarrhini Rhinopithecus bieti 61621 \ 097 Burmese snub-nosed monkey primates catarrhini Rhinopithecus strykeri 1194336 \ 098 Angolan colobus primates catarrhini Colobus angolensis
colAng154131 \ 099 Pileated Gibbon primates catarrhini Hylobates pileatus 9589 \ 100 black-shanked douc langur primates catarrhini Pygathrix nigripes 310352 \ 101 Milne-edwards' Macaque primates catarrhini Macaca thibetana 54602 \ 102 Phayre's Leaf-monkey primates catarrhini Trachypithecus phayrei 61618 \ 103 Assam macaque primates catarrhini Macaca assamensis 9551 \ 104 Eastern hoolock gibbon primates catarrhini Hoolock leuconedys 61851 \ 105 mandrill primates catarrhini Mandrillus sphinx 9561 \ 106 White-faced Saki primates platyrrhini Pithecia chrysocephala 2946515 \ 107 Monk Saki primates platyrrhini Pithecia hirsuta 2946516 \ 108 white-faced saki primates platyrrhini Pithecia pithecia 43777 \ 109 Mittermeier's Tapajós saki primates platyrrhini Pithecia mittermeieri 2946517 \ 110 Buffy Saki primates platyrrhini Pithecia albicans 2946514 \ 111 Pissinatti's saki primates platyrrhini Pithecia pissinattii (Pithecia pissinatti) 2946518 \ 112 Vanzolini's Bald-faced Saki primates platyrrhini Pithecia vanzolinii 2946519 \ 113 Bald-headed Uacari primates platyrrhini Cacajao calvus 30596 \ 114 Ayres Black Uakari primates platyrrhini Cacajao ayresi 535896 \ 115 Black-headed Uacari primates platyrrhini Cacajao melanocephalus 70825 \ 116 Black-headed Uacari primates platyrrhini Cacajao hosomi 535897 \ 117 Reddish-brown bearded saki primates platyrrhini Chiropotes sagulatus (Chiropotes chiropotes) 658221 \ 118 brown-backed bearded saki primates platyrrhini Chiropotes israelita 280163 \ 119 Collared Titi Monkey primates platyrrhini Cheracebus lugens 210166 \ 120 Brown Titi Monkey primates platyrrhini Plecturocebus brunneus 1812042 \ 121 Hoffmanns's titi monkey primates platyrrhini Plecturocebus hoffmannsi 78255 \ 122 Milton's Titi Monkey primates platyrrhini Plecturocebus miltoni 1812038 \ 123 Widow Monkey primates platyrrhini Cheracebus torquatus 30592 \ 124 Ashy Black Titi Monkey primates platyrrhini Plecturocebus cinerascens 1812037 \ 125 Prince Bernhard's Titi Monkey primates platyrrhini Plecturocebus bernhardi 1812036 \ 126 Yellow-handed Titi Monkey primates platyrrhini Cheracebus lucifer 2487712 \ 127 Coppery Titi Monkey primates platyrrhini Plecturocebus cupreus 202457 \ 128 Chestnut-bellied Titi primates platyrrhini Plecturocebus caligatus 867332 \ 129 Hershkovitzs Titi primates platyrrhini Plecturocebus dubius 2946520 \ 130 Red-bellied Titi Monkey primates platyrrhini Plecturocebus moloch 9523 \ 131 Groves' Titi primates platyrrhini Plecturocebus grovesi 2488670 \ 132 black-handed spider monkey primates platyrrhini Ateles geoffroyi_a 9509 \ 133 Widow Monkey primates platyrrhini Cheracebus regulus 1812110 \ 134 Guiana Spider Monkey primates platyrrhini Ateles paniscus 9510 \ 135 Black-faced Black Spider Monkey primates platyrrhini Ateles chamek 118643 \ 136 White-cheeked Spider Monkey primates platyrrhini Ateles marginatus 1529884 \ 137 White-bellied Spider Monkey primates platyrrhini Ateles belzebuth 9507 \ 138 Common Woolly Monkey primates platyrrhini Lagothrix lagothricha (Lagothrix lagotricha) 9519 \ 139 large-headed capuchin primates platyrrhini Sapajus macrocephalus (Sapajus apella macrocephalus) 1547595 \ 140 Spixs White-fronted Capuchin primates platyrrhini Cebus unicolor 1985288 \ 141 Central American spider monkey primates platyrrhini Ateles geoffroyi_b 9509 \ 142 Guinan Weeper Capuchin primates platyrrhini Cebus olivaceus 37295 \ 143 mantled howler monkey primates platyrrhini Alouatta palliata 30589 \ 144 white-fronted capuchin primates platyrrhini Cebus albifrons 9514 \ 145 Northern Night Monkey primates platyrrhini Aotus trivirgatus 9505 \ 146 Grey-handed Night Monkey primates platyrrhini Aotus griseimembra 292213 \ 147 Black-and-gold Howler Monkey primates platyrrhini Alouatta caraya 9502 \ 148 Spixs Night Monkey primates platyrrhini Aotus vociferans 57176 \ 149 Red-handed Howler Monkey primates platyrrhini Alouatta belzebul 30590 \ 150 Red-handed Howler Monkey primates platyrrhini Alouatta discolor 2905217 \ 151 Azara's Night Monkey primates platyrrhini Aotus azarae (Aotus azarai) 30591 \ 152 Purús Red Howler Monkey primates platyrrhini Alouatta puruensis (Alouatta seniculus puruensis) 1347729 \ 153 Black Howler Monkey primates platyrrhini Alouatta nigerrima (Alouatta belzebul) 30590 \ 154 Guianan Red Howler Monkey primates platyrrhini Alouatta macconnelli 198115 \ 155 Colombian Red Howler Monkey primates platyrrhini Alouatta juara 2946512 \ 156 Colombian Red Howler Monkey primates platyrrhini Alouatta seniculus 9503 \ 157 tufted capuchin primates platyrrhini Sapajus apella 9515 \ 158 Ma's night monkey primates platyrrhini Aotus nancymaae
GCA_000952055.2_Anan_2.037293 \ 159 Bolivian squirrel monkey primates platyrrhini Saimiri boliviensis
GCF_016699345.1_BCM_Sbol_2.027679 \ 160 White-nosed Saki primates platyrrhini Chiropotes albinasus 198627 \ 161 Black Mantle Tamarin primates platyrrhini Leontocebus nigricollis 9489 \ 162 brown-mantled tamarin primates platyrrhini Leontocebus fuscicollis 9487 \ 163 Illiger's saddle-back tamarin primates platyrrhini Leontocebus illigeri (Leontocebus fuscicollis illigeri) 881947 \ 164 Cotton-headed Tamarin primates platyrrhini Saguinus oedipus 9490 \ 165 Pied Tamarin primates platyrrhini Saguinus bicolor 37588 \ 166 Geoffroy's Tamarin primates platyrrhini Saguinus geoffroyi 43778 \ 167 White-fronted Titi Monkey primates platyrrhini Saguinus inustus 1079039 \ 168 Moustached Tamarin primates platyrrhini Saguinus mystax 9488 \ 169 tamarin primates platyrrhini Saguinus imperator 9491 \ 170 Guianan Squirrel Monkey primates platyrrhini Saimiri sciureus 9521 \ 171 Red-chested Mustached Tamarin primates platyrrhini Saguinus labiatus 78454 \ 172 Goeldi's Monkey primates platyrrhini Callimico goeldii 9495 \ 173 Black-crowned Central American Squirrel Monkey primates platyrrhini Saimiri oerstedii 70928 \ 174 Golden-headed Lion Tamarin primates platyrrhini Leontopithecus chrysomelas 57374 \ 175 golden lion tamarin primates platyrrhini Leontopithecus rosalia 30588 \ 176 Humboldt's Squirrel Monkey primates platyrrhini Saimiri cassiquiarensis 2946521 \ 177 bare-eared squirrel monkey primates platyrrhini Saimiri ustus 66265 \ 178 Ecuadorian squirrel monkey primates platyrrhini Saimiri macrodon 2946522 \ 179 white-tufted-ear marmoset primates platyrrhini Callithrix jacchus 9483 \ 180 Eastern Pygmy Marmoset primates platyrrhini Cebuella niveiventris 2826950 \ 181 Western Pygmy Marmoset primates platyrrhini Cebuella pygmaea 9493 \ 182 Black And White Tassel-ear Marmoset primates platyrrhini Mico humeralifer 52232 \ 183 Black-crowned Dwarf Marmoset primates platyrrhini Callibella humilis (Mico humilis) 666519 \ 184 Mico schneideri primates platyrrhini Mico schneideri n/a \ 185 Silvery Marmoset primates platyrrhini Mico argentatus 9482 \ 186 Midas tamarin primates platyrrhini Saguinus midas 30586 \ 187 Wieds Marmoset primates platyrrhini Callithrix kuhlii 867363 \ 188 Geoffroy's Tufted-ear Marmoset primates platyrrhini Callithrix geoffroyi 52231 \ 189 Horsfield's tarsier primates tarsiidae Cephalopachus bancanus 9477 \ 190 Philippine tarsier primates tarsiidae Carlito syrichta
tarSyr21868482 \ 191 Lariang Tarsier primates tarsiidae Tarsius lariang 630277 \ 192 Wallace's Tarsier primates tarsiidae Tarsius wallacei 981131 \ 193 aye-aye primates strepsirrhini Daubentonia madagascariensis 31869 \ 194 Crowned Sifaka primates strepsirrhini Propithecus coronatus (Propithecus deckenii coronatus) 475619 \ 195 Perrier's Sifaka primates strepsirrhini Propithecus perrieri 989338 \ 196 ruffed lemur primates strepsirrhini Varecia variegata 9455 \ 197 Diademed Sifaka primates strepsirrhini Propithecus diadema 83281 \ 198 Milne-Edwards Sifaka primates strepsirrhini Propithecus edwardsi 543559 \ 199 babakoto primates strepsirrhini Indri indri 34827 \ 200 Golden-crowned Sifaka primates strepsirrhini Propithecus tattersalli 30601 \ 201 Eastern Woolly Lemur primates strepsirrhini Avahi laniger 122246 \ 202 Verreauxs Sifaka primates strepsirrhini Propithecus verreauxi 34825 \ 203 Peyrieras Woolly Lemur primates strepsirrhini Avahi peyrierasi 1313323 \ 204 Red Ruffed Lemur primates strepsirrhini Varecia rubra 554167 \ 205 greater bamboo lemur primates strepsirrhini Prolemur simus 1328070 \ 206 Red-bellied Lemur primates strepsirrhini Eulemur rubriventer 34829 \ 207 mongoose lemur primates strepsirrhini Eulemur mongoz 34828 \ 208 Geoffroys Dwarf Lemur primates strepsirrhini Cheirogaleus major 47177 \ 209 Crowned Lemur primates strepsirrhini Eulemur coronatus 13514 \ 210 black lemur primates strepsirrhini Eulemur macaco 30602 \ 211 lesser dwarf lemur primates strepsirrhini Cheirogaleus medius 9460 \ 212 Sclater's lemur primates strepsirrhini Eulemur flavifrons 87288 \ 213 Coquerel's sifaka primates strepsirrhini Propithecus coquerelli (Propithecus coquereli)
proCoq1379532 \ 214 Collared Brown Lemur primates strepsirrhini Eulemur collaris (Eulemur fulvus collaris) 47178 \ 215 Red-tailed Sportive Lemur primates strepsirrhini Lepilemur ruficaudatus 78866 \ 216 Red Brown Lemur primates strepsirrhini Eulemur rufus 859983 \ 217 Sanfords Brown Lemur primates strepsirrhini Eulemur sanfordi 122225 \ 218 White-fronted Lemur primates strepsirrhini Eulemur albifrons 1215604 \ 219 Gray's Sportive Lemur primates strepsirrhini Lepilemur dorsalis 78583 \ 220 brown lemur primates strepsirrhini Eulemur fulvus 13515 \ 221 Sahafary Sportive Lemur primates strepsirrhini Lepilemur septentrionalis 78584 \ 222 Sambirano Lesser Bamboo Lemur primates strepsirrhini Hapalemur occidentalis 867377 \ 223 Alaotra Reed Lemur primates strepsirrhini Hapalemur alaotrensis (Hapalemur griseus alaotrensis) 122220 \ 224 Eastern Lesser Bamboo Lemur primates strepsirrhini Hapalemur griseus 13557 \ 225 Ankarana Sportive Lemur primates strepsirrhini Lepilemur ankaranensis 342401 \ 226 ring-tailed lemur primates strepsirrhini Lemur catta 9447 \ 227 gray bamboo lemur primates strepsirrhini Hapalemur gilberti 3043110 \ 228 Rusty-gray Lesser Bamboo Lemur primates strepsirrhini Hapalemur meridionalis 3043112 \ 229 Demidoffs Dwarf Galago primates strepsirrhini Galagoides demidoff 89672 \ 230 northern giant mouse lemur primates strepsirrhini Mirza zaza 339999 \ 231 gray mouse lemur primates strepsirrhini Microcebus murinus
GCA_000165445.3_Mmur_3.030608 \ 232 small-eared galago primates strepsirrhini Otolemur garnettii
otoGar330611 \ 233 Northern Lesser Galago primates strepsirrhini Galago senegalensis 9465 \ 234 Thick-tailed Greater Galago primates strepsirrhini Otolemur crassicaudatus 9463 \ 235 Grey Slender Loris primates strepsirrhini Loris lydekkerianus 300163 \ 236 slender loris primates strepsirrhini Loris tardigradus 9468 \ 237 West African Potto primates strepsirrhini Perodicticus potto 9472 \ 238 East African Potto primates strepsirrhini Perodicticus ibeanus (Perodicticus potto ibeanus) 261737 \ 239 Moholi bushbaby primates strepsirrhini Galago moholi 30609 \ 240 Pygmy Slow Loris primates strepsirrhini Nycticebus pygmaeus (Xanthonycticebus pygmaeus) 101278 \ 241 Bengal slow loris primates strepsirrhini Nycticebus bengalensis 261741 \ 242 Calabar Angwantibo primates strepsirrhini Arctocebus calabarensis 261739 \ 243 slow loris primates strepsirrhini Nycticebus coucang 9470 \ 244 jaguar carnivora Panthera onca
GCA_004023805.1_PanOnc_v1_BIUU9690 \ 245 leopard carnivora Panthera pardus
GCA_001857705.1_PanPar1.09691 \ 246 giant panda carnivora Ailuropoda melanoleuca
GCA_002007445.1_ASM200744v19646 \ 247 Hawaiian monk seal carnivora Neomonachus schauinslandi
GCA_002201575.1_ASM220157v129088 \ 248 California sea lion carnivora Zalophus californianus
GCA_004024565.1_ZalCal_v1_BIUU9704 \ 249 Greenland wolf carnivora Canis lupus orion
GCA_905319855.2_mCanLor1.22605939 \ 250 Pacific walrus carnivora Odobenus rosmarus
odoRosDiv19707 \ 251 domestic cat (Fca126) carnivora Felis catus fca126 (Felis catus)
GCF_018350175.1_F.catus_Fca126_mat1.09685 \ 252 northern elephant seal carnivora Mirounga angustirostris
GCA_004023865.1_MirAng_v1_BIUU9716 \ 253 domestic cat carnivora Felis catus
felCat89685 \ 254 domestic dog (BS72/Village Dog) carnivora Canis lupus familiaris
GCA_004027395.1_CanFam_VD_v1_BIUU\ 255 German Shepherd dog (Mischka) carnivora Canis lupus familiaris (CanFam4) (Canis lupus familiaris)
canFam4\ 256 dingo carnivora Canis lupus dingo 286419 \ 257 raccoon dog carnivora Nyctereutes procyonoides 34880 \ 258 fossa carnivora Cryptoprocta ferox 94188 \ 259 polar bear carnivora Ursus maritimus
GCA_000687225.1_UrsMar_1.029073 \ 260 Asian palm civet carnivora Paradoxurus hermaphroditus
GCA_004024585.1_ParHer_v1_BIUU71117 \ 261 African hunting dog carnivora Lycaon pictus
GCA_001887905.1_LycPicSAfr1.09622 \ 262 Arctic fox carnivora Vulpes lagopus
GCA_004023825.1_VulLag_v1_BIUU494514 \ 263 dog carnivora Canis lupus familiaris
GCF_000002285.3_CanFam3.19615 \ 264 striped hyena carnivora Hyaena hyaena
GCA_004023945.1_HyaHya_v1_BIUU95912 \ 265 n/a carnivora Acinonyx jubatus
GCA_001443585.1_aciJub132536 \ 266 tiger carnivora Panthera tigris
GCA_000464555.1_PanTig1.09694 \ 267 Sea otter carnivora Enhydra lutris
GCA_002288905.2_ASM228890v234882 \ 268 giant otter carnivora Pteronura brasiliensis 9672 \ 269 bat-eared fox carnivora Otocyon megalotis 9624 \ 270 Weddell seal carnivora Leptonychotes weddellii
GCA_000349705.1_LepWed1.09713 \ 271 Lesser panda carnivora Ailurus fulgens
GCA_002007465.1_ASM200746v19649 \ 272 ratel carnivora Mellivora capensis
GCA_004024625.1_MelCap_v1_BIUU9664 \ 273 banded mongoose carnivora Mungos mungo
GCA_004023785.1_MunMun_v1_BIUU210652 \ 274 dwarf mongoose carnivora Helogale parvula
GCA_004023845.1_HelPar_v1_BIUU210647 \ 275 meerkat carnivora Suricata suricatta
GCA_004023905.1_SurSur_v1_BIUU37032 \ 276 puma carnivora Puma concolor
GCA_003327715.1_PumCon1.09696 \ 277 black-footed cat carnivora Felis nigripes
GCA_004023925.1_FelNig_v1_BIUU61379 \ 278 European polecat carnivora Mustela putorius
GCA_000239315.1_MusPutFurMale1.09668 \ 279 western spotted skunk carnivora Spilogale gracilis
GCA_004023965.1_SpiGra_v1_BIUU30551 \ 280 Sumatran rhinoceros laurasiatheria Dicerorhinus sumatrensis
GCA_002844835.1_ASM284483v189632 \ 281 black rhinoceros laurasiatheria Diceros bicornis
GCA_004027315.1_DicBicMic_v1_BIUU9805 \ 282 Asiatic tapir laurasiatheria Tapirus indicus
GCA_004024905.1_TapInd_v1_BIUU9802 \ 283 Brazilian tapir laurasiatheria Tapirus terrestris
GCA_004025025.1_TapTer_v1_BIUU9801 \ 284 northern white rhinoceros laurasiatheria Ceratotherium simum cottoni 310713 \ 285 ass laurasiatheria Equus asinus
GCA_001305755.1_ASM130575v19793 \ 286 Southern white rhinoceros laurasiatheria Ceratotherium simum
GCA_000283155.1_CerSimSim1.09807 \ 287 Przewalski's horse laurasiatheria Equus przewalskii
GCA_000696695.1_Burgud9798 \ 288 horse laurasiatheria Equus caballus
GCA_000002305.1_EquCab2.09796 \ 289 Malayan pangolin laurasiatheria Manis javanica
GCA_001685135.1_ManJav1.09974 \ 290 Chinese pangolin laurasiatheria Manis pentadactyla
GCA_000738955.1_M_pentadactyla-1.1.1143292 \ 291 Hispaniolan solenodon laurasiatheria Solenodon paradoxus 79805 \ 292 eastern mole laurasiatheria Scalopus aquaticus
GCA_004024925.1_ScaAqu_v1_BIUU71119 \ 293 gracile shrew mole laurasiatheria Uropsilus gracilis
GCA_004024945.1_UroGra_v1_BIUU182669 \ 294 star-nosed mole laurasiatheria Condylura cristata
GCF_000260355.1_ConCri1.0143302 \ 295 western European hedgehog laurasiatheria Erinaceus europaeus
GCA_000296755.1_EriEur2.09365 \ 296 European shrew laurasiatheria Sorex araneus
sorAra242254 \ 297 Indochinese shrew laurasiatheria Crocidura indochinensis
GCA_004027635.1_CroInd_v1_BIUU876679 \ 298 Hoffmann's two-fingered sloth xenarthra Choloepus hoffmanni
GCA_000164785.2_C_hoffmanni-2.0.19358 \ 299 nine-banded armadillo xenarthra Dasypus novemcinctus
GCA_000208655.2_Dasnov3.09361 \ 300 giant anteater xenarthra Myrmecophaga tridactyla
GCA_004026745.1_MyrTri_v1_BIUU71006 \ 301 southern tamandua xenarthra Tamandua tetradactyla
GCA_004025105.1_TamTet_v1_BIUU48850 \ 302 placentals xenarthra Tolypeutes matacus 183749 \ 303 southern two-toed sloth xenarthra Choloepus didactylus
GCA_004027855.1_ChoDid_v1_BIUU27675 \ 304 screaming hairy armadillo xenarthra Chaetophractus vellerosus
GCA_004027955.1_ChaVel_v1_BIUU340076 \ 305 North Pacific right whale artiodactyla Eubalaena japonica 302098 \ 306 grey whale artiodactyla Eschrichtius robustus 9764 \ 307 hippopotamus artiodactyla Hippopotamus amphibius
GCA_004027065.1_HipAmp_v1_BIUU9833 \ 308 Minke whale artiodactyla Balaenoptera acutorostrata
GCA_000493695.1_BalAcu1.09767 \ 309 beluga whale artiodactyla Delphinapterus leucas
GCA_002288925.2_ASM228892v29749 \ 310 Antarctic minke whale artiodactyla Balaenoptera bonaerensis
GCA_000978805.1_ASM97880v133556 \ 311 boutu artiodactyla Inia geoffrensis 9725 \ 312 harbor porpoise artiodactyla Phocoena phocoena 9742 \ 313 narwhal artiodactyla Monodon monoceros
GCA_004026685.1_MonMon_M_v1_BIUU40151 \ 314 Yangtze River dolphin artiodactyla Lipotes vexillifer
GCA_000442215.1_Lipotes_vexillifer_v1118797 \ 315 killer whale artiodactyla Orcinus orca
orcOrc19733 \ 316 Ganges River dolphin artiodactyla Platanista gangetica 118798 \ 317 Yangtze finless porpoise artiodactyla Neophocaena asiaeorientalis
GCA_003031525.1_Neophocaena_asiaeorientalis_V1189058 \ 318 Sowerby's beaked whale artiodactyla Mesoplodon bidens 48745 \ 319 alpaca artiodactyla Vicugna pacos
GCA_000767525.1_Vi_pacos_V1.030538 \ 320 Cuvier's beaked whale" artiodactyla Ziphius cavirostris 9760 \ 321 Bactrian camel artiodactyla Camelus bactrianus
GCA_000767855.1_Ca_bactrianus_MBC_1.09837 \ 322 Arabian camel artiodactyla Camelus dromedarius
GCA_000767585.1_PRJNA234474_Ca_dromedarius_V1.09838 \ 323 wild Bactrian camel artiodactyla Camelus ferus
GCA_000311805.2_CB1419612 \ 324 pygmy sperm whale artiodactyla Kogia breviceps 27615 \ 325 Chacoan peccary artiodactyla Catagonus wagneri
GCA_004024745.1_CatWag_v1_BIUU51154 \ 326 reindeer artiodactyla Rangifer tarandus
GCA_004026565.1_RanTarSib_v1_BIUU9870 \ 327 Pere David's deer artiodactyla Elaphurus davidianus
GCA_002443075.1_Milu1.043332 \ 328 okapi artiodactyla Okapia johnstoni
GCA_001660835.1_ASM166083v186973 \ 329 Masai giraffe artiodactyla Giraffa tippelskirchi
GCA_001651235.1_ASM165123v1439328 \ 330 Siberian musk deer artiodactyla Moschus moschiferus
GCA_004024705.1_MosMos_v1_BIUU68415 \ 331 water buffalo artiodactyla Bubalus bubalis
GCA_000471725.1_UMD_CASPUR_WB_2.089462 \ 332 cow artiodactyla Bos taurus
GCA_000003205.6_Btau_5.0.19913 \ 333 pronghorn artiodactyla Antilocapra americana
GCA_004027515.1_AntAmePen_v1_BIUU9891 \ 334 white-tailed deer artiodactyla Odocoileus virginianus
GCA_002102435.1_Ovir.te_1.09874 \ 335 aoudad artiodactyla Ammotragus lervia
GCA_002201775.1_ALER1.09899 \ 336 bighorn sheep artiodactyla Ovis canadensis
GCA_004026945.1_OviCan_v1_BIUU37174 \ 337 goat artiodactyla Capra hircus
GCA_001704415.1_ARS19925 \ 338 Nilgiri tahr artiodactyla Hemitragus hylocrius
GCA_004026825.1_HemHyl_v1_BIUU330464 \ 339 hirola artiodactyla Beatragus hunteri
GCA_004027495.1_BeaHun_v1_BIUU59527 \ 340 wild yak artiodactyla Bos mutus
bosMut172004 \ 341 American bison artiodactyla Bison bison
GCA_000754665.1_Bison_UMD1.09901 \ 342 sheep artiodactyla Ovis aries
GCA_000298735.2_Oar_v4.09940 \ 343 chiru artiodactyla Pantholops hodgsonii
GCA_000400835.1_PHO1.059538 \ 344 wild goat artiodactyla Capra aegagrus
GCA_000978405.1_CapAeg_1.09923 \ 345 Java mouse-deer artiodactyla Tragulus javanicus
GCA_004024965.1_TraJav_v1_BIUU9849 \ 346 pig artiodactyla Sus scrofa
susScr39823 \ 347 zebu cattle artiodactyla Bos indicus
GCA_000247795.2_Bos_indicus_1.09915 \ 348 common bottlenose dolphin artiodactyla Tursiops truncatus
GCA_001922835.1_NIST_Tur_tru_v19739 \ 349 Saiga antelope artiodactyla Saiga tatarica
GCA_004024985.1_SaiTat_v1_BIUU34875 \ 350 Chinese rufous horseshoe bat chiroptera Rhinolophus sinicus
GCA_001888835.1_ASM188883v189399 \ 351 black flying fox chiroptera Pteropus alecto
pteAle19402 \ 352 Cantor's roundleaf bat chiroptera Hipposideros galeritus 58069 \ 353 Egyptian rousette chiroptera Rousettus aegyptiacus
GCA_004024865.1_RouAeg_v1_BIUU9407 \ 354 long-tongued fruit bat chiroptera Macroglossus sobrinus 326083 \ 355 large flying fox chiroptera Pteropus vampyrus
GCF_000151845.1_Pvam_2.0132908 \ 356 Brazilian free-tailed bat chiroptera Tadarida brasiliensis
GCA_004025005.1_TadBra_v1_BIUU9438 \ 357 great roundleaf bat chiroptera Hipposideros armiger
GCA_001890085.1_ASM189008v1186990 \ 358 straw-colored fruit bat chiroptera Eidolon helvum
eidHel177214 \ 359 Antillean ghost-faced bat chiroptera Mormoops blainvillei
GCA_004026545.1_MorMeg_v1_BIUU118852 \ 360 tailed tailless bat chiroptera Anoura caudifer
GCA_004027475.1_AnoCau_v1_BIUU27642 \ 361 common vampire bat chiroptera Desmodus rotundus
GCA_002940915.2_ASM294091v29430 \ 362 hairy big-eared bat chiroptera Micronycteris hirsuta
GCA_004026765.1_MicHir_v1_BIUU148065 \ 363 stripe-headed round-eared bat chiroptera Tonatia saurophila
GCA_004024845.1_TonSau_v1_BIUU171122 \ 364 Seba's short-tailed bat chiroptera Carollia perspicillata
GCA_004027735.1_CarPer_v1_BIUU40233 \ 365 Jamaican fruit-eating bat chiroptera Artibeus jamaicensis
GCA_004027435.1_ArtJam_v1_BIUU9417 \ 366 Indian false vampire chiroptera Megaderma lyra
GCA_004026885.1_MegLyr_v1_BIUU9413 \ 367 Schreibers' long-fingered bat chiroptera Miniopterus schreibersii
GCA_004026525.1_MinSch_v1_BIUU9433 \ 368 greater bulldog bat chiroptera Noctilio leporinus
GCA_004026585.1_NocLep_v1_BIUU94963 \ 369 Natal long-fingered bat chiroptera Miniopterus natalensis
GCF_001595765.1_Mnat.v1291302 \ 370 hog-nosed bat chiroptera Craseonycteris thonglongyai
GCA_004027555.1_CraTho_v1_BIUU208972 \ 371 Parnell's mustached bat chiroptera Pteronotus parnellii
ptePar159476 \ 372 greater mouse-eared bat chiroptera Myotis myotis
GCA_004026985.1_MyoMyo_v1_BIUU51298 \ 373 Ashy-gray tube-nosed bat chiroptera Murina feae (Murina aurata feae)
GCA_004026665.1_MurFea_v1_BIUU1453894 \ 374 David's myotis chiroptera Myotis davidii
myoDav1225400 \ 375 Brandt's bat chiroptera Myotis brandtii
myoBra1109478 \ 376 big brown bat chiroptera Eptesicus fuscus
GCF_000308155.1_EptFus1.029078 \ 377 red bat chiroptera Lasiurus borealis
GCA_004026805.1_LasBor_v1_BIUU258930 \ 378 little brown bat chiroptera Myotis lucifugus
myoLuc259463 \ 379 common pipistrelle chiroptera Pipistrellus pipistrellus
GCA_004026625.1_PipPip_v1_BIUU59474 \ 380 African savanna elephant afrotheria Loxodonta africana
GCA_000001905.1_Loxafr3.09785 \ 381 Florida manatee afrotheria Trichechus manatus
GCA_000243295.1_TriManLat1.09778 \ 382 yellow-spotted hyrax afrotheria Heterohyrax brucei
GCA_004026845.1_HetBruBak_v1_BIUU77598 \ 383 Cape rock hyrax afrotheria Procavia capensis
GCA_004026925.1_ProCapCap_v1_BIUU9813 \ 384 aardvark afrotheria Orycteropus afer 9818 \ 385 Cape golden mole afrotheria Chrysochloris asiatica
GCA_004027935.1_ChrAsi_v1_BIUU185453 \ 386 Cape elephant shrew afrotheria Elephantulus edwardii
eleEdw128737 \ 387 Talazac's shrew tenrec afrotheria Microgale talazaci (Nesogale talazaci)
GCA_004026705.1_MicTal_v1_BIUU2583312 \ 388 small Madagascar hedgehog afrotheria Echinops telfairi
GCA_000313985.1_EchTel2.09371 \ 389 Sunda flying lemur euarchontoglires Galeopterus variegatus
GCA_004027255.1_GalVar_v1_BIUU482537 \ 390 Chinese tree shrew euarchontoglires Tupaia chinensis
tupChi1246437 \ 391 South African ground squirrel euarchontoglires Xerus inauris
GCA_004024805.1_XerIna_v1_BIUU234690 \ 392 large tree shrew euarchontoglires Tupaia tana 70687 \ 393 mountain beaver euarchontoglires Aplodontia rufa
GCA_004027875.1_AplRuf_v1_BIUU51342 \ 394 Alpine marmot euarchontoglires Marmota marmota
GCF_001458135.1_marMar2.19993 \ 395 Daurian ground squirrel euarchontoglires Spermophilus dauricus
GCA_002406435.1_ASM240643v199837 \ 396 crested porcupine euarchontoglires Hystrix cristata
GCA_004026905.1_HysCri_v1_BIUU10137 \ 397 thirteen-lined ground squirrel euarchontoglires Ictidomys tridecemlineatus
speTri243179 \ 398 American beaver euarchontoglires Castor canadensis
GCA_004027675.1_CasCan_v1_BIUU51338 \ 399 long-tailed chinchilla euarchontoglires Chinchilla lanigera
chiLan134839 \ 400 punctate agouti euarchontoglires Dasyprocta punctata 34846 \ 401 pacarana euarchontoglires Dinomys branickii
GCA_004027595.1_DinBra_v1_BIUU108858 \ 402 fat dormouse euarchontoglires Glis glis
GCA_004027185.1_GliGli_v1_BIUU41261 \ 403 northern gundi euarchontoglires Ctenodactylus gundi
GCA_004027205.1_CteGun_v1_BIUU10166 \ 404 naked mole-rat euarchontoglires Heterocephalus glaber
GCA_000247695.1_HetGla_female_1.010181 \ 405 Patagonian cavy euarchontoglires Dolichotis patagonum
GCA_004027295.1_DolPat_v1_BIUU29091 \ 406 capybara euarchontoglires Hydrochoerus hydrochaeris
GCA_004027455.1_HydHyd_v1_BIUU10149 \ 407 Montane guinea pig euarchontoglires Cavia tschudii
GCA_004027695.1_CavTsc_v1_BIUU143287 \ 408 domestic guinea pig euarchontoglires Cavia porcellus
GCA_000151735.1_Cavpor3.010141 \ 409 degu euarchontoglires Octodon degus
GCA_000260255.1_OctDeg1.010160 \ 410 lowland paca euarchontoglires Cuniculus paca 108852 \ 411 social tuco-tuco euarchontoglires Ctenomys sociabilis
GCA_004027165.1_CteSoc_v1_BIUU43321 \ 412 Damara mole-rat euarchontoglires Fukomys damarensis
fukDam1885580 \ 413 woodland dormouse euarchontoglires Graphiurus murinus 51346 \ 414 Desmarest's hutia euarchontoglires Capromys pilorides
GCA_004027915.1_CapPil_v1_BIUU34842 \ 415 Upper Galilee mountains blind mole rat euarchontoglires Nannospalax galili
GCA_000622305.1_S.galili_v1.01026970 \ 416 nutria euarchontoglires Myocastor coypus
GCA_004027025.1_MyoCoy_v1_BIUU10157 \ 417 hazel dormouse euarchontoglires Muscardinus avellanarius
GCA_004027005.1_MusAve_v1_BIUU39082 \ 418 dassie-rat euarchontoglires Petromus typicus
GCA_004026965.1_PetTyp_v1_BIUU10183 \ 419 greater cane rat euarchontoglires Thryonomys swinderianus
GCA_004025085.1_ThrSwi_v1_BIUU10169 \ 420 snowshoe hare euarchontoglires Lepus americanus
GCA_004026855.1_LepAme_v1_BIUU48086 \ 421 Gambian giant pouched rat euarchontoglires Cricetomys gambianus
GCA_004027575.1_CriGam_v1_BIUU10085 \ 422 Prairie deer mouse euarchontoglires Peromyscus maniculatus
GCF_000500345.1_Pman_1.010042 \ 423 southern grasshopper mouse euarchontoglires Onychomys torridus
GCA_004026725.1_OnyTor_v1_BIUU38674 \ 424 rabbit euarchontoglires Oryctolagus cuniculus
GCA_000003625.1_OryCun2.09986 \ 425 muskrat euarchontoglires Ondatra zibethicus
GCA_004026605.1_OndZib_v1_BIUU10060 \ 426 northern mole vole euarchontoglires Ellobius talpinus
GCA_001685095.1_ETalpinus_0.1329620 \ 427 Mongolian gerbil euarchontoglires Meriones unguiculatus
GCA_004026785.1_MerUng_v1_BIUU10047 \ 428 fat sand rat euarchontoglires Psammomys obesus
GCA_002215935.1_ASM221593v148139 \ 429 house mouse euarchontoglires Mus musculus
mm1010090 \ 430 Chinese hamster euarchontoglires Cricetulus griseus
GCA_900186095.1_CHOK1S_HZDv110029 \ 431 Norway rat euarchontoglires Rattus norvegicus
GCF_000001895.5_Rnor_6.010116 \ 432 western wild mouse euarchontoglires Mus spretus
GCA_001624865.1_SPRET_EiJ_v110096 \ 433 meadow jumping mouse euarchontoglires Zapus hudsonius
GCA_004024765.1_ZapHud_v1_BIUU160400 \ 434 prairie vole euarchontoglires Microtus ochrogaster
micOch179684 \ 435 Ryukyu mouse euarchontoglires Mus caroli
GCA_900094665.2_CAROLI_EIJ_v1.110089 \ 436 Egyptian spiny mouse euarchontoglires Acomys cahirinus
GCA_004027535.1_AcoCah_v1_BIUU10068 \ 437 Gobi jerboa euarchontoglires Allactaga bullata (Orientallactaga bullata)
GCA_004027895.1_AllBul_v1_BIUU1041416 \ 438 shrew mouse euarchontoglires Mus pahari
GCF_900095145.1_PAHARI_EIJ_v1.110093 \ 439 Transcaucasian mole vole euarchontoglires Ellobius lutescens
GCA_001685075.1_ASM168507v139086 \ 440 hispid cotton rat euarchontoglires Sigmodon hispidus
GCA_004025045.1_SigHis_v1_BIUU42415 \ 441 lesser Egyptian jerboa euarchontoglires Jaculus jaculus
GCA_000280705.1_JacJac1.051337 \ 442 Brazilian guinea pig euarchontoglires Cavia aperea
cavApe137548 \ 443 golden hamster euarchontoglires Mesocricetus auratus
GCA_000349665.1_MesAur1.010036 \ 444 Stephens's kangaroo rat euarchontoglires Dipodomys stephensi
GCA_004024685.1_DipSte_v1_BIUU323379 \ 445 American pika euarchontoglires Ochotona princeps
GCA_000292845.1_OchPri3.09978 \ 446 Ord's kangaroo rat euarchontoglires Dipodomys ordii
dipOrd210020 \ 447 little pocket mouse euarchontoglires Perognathus longimembris 38669
\ Table 1. Genome assemblies included in the 447-way Conservation track.\
\ Pollard KS, Hubisz MJ, Rosenbloom KR, Siepel A.\ \ Detection of nonneutral substitution rates on mammalian phylogenies.\ Genome Res. 2010 Jan;20(1):110-21.\ PMID: 19858363;\ PMC: PMC2798823\
\\ Kuderna LFK, Ulirsch JC, Rashid S, Ameen M, Sundaram L, Hickey G, Cox AJ, Gao H, Kumar A, Aguet F\ et al.\ \ Identification of constrained sequence elements across 239 primate genomes.\ Nature. 2023 Nov 29;.\ DOI: 10.1038/s41586-023-06798-8; PMID: 38030727\
\\ Kuderna LFK, Gao H, Janiak MC, Kuhlwilm M, Orkin JD, Bataillon T, Manu S, Valenzuela A, Bergman J,\ Rousselle M et al.\ \ A global catalog of whole-genome diversity from 233 primate species.\ Science. 2023 Jun 2;380(6648):906-913.\ DOI: 10.1126/science.abn7829;\ PMID: 37262161\
\\ Zoonomia Consortium.\ \ A comparative genomics multitool for scientific discovery and conservation.\ Nature. 2020 Nov;587(7833):240-245.\ DOI: 10.1038/s41586-020-2876-6; PMID: 33177664; PMC: PMC7759459\
\\ Feng S, Stiller J, Deng Y, Armstrong J, Fang Q, Reeve AH, Xie D, Chen G, Guo C, Faircloth BC et\ al.\ \ Dense sampling of bird diversity increases power of comparative genomics.\ Nature. 2020 Nov;587(7833):252-257.\ DOI: 10.1038/s41586-020-2873-9; PMID: 33177665; PMC: PMC7759463\
\\ Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, Fang Q, Xie D, Feng S, Stiller J\ et al.\ \ Progressive Cactus is a multiple-genome aligner for the thousand-genome era.\ Nature. 2020 Nov;587(7833):246-251.\ DOI: 10.1038/s41586-020-2871-y; PMID: 33177663; PMC: PMC7673649\
\ compGeno 1 compositeTrack on\ dragAndDrop subTracks\ group compGeno\ html cactus447way\ longLabel Zoonomia+Primates 447 - 447 mammals, including 233 primates, aligned with Cactus, for Kuderna et al. 2023\ shortLabel Zoonomia+Primates 447\ subGroup1 view Views align=Multiz_Alignments phyloP=Basewise_Conservation_(phyloP)\ track cons447way\ type bed 4\ visibility hide\ AorticSmoothMuscleCellResponseToIL1b06hrBiolRep3LK60_CNhs13586_ctss_fwd AorticSmsToIL1b_06hrBr3+ bigWig Aortic smooth muscle cell response to IL1b, 06hr, biol_rep3 (LK60)_CNhs13586_12857-137D4_forward 0 101 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12857-137D4 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20IL1b%2c%2006hr%2c%20biol_rep3%20%28LK60%29.CNhs13586.12857-137D4.hg38.ctss.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to IL1b, 06hr, biol_rep3 (LK60)_CNhs13586_12857-137D4_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12857-137D4 sequence_tech=hCAGE\ parent TSS_activity_read_counts off\ shortLabel AorticSmsToIL1b_06hrBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_IL1b strand=forward\ track AorticSmoothMuscleCellResponseToIL1b06hrBiolRep3LK60_CNhs13586_ctss_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12857-137D4\ urlLabel FANTOM5 Details:\ AorticSmoothMuscleCellResponseToIL1b06hrBiolRep3LK60_CNhs13586_tpm_fwd AorticSmsToIL1b_06hrBr3+ bigWig Aortic smooth muscle cell response to IL1b, 06hr, biol_rep3 (LK60)_CNhs13586_12857-137D4_forward 1 101 255 0 0 255 127 127 0 0 0 http://fantom.gsc.riken.jp/5/sstar/FF:12857-137D4 regulation 0 bigDataUrl /gbdb/hg38/fantom5/Aortic%20smooth%20muscle%20cell%20response%20to%20IL1b%2c%2006hr%2c%20biol_rep3%20%28LK60%29.CNhs13586.12857-137D4.hg38.tpm.fwd.bw\ color 255,0,0\ longLabel Aortic smooth muscle cell response to IL1b, 06hr, biol_rep3 (LK60)_CNhs13586_12857-137D4_forward\ maxHeightPixels 100:8:8\ metadata ontology_id=12857-137D4 sequence_tech=hCAGE\ parent TSS_activity_TPM off\ shortLabel AorticSmsToIL1b_06hrBr3+\ subGroups sequenceTech=hCAGE category=AoSMC_response_to_IL1b strand=forward\ track AorticSmoothMuscleCellResponseToIL1b06hrBiolRep3LK60_CNhs13586_tpm_fwd\ type bigWig\ url http://fantom.gsc.riken.jp/5/sstar/FF:12857-137D4\ urlLabel FANTOM5 Details:\ ENCFF978IHV_ENCFF221TSA_ENCFF619JXN_ENCFF227NGR ENCFF978IHV_ENCFF221TSA_ENCFF619JXN_ENCFF227NGR bigBed 9 + 5 Caco-2: (1) cCREs 4 101 0 0 0 127 127 127 0 0 0 https://screen.wenglab.org/search?assembly=GRCh38&accessions=$$ regulation 1 bigDataUrl /gbdb/hg38/encode4/ccre/coreCollection/ENCFF978IHV_ENCFF221TSA_ENCFF619JXN_ENCFF227NGR.bb\ longLabel Caco-2: (1) cCREs\ mouseOver ID: ${name}