Abstract
Despite whole-genome sequencing (WGS), many cases of single-gene disorders remain unsolved, impeding diagnosis and preventative care for people whose disease-causing variants escape detection. Since early WGS data analytic steps prioritize protein-coding sequences, to simultaneously prioritize variants in non-coding regions rich in transcribed and critical regulatory sequences, we developed GROFFFY, an analytic tool that integrates coordinates for regions with experimental evidence of functionality. Applied to WGS data from solved and unsolved hereditary hemorrhagic telangiectasia (HHT) recruits to the 100,000 Genomes Project, GROFFFY-based filtration reduced the mean number of variants/DNA from 4,867,167 to 21,486, without deleting disease-causal variants. In three unsolved cases (two related), GROFFFY identified ultra-rare deletions within the 3' untranslated region (UTR) of the tumor suppressor SMAD4, where germline loss-of-function alleles cause combined HHT and colonic polyposis (MIM: 175050). Sited >5.4 kb distal to coding DNA, the deletions did not modify or generate microRNA binding sites, but instead disrupted the sequence context of the final cleavage and polyadenylation site necessary for protein production: By iFoldRNA, an AAUAAA-adjacent 16-nucleotide deletion brought the cleavage site into inaccessible neighboring secondary structures, while a 4-nucleotide deletion unfolded the downstream RNA polymerase II roadblock. SMAD4 RNA expression differed to control-derived RNA from resting and cycloheximide-stressed peripheral blood mononuclear cells. Patterns predicted the mutational site for an unrelated HHT/polyposis-affected individual, where a complex insertion was subsequently identified. In conclusion, we describe a functional rare variant type that impacts regulatory systems based on RNA polyadenylation. Extension of coding sequence-focused gene panels is required to capture these variants.
| Original language | English |
|---|---|
| Pages (from-to) | 1903-1918 |
| Number of pages | 16 |
| Journal | American Journal of Human Genetics |
| Volume | 110 |
| Issue number | 11 |
| Early online date | 9 Oct 2023 |
| DOIs | |
| Publication status | Published - 2 Nov 2023 |
Bibliographical note
Funding Information:This research was made possible through access to the data and findings generated by the 100,000 Genomes Project. The work was cofounded by the National Institute for Health Research Imperial Biomedical Research Centre, the D'Almeida Charitable Trust, and Imperial College Healthcare NHS Trust. A.A. was supported by Prince Sultan Military Medical City, Saudi Arabia. M.A.A. was supported by the National Institutes of Health (grant R35HL140019). The 100,000 Genomes Project is managed by Genomics England Limited (a wholly owned company of the Department of Health and Social Care). The 100,000 Genomes Project uses data provided by patients and collected by the National Health Service as part of their care and support. We thank the National Health Service staff of the UK Genomic Medicine Centres and the participants for their willing participation; the Genomics England Clinical Research Interface team, specifically Susan Walker, for separately reviewing bam file variant sequences; Charlotte Bevan, Michael Hubank, and Santiago Vernia for helpful discussions and manuscript review; and our academic and public partners within the NIHR Imperial BRC's Social Genetic and Environmental Determinants of Health (SGE) theme. We specifically thank the presented families for confirmation of their clinical phenotypes and consent to share in this manuscript. The views expressed are those of the authors and not necessarily those of funders, the NHS, the NIHR, or the Department of Health and Social Care. Conceptualization, S.X. C.L.S.; methodology, S.X. Z.K. D.M. D.P. A.M.B. M.E.B.-H. A.A. M.A.A. N.V. M.J.C. GERC, C.L.S.; investigation, S.X. Z.K. D.M. D.L. A.D.M. S.K.W. C.L.S.; visualization, S.X. C.L.S.; funding acquisition, C.L.S.; project administration, GERC, C.L.S.; supervision, D.P. M.E.B.-H. M.A.A. C.L.S.; writing – original draft, C.L.S.; writing – review & editing, S.X. Z.K. D.M. D.L. D.P. A.M.B. M.E.B.-H. A.A. A.D.M. S.K.W. M.A.A. N.V. M.J.C. GERC, C.L.S. S.X. devised and generated the GROFFFY approach, devised all scripts to generate GROFFFY, and generated all GROFFFY numeric data, Figures 1, 2, 3, S1, and S2, and Tables S1, S2, S3, S4, S5, and S6. Z.K. advised on Linux and script generation. D.M. interrogated donor 3 bam files. D.L. assisted in PBMC cultures. D.P. A.M.B. M.E.B.-H. and M.A.A. performed BOEC cultures and RNA preparations. A.A. designed primers for validations. A.D.M. contributed to recruitment of affected individuals. S.K.W. contributed to clinical correlations. N.V. advised on SMAD4 regulation. M.J.C. contributed to specific project set up at Genomics England. GERC performed all whole-genome sequencing and alignments. C.L.S. recruited patients and performed clinical correlations; devised concepts and advised on GROFFFY approaches; devised and performed PBMC cultures; devised and performed in-house endothelial and PBMC RNA-seq and variant level data analyses; generated Figure 3, 4, 5, 6, 7, 8, S3, S4, S5, S6, and S7, and Tables S6–S8, and wrote the manuscript. All authors have reviewed and approved the final manuscript. The authors declare no competing interests.
Funding Information:
This research was made possible through access to the data and findings generated by the 100,000 Genomes Project. The work was cofounded by the National Institute for Health Research Imperial Biomedical Research Centre, the D’Almeida Charitable Trust, and Imperial College Healthcare NHS Trust. A.A. was supported by Prince Sultan Military Medical City , Saudi Arabia. M.A.A. was supported by the National Institutes of Health (grant R35HL140019 ). The 100,000 Genomes Project is managed by Genomics England Limited (a wholly owned company of the Department of Health and Social Care). The 100,000 Genomes Project uses data provided by patients and collected by the National Health Service as part of their care and support. We thank the National Health Service staff of the UK Genomic Medicine Centres and the participants for their willing participation; the Genomics England Clinical Research Interface team, specifically Susan Walker, for separately reviewing bam file variant sequences; Charlotte Bevan, Michael Hubank, and Santiago Vernia for helpful discussions and manuscript review; and our academic and public partners within the NIHR Imperial BRC’s Social Genetic and Environmental Determinants of Health (SGE) theme. We specifically thank the presented families for confirmation of their clinical phenotypes and consent to share in this manuscript. The views expressed are those of the authors and not necessarily those of funders, the NHS, the NIHR, or the Department of Health and Social Care.
Publisher Copyright:
© 2023 The Authors
Fingerprint
Dive into the research topics of 'Functional filter for whole-genome sequencing data identifies HHT and stress-associated non-coding SMAD4 polyadenylation site variants >5 kb from coding DNA'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver