Showing posts with label DNA Sequences. Show all posts
Showing posts with label DNA Sequences. Show all posts

Aug 6, 2024

Cracking the code of life: new AI model learns DNA's hidden language

DNA contains foundational information needed to sustain life. Understanding how this information is stored and organized has been one of the greatest scientific challenges of the last century. With GROVER, a new large language model trained on human DNA, researchers could now attempt to decode the complex information hidden in our genome. Developed by a team at the Biotechnology Center (BIOTEC) of Dresden University of Technology, GROVER treats human DNA as a text, learning its rules and context to draw functional information about the DNA sequences. This new tool, published in Nature Machine Intelligence, has the potential to transform genomics and accelerate personalized medicine.

Since the discovery of the double helix, scientists have sought to understand the information encoded in DNA. 70 years later, it is clear that the information hidden in the DNA is multilayered. Only 1-2 % of the genome consists of genes, the sequences that code for proteins.

"DNA has many functions beyond coding for proteins. Some sequences regulate genes, others serve structural purposes, most sequences serve multiple functions at once. Currently, we don't understand the meaning of most of the DNA. When it comes to understanding the non-coding regions of the DNA, it seems that we have only started to scratch the surface. This is where AI and large language models can help," says Dr. Anna Poetsch, research group leader at the BIOTEC.

DNA as a Language


Large language models, like GPT, have transformed our understanding of language. Trained exclusively on text, the large language models developed the ability to use the language in many contexts.

"DNA is the code of life. Why not treat it like a language?" says Dr. Poetsch. The Poetsch team trained a large language model on a reference human genome. The resulting tool named GROVER, or "Genome Rules Obtained via Extracted Representations," can be used to extract biological meaning from the DNA.

"GROVER learned the rules of DNA. In terms of language, we are talking about grammar, syntax, and semantics. For DNA this means learning the rules governing the sequences, the order of the nucleotides and sequences, and the meaning of the sequences. Like GPT models learning human languages, GROVER has basically learned how to 'speak' DNA," explains Dr. Melissa Sanabria, the researcher behind the project.

The team showed that GROVER can not only accurately predict the following DNA sequences but can also be used to extract contextual information that has biological meaning, e.g., identify gene promoters or protein binding sites on DNA. GROVER also learns processes that are generally considered to be "epigenetic," i.e., regulatory processes that happen on top of the DNA rather than being encoded.

"It is fascinating that by training GROVER with only the DNA sequence, without any annotations of functions, we are actually able to extract information on biological function. To us, it shows that the function, including some of the epigenetic information, is also encoded in the sequence," says Dr. Sanabria.

The DNA Dictionary

"DNA resembles language. It has four letters that build sequences and the sequences carry a meaning. However, unlike a language, DNA has no defined words," says Dr. Poetsch. DNA consists of four letters (A, T, G, and C) and genes, but there are no predefined sequences of different lengths that combine to build genes or other meaningful sequences.

To train GROVER, the team had to first create a DNA dictionary. They used a trick from compression algorithms. "This step is crucial and sets our DNA language model apart from the previous attempts," says Dr. Poetsch.

"We analyzed the whole genome and looked for combinations of letters that occur most often. We started with two letters and went over the DNA, again and again, to build it up to the most common multi-letter combinations. In this way, in about 600 cycles, we have fragmented the DNA into 'words' that let GROVER perform the best when it comes to predicting the next sequence," explains Dr. Sanabria.

Read more at Science Daily

Aug 24, 2023

Researchers fully sequence the Y chromosome for the first time

What was once the final frontier of the human genome -- the Y chromosome -- has just been mapped out in its entirety.

Led by the National Human Genome Research Institute (NHGRI), a team of researchers at the National Institute of Standards and Technology (NIST) and many other organizations used advanced sequencing technologies to read out the full DNA sequence of the Y chromosome -- a region of the genome that typically drives male reproductive development. The results of a study published in Nature demonstrate that this advance improves DNA sequencing accuracy for the chromosome, which could help identify certain genetic disorders and potentially uncover the genetic roots of others.

DNA sequencing isn't as simple as reading genetic material from a genome's beginning to its end. DNA gets chopped up when it is extracted from cells, plus even the best sequencing equipment can only handle relatively small bits of DNA at a time. So, researchers and clinicians rely on special software to piece together fragments of sequenced code in the correct order like a puzzle.

A reference genome is a separate, already pieced-together genome that serves as a guide, similar to the pictures on the front of puzzle boxes. And because 99.9% of our species' genetic code is shared, any human genome would closely match a reference.

Last year, a team from the Telomere-to-Telomere (T2T) consortium, which is made up of experts from dozens of organizations such as NIST, generated the most complete reference genome at the time by using new sequencing technologies to crack previously indecipherable regions of the genome. But cells used in that work did not contain the most puzzling of all, the Y chromosome.

"Chromosomes all contain sections of very repetitive DNA, but well over half of the Y chromosome is like that," said study co-author Justin Zook, who leads NIST's Genome in a Bottle (GIAB) consortium. "If you use the puzzle analogy, a lot of the Y chromosome looks like the backgrounds often do, where all the pieces look really similar."

With this new endeavor, T2T was not starting at zero as the GIAB had already gotten the ball rolling.

The GIAB's mission is to produce test materials, or benchmarks, that can be used to evaluate sequencing technologies or methods. The materials themselves are highly accurate readouts of specific genes that can act as an answer key for checking the results of a particular sequencing method.

NIST has rigorously analyzed several individual human genomes to create their benchmarks. While GIAB has not yet produced a benchmark for the Y chromosome specifically, the consortium has studied one genome extensively, accumulating the largest collection of Y chromosome data prior to the new study.

That data served as a jumping-off point for the new study's authors, who focused their analysis on the best understood GIAB Y chromosome. They examined the sample with a combination of cutting-edge technologies -- namely high fidelity and nanopore sequencing -- that make the DNA fragment puzzle pieces larger and thus easier to assemble.

A machine-learning analysis tool and gamut of other advanced programs helped the team identify and assemble the pieces of the chromosome. More than 62 million letters of genetic code later, the authors had spelled out the GIAB Y chromosome front to back.

The researchers pitted their complete Y chromosome sequence, named T2T-Y, against the most widely used reference genome's Y chromosome parts, which are riddled with stretches of absent code. Using them both as guides for sequencing a diverse group of over 1,200 separate genomes, they found that T2T-Y drastically improved the outcomes.

T2T-Y, in combination with the group's previous reference genome, T2T-CHM13, represents the world's first complete genome for the half of the population with a Y chromosome.

The newest addition could be useful in identifying and diagnosing the few known conditions related to genes in the Y chromosome. But what's more is the new reference's potential to shed light on new genes and their function.

"There are certainly aspects of fertility and some genetic disorders that are connected to genes in the Y chromosome," Zook said. "But because it's been so hard to analyze up to this point, we may not even know yet just how important the Y chromosome is."

Read more at Science Daily

May 27, 2023

River erosion can shape fish evolution

New findings could explain biodiversity hotspots in tectonically quiet regions.

If we could rewind the tape of species evolution around the world and play it forward over hundreds of millions of years to the present day, we would see biodiversity clustering around regions of tectonic turmoil. Tectonically active regions such as the Himalayan and Andean mountains are especially rich in flora and fauna due to their shifting landscapes, which act to divide and diversify species over time.

But biodiversity can also flourish in some geologically quieter regions, where tectonics hasn't shaken up the land for millennia. The Appalachian Mountains are a prime example: The range has not seen much tectonic activity in hundreds of millions of years, and yet the region is a notable hotspot of freshwater biodiversity.

Now, an MIT study identifies a geological process that may shape the diversity of species in tectonically inactive regions. In a paper appearing in Science, the researchers report that river erosion can be a driver of biodiversity in these older, quieter environments.

They make their case in the southern Appalachians, and specifically the Tennessee River Basin, a region known for its huge diversity of freshwater fishes. The team found that as rivers eroded through different rock types in the region, the changing landscape pushed a species of fish known as the greenfin darter into different tributaries of the river network. Over time, these separated populations developed into their own distinct lineages.

The team speculates that erosion likely drove the greenfin darter to diversify. Although the separated populations appear outwardly similar, with the greenfin darter's characteristic green-tinged fins, they differ substantially in their genetic makeup. For now, the separated populations are classified as one single species.

"Give this process of erosion more time, and I think these separate lineages will become different species," says Maya Stokes PhD '21, who carried out part of the work as a graduate student in MIT's Department of Earth, Atmospheric and Planetary Sciences (EAPS).

The greenfin darter may not be the only species to diversify as a consequence of river erosion. The researchers suspect that erosion may have driven many other species to diversify throughout the basin, and possibly other tectonically inactive regions around the world.

"If we can understand the geologic factors that contribute to biodiversity, we can do a better job of conserving it," says Taylor Perron, the Cecil and Ida Green Professor of Earth, Atmospheric, and Planetary Sciences at MIT.

The study's co-authors include collaborators at Yale University, Colorado State University, the University of Tennessee, the University of Massachusetts at Amherst, and the Tennessee Valley Authority (TVA). Stokes is currently an assistant professor at Florida State University.

Fish in trees


The new study grew out of Stokes' PhD work at MIT, where she and Perron were exploring connections between geomorphology (the study of how landscapes evolve) and biology. They came across work at Yale by Thomas Near, who studies lineages of North American freshwater fishes. Near uses DNA sequence data collected from freshwater fishes across various regions of North America to show how and when certain species evolved and diverged in relation to each other.

Near brought a curious observation to the team: a habitat distribution map of the greenfin darter showing that the fish was found in the Tennessee River Basin -- but only in the southern half. What's more, Near had mitochondrial DNA sequence data showing that the fish's populations appeared to be different in their genetic makeup depending on the tributary in which they were found.

To investigate the reasons for this pattern, Stokes gathered greenfin darter tissue samples from Near's extensive collection at Yale, as well as from the field with help from TVA colleagues. She then analyzed DNA sequences from across the entire genome, and compared the genes of each individual fish to every other fish in the dataset. The team then created a phylogenetic tree of the greenfin darter, based on the genetic similarity between fish.

From this tree, they observed that fish within a tributary were more related to each other than to fish in other tributaries. What's more, fish within neighboring tributaries were more similar to each other than fish from more distant tributaries.

"Our question was, could there have been a geological mechanism that, over time, took this single species, and splintered it into different, genetically distinct groups?" Perron says.

A changing landscape

Stokes and Perron started to observe a "tight correlation" between greenfin darter habitats and the type of rock where they are found. In particular, much of the southern half of the Tennessee River Basin, where the species abounds, is made of metamorphic rock, whereas the northern half consists of sedimentary rock, where the fish are not found.

They also observed that the rivers running through metamorphic rock are steeper and more narrow, which generally creates more turbulence, a characteristic greenfin darters seem to prefer. The team wondered: Could the distribution of greenfin darter habitat have been shaped by a changing landscape of rock type, as rivers eroded into the land over time?

To check this idea, the researchers developed a model to simulate how a landscape evolves as rivers erode through various rock types. They fed the model information about the rock types in the Tennessee River Basin today, then ran the simulation back to see how the same region may have looked millions of years ago, when more metamorphic rock was exposed.

They then ran the model forward and observed how the exposure of metamorphic rock shrank over time. They took special note of where and when connections between tributaries crossed into non-metamorphic rock, blocking fish from passing between those tributaries. They drew up a simple timeline of these blocking events and compared this to the phylogenetic tree of diverging greenfin darters. The two were remarkably similar: The fish seemed to form separate lineages in the same order as when their respective tributaries became separated from the others.

"It means it's plausible that erosion through different rock layers caused isolation between different populations of the greenfin darter and caused lineages to diversify," Stokes says.

Read more at Science Daily

Jan 24, 2023

Genome editing procedures optimized

In the course of optimising key procedures of genome editing, researchers from the department of Developmental Biology / Physiology at the Centre for Organismal Studies of Heidelberg University have succeeded in substantially improving the efficiency of molecular genetic methods such as CRISPR/Cas9 and related systems, and in broadening their areas of application. Together with colleagues from other disciplines, the life scientists fine-tuned these tools to enable, inter alia, effective genetic screening for modelling specific gene mutations. In addition, initially inaccessible DNA sequences can now be modified. According to Prof. Dr Joachim Wittbrodt, this opens up extensive new areas of work in basic research and, potentially, therapeutic application.

Genome editing means the deliberate altering of DNA with molecular genetic methods. It is used to breed plants and animals, but also in basic medical and biological research. The most common procedures include the "gene scissors" CRISPR/Cas9 and its variants known as base editors. In both cases, enzymes have to be transported into the nucleus of the target cell. Upon arrival, the CRISPR/Cas9 system cuts the DNA at specific sites, which causes a double strand break. New DNA segments can then be inserted at that site. Base editors use a similar molecular mechanism but they do not cut the DNA double strand. Instead, an enzyme coupled with the Cas9 protein performs a targeted exchange of nucleotides -- the basic building blocks of the genome. In three successive studies, Prof. Wittbrodt's team succeeded in considerably enhancing the efficiency and applicability of these methods.

A challenge when using CRISPR/Cas9 consists in the efficient delivery of the required Cas9 enzymes to the nucleus. "The cell has an elaborate 'bouncer' mechanism. It distinguishes between proteins that are allowed to translocate into the nucleus and those that are supposed to stay in the cytoplasm," explains Dr Tinatini Tavhelidse-Suck from Prof. Wittbrodt's team. Access is enabled here by a tag made up of a few amino acids that functions like an "admission ticket." The scientists have now come up with a kind of generally valid "VIP admission ticket" which lets enzymes equipped with it into the nucleus very quickly. They have named it "high efficiency-tag," "hei-tag" for short. "Other proteins that have to penetrate the cell nucleus are also more successful with 'hei-tag'," concludes Dr Thomas Thumberger, who is also a researcher at the Centre for Organismal Studies (COS). In cooperation with pharmacologists from Heidelberg University, the team could show that Cas9 in connection with the "hei-tag" ticket can enable highly efficient, targeted genome alterations not only in the model organism medaka, the Japanese ricefish (Oryzias latipes), but also in mammalian cell cultures and mouse embryos.

In a further study, the Heidelberg scientists showed that base editors operate highly efficiently in the living organism and are even suited to genetic screening. In an experiment with Japanese rice fish, they were able to show that these locally limited, targeted modifications in individual buildings blocks of the DNA achieve an outcome that is otherwise only obtained by the comparatively laborious breeding of organisms with altered genes. The research team at COS, in cooperation with Dr Dr Jakob Gierten, a paediatric cardiologist at Heidelberg University Hospital, focused on certain genetic mutations. These mutations were suspected of triggering congenital heart defects in humans. Through modifying individual building blocks of the DNA of the relevant genes in the model organism, the scientists were able to imitate and study fish embryos with the described heart defects. The targeted intervention led to visible changes in the heart already during early stages of fish embryonic development, say Bettina Welz and Dr Alex Cornean, two of the first authors of the study from Prof. Wittbrodt's team. That enabled the researchers to confirm the original suspicion and establish a causal connection between genetic alteration and clinical symptoms.

The precise intervention in the genome of the fish embryos was made possible through especially developed software ACEofBASEs, which is available online. It allows for identifying genetic locations that very efficiently lead to desired changes in the target genes and the resultant proteins. The scientists say that the Japanese ricefish is an excellent genetic model organism for modelling mutations like those identified from the respective patients. "Our method enables an efficient screening analysis and could therefore offer a starting point for developing individualised medical treatment," according to Jakob Gierten.

A third study, again from the Wittbrodt group, deals with the limitations of base editors. For such an editor to bind the DNA of a target cell, there has to be a certain sequence motif. It is called Protospacer Adjacent Motif, PAM for short. "If this motif is lacking near the DNA building block to be changed, it is impossible to exchange nucleotides," explains Dr Thumberger. A team under his direction has now found a way to get around this limitation. Two base editors in a single cell are used in succession. In an initial step, a new DNA binding motif for a further base editor is generated, upon which this second editor, which is applied simultaneously, can edit a site that was inaccessible before. This staggered use turned out to be highly efficient, explains Kaisa Pakari, the first author of the study. With this trick, the Heidelberg scientists were able to increase the number of possible application sites of established base editors by 65 percent. Now DNA sequences that were initially inaccessible can also be modified.

"Optimising the existing tools for genome editing and their fine-tuned application results in enormously varied possibilities for basic research and, potentially, novel therapeutic approaches," Joachim Wittbrodt underlines.

Read more at Science Daily

Nov 29, 2022

DNA sequence enhances understanding origins of jaws

Researchers at Uppsala University have discovered and characterised a DNA sequence found in jawed vertebrates, such as sharks and humans, but absent in jawless vertebrates, such as lampreys. This DNA is important for the shaping of the joint surfaces during embryo development.

The vast majority of vertebrate species living today, including humans, belong to the jawed vertebrate group. The development of articulating jaws during vertebrate evolution was one of the most significant evolutionary transitions from jawless to jawed vertebrates, taking place at least 423 million years ago. The lower and upper jaws were initially connected by the primary jaw joint. However, during the evolution of mammals this moved to the middle ear to enhance hearing and was replaced by the secondary jaw joint, which is how humans are constructed today.

The primary jaw joint is formed during embryonic development and has an active gene which contains sequence information for a specific protein -- transcription factor Nkx3.2. This protein has long been thought to have played a major role in the evolution of this jaw joint, but little was known before about how its gene activity is regulated in the jaw joint cells.

Typically, genes are activated with help from DNA sequences, known as enhancers, that do not contain gene sequence information. Furthermore, such 'regulatory' DNA can contribute to the activation of the gene only in a certain cell type and can be conserved among different animal species.

"We searched through the genome sequences of many different vertebrate species and only found the DNA sequence near the Nkx3.2 gene in jawed vertebrates -- not in jawless ones. When we injected these DNA sequences from jawed vertebrates into zebrafish embryos, they were all activated in the jaw joint cells. The fact that their ability to activate has been preserved for over 400 million years shows how important it is for jawed vertebrates," notes Tatjana Haitina, researcher at Uppsala University, who led the study.

"In experiments where we deleted the newly discovered DNA sequence from the zebrafish genome using the CRISPR/Cas9 technique, we saw that the early activation of the Nkx3.2 gene was reduced, which caused defects in the jaw joint shape. It turned out that these defects were later repaired, suggesting that there is additional regulatory DNA somewhere in the genome that controls the activation of the Nkx3.2 gene and is waiting to be discovered," adds Jake Leyhr, doctoral student student in the research team.

The researchers hope that their discovery is an important step towards eventually understanding the process behind the origins of vertebrate jaws.

Read more at Science Daily

Sep 1, 2022

Corals pass mutations acquired during their lifetimes to offspring

In a discovery that challenges over a century of evolutionary conventional wisdom, corals have been shown to pass somatic mutations -- changes to the DNA sequence that occur in non-reproductive cells -- to their offspring. The finding, by an international team of scientists led by Penn State biologists, demonstrates a potential new route for the generation of genetic diversity, which is the raw material for evolutionary adaptation, and could be vital for allowing endangered corals to adapt to rapidly changing environmental conditions.

"For a trait, such as growth rate, to evolve, the genetic basis of that trait must be passed from generation to generation," said Iliana Baums, professor of biology at Penn State and leader of the research team. "For most animals, a new genetic mutation can only contribute to evolutionary change if it occurs in a germline or reproductive cell, for example in an egg or sperm cell. Mutations that occur in the rest of the body, in the somatic cells, were thought to be evolutionarily irrelevant because they do not get passed on to offspring. However, corals appear to have a way around this barrier that seems to allow them to break this evolutionary rule."

Since the time of Darwin, our understanding of evolution has become ever more detailed. We now know that an organism's traits are heavily determined by the sequence of their DNA. Individuals in a population vary in their DNA sequence, and this genetic variation can lead to the variation in traits, such as body size, that could give an individual a reproductive advantage. Only rarely does a new genetic mutation occur that gives an individual such a reproductive advantage and evolution can only proceed further if -- and this is the key -- the individual can pass the change to its offspring.

"In most animals, reproductive cells are segregated from body cells early in development," said Kate Vasquez Kuntz, a graduate student at Penn State and the co-lead author of the study. "So only genetic mutations that occur in the reproductive cells have the potential to contribute to the evolution of the species. This slow process of waiting for rare mutations in a particular set of cells can be particularly problematic given the rapid nature of climate change. However, for some organisms, like corals, the segregation of reproductive cells from all other cells may occur later in development or may never occur at all, allowing a path for genetic mutations to travel from a parent's body to its offspring. This would increase genetic variation and potentially even serve as a 'pre-screening' system for advantageous mutations."

Corals can reproduce both asexually (through budding and colony fragmentation) and sexually, by producing egg and sperm cells. For the Elkhorn corals studied here, which broadcast their egg and sperm cells into the water in spawning events, eggs from one coral colony are usually fertilized by sperm from a neighboring colony. However, the research team found that some Elkhorn coral eggs developed into viable offspring without a second coral being involved, a kind of single-parent sexual reproduction.

"This single-parent reproduction allowed us to more easily search for potential somatic mutations from the parent coral and track them into the offspring by simplifying the total number of genetic possibilities that could occur in the offspring," said Sheila Kitchen, co-lead author of the study, a postdoctoral researcher at Penn State and the California Institute of Technology co-lead author of the study.

The research team genotyped samples -- using a high-resolution molecular tool called a microarray to investigate DNA differences between the samples -- from ten different locations on a large Elkhorn coral colony that had produced single-parent offspring, and samples from five neighboring colonies at nearly 20,000 genetic locations. The results showed that all six of the separate coral colonies belonged to the same original coral genotype (known as a "genet"), meaning essentially that they were clones derived from a single original colony through asexual reproduction and colony fragmentation. Thus, any genetic variation found in these corals would have been the result of somatic mutation. The team found a total of 268 somatic mutations in the samples, with each coral sample harboring between 2 and 149 somatic mutations.

The team then looked at the single-parent offspring from the parent Elkhorn coral colony and found that 50% of the somatic mutations had been inherited. The exact mechanism of how the somatic mutations make their way into germline cells in the corals is still unknown, but the researchers suspect that the segregation between body and germline cells in corals may be incomplete and some body cells may retain the capacity to form germ cells, allowing somatic mutations to make their way into offspring. They also found evidence for the inheritance of somatic mutations in some offspring from the mating of two separate coral parents but will need additional studies to confirm this.

Read more at Science Daily

Jun 10, 2022

Most 'silent' genetic mutations are harmful, not neutral -- a finding with broad implications

In the early 1960s, University of Michigan alumnus Marshall Nirenberg and a few other scientists deciphered the genetic code of life, determining the rules by which information in DNA molecules is translated into proteins, the working parts of living cells.

They identified three-letter units in DNA sequences, known as codons, that specify each of the 20 amino acids that make up proteins, work for which Nirenberg later shared a Nobel Prize with two others.

Occasionally, single-letter misspellings in the genetic code, known as point mutations, occur. Point mutations that alter the resulting protein sequences are called nonsynonymous mutations, while those that do not alter protein sequences are called silent or synonymous mutations.

Between one-quarter and one-third of point mutations in protein-coding DNA sequences are synonymous. Ever since the genetic code was cracked, those mutations have generally been assumed to be neutral, or nearly so.

But in a study scheduled for online publication June 8 in the journal Nature that involved the genetic manipulation of yeast cells in the laboratory, University of Michigan biologists show that most synonymous mutations are strongly harmful.

The strong nonneutrality of most synonymous mutations -- if found to be true for other genes and in other organisms -- would have major implications for the study of human disease mechanisms, population and conservation biology, and evolutionary biology, according to the study authors.

"Since the genetic code was solved in the 1960s, synonymous mutations have been generally thought to be benign. We now show that this belief is false," said study senior author Jianzhi "George" Zhang, the Marshall W. Nirenberg Collegiate Professor in the U-M Department of Ecology and Evolutionary Biology.

"Because many biological conclusions rely on the presumption that synonymous mutations are neutral, its invalidation has broad implications. For example, synonymous mutations are generally ignored in the study of disease-causing mutations, but they might be an underappreciated and common mechanism."

In the past decade, anecdotal evidence has suggested that some synonymous mutations are nonneutral. Zhang and his colleagues wanted to know if such cases are the exception or the rule.

They chose to address this question in budding yeast (Saccharomyces cerevisiae) because the organism's short generation time (about 80 minutes) and small size allowed them to measure the effects of a large number of synonymous mutations relatively quickly, precisely and conveniently.

They used CRISPR/Cas9 genome editing to construct more than 8,000 mutant yeast strains, each carrying a synonymous, nonsynonymous or nonsense mutation in one of 21 genes the researchers targeted.

Then they quantified the "fitness" of each mutant strain by measuring how quickly it reproduced relative to the nonmutant strain. Darwinian fitness, simply put, refers to the number of offspring an individual has. In this case, measuring the reproductive rates of the yeast strains showed whether the mutations were beneficial, harmful or neutral.

To their surprise, the researchers found that 75.9% of synonymous mutations were significantly deleterious, while 1.3% were significantly beneficial.

"The previous anecdotes of nonneutral synonymous mutations turned out to be the tip of the iceberg," said study lead author Xukang Shen, a graduate student research assistant in Zhang's lab.

"We also studied the mechanisms through which synonymous mutations affect fitness and found that at least one reason is that both synonymous and nonsynonymous mutations alter the gene-expression level, and the extent of this expression effect predicts the fitness effect."

Zhang said the researchers knew beforehand, based on the anecdotal reports, that some synonymous mutations would likely turn out to be nonneutral.

"But we were shocked by the large number of such mutations," he said. "Our results imply that synonymous mutations are nearly as important as nonsynonymous mutations in causing disease and call for strengthened effort in predicting and identifying pathogenic synonymous mutations."

Read more at Science Daily