Curated by Dewpoint Therapeutics
Contribute Contact
search icon
Condensates Logo
  • Condensates 101
    • Introduction to Condensates
    • Key Publications
    • Condensates Glossary
  • Publications & Events
    • Publications
    • Kitchen Table Talks
    • Community Events
  • Resources
    • CD-CODE
    • C-mods
    • Condensatopathies
  • Expert Voice
  • Subscribe for Free
  • Log In
  • My Account
  • Subscribe for Free
  • Log in

Resources

    View Your Favorites

    Access your saved content with one easy click.

    My Favorites

    Save Your Favorites

    Subscribe for free to save your favorites and view a personalized view of content.

    Subscribe for Free

    Be a Contributor

    Submit new content suggestions for Condensates.com.

    Contribute

Home

VIDEO: Kitchen Table Talk: Agnes Toth-Petroczy and Francis Carpenter on Transforming Condensate Biology into Therapeutics: The CD-CODE Knowledge Platform

Type Kitchen Table Talk
Topics
  • Artificial intelligence
  • Biology and Physics of Condensates
  • Biotechnology and engineering
  • Cancer
  • Cardiopulmonology
  • Drug Discovery
  • Infection
  • Metabolism
  • Neurology
  • Technology
Tags
  • Biomolecular condensates
  • Intrinsically disordered proteins
  • Membraneless organelles
  • Nuclear speckles
  • Nucleoli
  • P bodies
  • Phase separation
  • Reviews
  • Stress granules
Share

Condensates.com welcomed Agnes Toth-Petroczy, Research Group Leader, Max-Planck Institute for Molecular Cell Biology and Genetics and Francis Carpenter, Head of Data Science and Engineering, Dewpoint Therapeutics, to deliver our first Kitchen Table Talk of the year.

Title: “Transforming Condensate Biology into Therapeutics: The CD-CODE Knowledge Platform”

Abstract: The exponential growth in biomolecular condensate research has created a need to systematically organize the vast and complex data publicly available on condensates. To address this need, the CrowDsourcing COndensate Database and Encyclopedia (CD-CODE) was launched in 2023. Developed by the Toth-Petroczy Lab in collaboration with Dewpoint Therapeutics and the Hyman Lab, CD-CODE.org is a unique platform integrating experimental data with crowd-sourced expertise from the condensate community. This centralized resource has become indispensable, supporting hypothesis generation, guiding experimental design, and enabling the training of machine learning models to predict condensate-forming proteins. This talk introduced CD-CODE and explored examples of how it powers condensate research.

Click here or on the video to view the engaging discussion and read the transcript below.

TRANSCRIPT

Diana Mitrea: Hello everyone and welcome to Kitchen Table Talk number 42. It is my great pleasure to introduce our first, back-to-back dual-speaker event. I’m Diana Mitrea, I’m your host, and I’m joining you from Dewpoint’s Kitchen Table in Boston, Massachusetts.

Before we get into today’s talk, a few housekeeping messages. This talk, as you know, is a live event and will be recorded, so feel free to turn on your camera and follow along with the speaker, but please, no on-camera shenanigans. We would prefer to keep the questions to the built-in Q&A breaks, so feel free to type your question in the chat at any time.

During the Q&A, I will call your name, ask you to unmute your microphone, and at that point, you can ask your question directly. If, for whatever reason, you’re unable to talk, please let me know in the chat, and I’ll ask the question for you. The event recording will be posted on Condensates.com within a week.

So, now to the part for which you’re here. It is my great pleasure to introduce today’s first speaker, Dr. Agnes Toth-Petroczy. She’s a Group Leader at Max Planck Institute of Molecular Cell Biology and Genetics in Dresden, Germany. Agnes is renowned for her contributions in computational biology and protein evolution. She has earned a master’s degree in chemistry from Eötvös Loránd University in Hungary, and a PhD in Life Sciences from Weizmann Institute of Science in Israel. She completed her postdoctoral training right here in Massachusetts at Harvard Medical School and Brigham and Women’s Hospital. Throughout her career, Agnes contributed to our understanding of evolution of disordered proteins and their contributions to Mendelian diseases.

She mentored numerous trainees and supported equality, sustainability, and community outreach efforts. In 2025, she was honored with the Ernst Schering and EMBO Young Investigator Awards. Since 2018, as a Group Leader, she contributed to the development of several bioinformatics tools that facilitate the study and advancement of intrinsically disordered proteins and biomolecular condensate science. One of these, CD-CODE, is the topic of today’s talk. So without further ado, Agnes, the floor is yours.

Agnes Toth-Petroczy: Thank you, Diana, for this very nice introduction, and also, thank you for giving me the opportunity to present CD-CODE. I still remember when we started brainstorming about the idea of creating a knowledge platform about condensates back in 2020, during the COVID lockdown. Back then, there wasn’t even this AI revolution in site, and now we know that AI is revolutionizing our lives and also drug discovery.

In my lab, particularly, we are interested in intrinsically disordered protein regions, or IDRs for short. These proteins and regions don’t have a well-defined structure, and because of this, they are also considered undruggable. And often, these disordered regions actually drive the formation of biomolecular condensates, which could be a way around this, and targeting, actually, with these condensates instead of the disordered regions.

And while AlphaFold, as one of the breakthroughs in structural biology, basically solves the problem of sequence-to-structure paradigm for ordered proteins, and it can predict the structure of the ordered regions, it still cannot predict the structure of the disordered parts of the proteins. It gives only low-confidence predictions, and this is something my lab is interested in. How can we actually advance the sequence-to-function paradigm of disordered proteins?

But let’s discuss, why does AlphaFold work, or what were the requirements for it to be developed? And, of course, the development in computer science and the breakthroughs in new AI algorithms contributed, as well as the hardware development of recent years. But very importantly, also the training data was absolutely necessary to develop this tool.

And here we are talking about two types of data in the case of AlphaFold. So it needs evolutionary data, so all the sequence information from hundreds and thousands of different organisms that we have sequenced in recent years, and the decades of structural biology efforts of actually collecting structures in PDB as a main resource. So the community basically agreed to deposit these structures that they have solved in the PDB, and also developed standards for what is a good structure. And if we just have a quick estimate of what is the cost of data actually in PDB, that I asked ChatGPT, it comes up with this 3 to 20 billion Euro estimate for actually the data that was deposited there.

But I would argue there is even one more ingredient that is important for success and for these breakthroughs, which is the challenges that the communities actually developed. And this is, for example, the Critical Assessment of Structure Prediction, or CASPP, that AlphaFold won in 2021, when the community actually agrees what are the questions that we want to tackle, what are the challenges to solve, and what are the metrics of success. And there are several other challenges for predicting interactions or functions, also predicting intrinsic disorder. There is no challenge yet for condensates. And all of these, of course, use the datasets and databases that were developed and can be used for training and also for evaluation, for structure, sequences, function annotations, and also disorder. And now, as of 2023, when we released the first version of CD-CODE, that is data also on condensates, that followed four previous databases that were mainly focusing on liquid-liquid phase separation and also collected proteins that are involved in those.

Our goal was that we would like to assemble the condensate data and kind of catalog it in a way that it’s useful for systematic discovery and also for AI modeling. And since the condensate field is very new, there is actually a zoo of different terms that you can find in the literature, and what the community actually defines as a condensate, and what experimental evidence we actually accept as a proof for a condensate to exist. So we decided we will go with the term biomolecular condensates. This was coined in 2017 in a landmark review.

And we distinguish, basically, biomolecular condensates from synthetic condensates, which would be condensates for phase separation actually assayed in vitro, when only a few, one or few components are actually studied, from condensates being observed in cells or in the organisms themselves.

And regarding the protein components, we distinguish drivers, which would be proteins that are essential for the formation of the condensates, and if we would knock out this protein, we wouldn’t actually see them. And often, these are scaffolds and kind of binding many, many different components. While members, on the other hand, or clients, they can partition in and out of the condensate without compromising its integrity, and they could potentially modify its function.

Because of this, all these many terms, we had to make some decisions. So we decided we were building a crowdsourcing functionality into the database, so the community can contribute and edit it, and we named it Crowdsourcing Condensate Database and Encyclopedia, or CD-CODE for short. This was truly a collaborative effort between Dewpoint, where from the beginning, Diana was our main contact and the main lead, and Tony Hyman’s group and my group, the first version headed by Nadia and Deep. And the first version already had around 10,000 proteins and 200 condensates, which since then actually doubled. In the background, there are many other people contributing, of course, both on the software side and on the data side. And we even organize these community events at my institute, at MPI-CBG, which we call “Condensathons” when we get together, have free pizza, or celebrate with a cake, and discuss papers, and decide on the manual curation of many of the entries. Although we can use AI, and we of course use LLMs to help with the curation of the papers and the data, there is need still for human checks and interactions. Our policy was that the curators actually become co-authors, so if you are interested in curating data, then join us, and you can be part of the next release and potentially the next publication.

While in the first version of CD-CODE, we focused mainly on the function and the components, in the second version, we decided also to include data that has biomedical relevance. So, we started curating infectious condensates that are condensates that form upon viral or bacterial infection, and drugs that impact condensates that are called c-mods and condensatopathies, which are diseases that are linked to condensate dysregulation or malfunction, and this second version was headed in by a group by Ksenia and Maxim.

We are also putting effort into improving the user experience, and you can see many people actually just browsing the website, and these are the terms that that are the most commonly searched terms on the website, like highlighting the most famous phase-separating proteins. But also, for programmatic access, we have an API that we have now improved documentation for many computational researchers or AI developers as well.

To give you an example here, you can see and select, for example, infectious condensates, that gives you a long list with the different components annotated, and then a description links to different other resources cross-referenced, and then always gives you, basically, an evidence for what type of experiments are supporting, let’s say, that this protein is part of that condensate.

And this is actually crucial, because depending on what type of evidence you’re looking at, for example, many times we just look at microscopy-based condensate identification and localization, and we actually don’t know what other proteins are in these condensates, so if we just look at the distribution of how many proteins a condensate contains, many of them look like they have only a few.

But a few condensates have been actually purified, and mass spectrometry analysis suggests that there are hundreds and even thousands of different members inside these condensates. Also, if we look at the proteomes of different condensates, you can see that they overlap. And, for example, in stress granules and p-bodies, which share 300 different members that I’ve shared between them, also potentially suggesting their different origin, so many proteins actually can partition into several different condensates.

We also looked at the disordered content, of course, and we can clearly see, if we compare to the human proteome the different length distribution of IDR regions to the condensate proteome in CD-CODE, that there is an enrichment for IDRs. However, there is still 21% of proteins that have no disorder content at all, and they are still found in condensates, which could be potentially modulated like kinases or, let’s say, clients. But at the end, it turns out that even drivers can be fully structured, so we can put them in this continuum where there are disordered proteins driving condensates like FUS, but there are also fully structured proteins, like SPOP, that are found in nuclear speckle, driving nuclear speckle formation.

So this data that we have now systematically in CD-CODE enables testing certain hypotheses that the field had, or also highlights the unknowns, let’s say the members that we still don’t know. And it also enables machine learning and AI applications. So in my lab, as a computational biologist, we also are interested in developing methods for condensates. And particularly, we are interested in predicting from a sequence of a protein if it would form condensates, and which particular condensate it would localize into, and also how would mutations in the protein sequence impact a condensate.

I will briefly introduce you to our method, that is PICNIC, that predicts condensate formation, but also CD-CODE data was used to develop ProtGPS, which is a tool that can predict the specific localization of a protein into 11 different condensates developed by Rick Young’s group and catGRANULE2.0 that is predicting the impacts of mutations on the protein sequence.

It’s great that we have CD-CODE, because this can be used as positive examples for a machine learning predictor. But for supervised learning, we also need negative examples, which are much more tricky to find. So Anna, in my group, thought about basically using the protein-protein interaction network, in human in this case, and define as negative those proteins that are not known to interact with already existing in condensate proteins that we have already proved that they form condensates. So they’re at least two steps away in this interaction network from non-condensate proteins, and this was our way of having kind of an unbiased negative data set for training a predictor. And, our predictor is called PICNIC, that stands for Proteins Involved in CondeNsates In Cells. It uses sequence and structure-based features, and then at the end basically classifies into “yes” or “no” forming condensates, regardless of also structural disorder. It predicts around 1,500 new condensate proteins in humans that have not been characterized yet.

We tested some of them experimentally with, Hari, Tony Hyman’s postdoc and we were happy to see that overall it has an 80% accuracy, as we can see here in blue, all the proteins that we tested, and indeed formed condensates, and this can have a wide range of structural content that can be highly disordered, but also highly structured.

PICNIC can also be used, of course, not only in human data, but across any organism, or even any synthetic sequence. So we wanted to see what are the condensate components of different organisms? And while we know that disordered content increases in evolution, so compared to bacteria and archaea, there is a lot more disorder in eukaryotes. There was an expansion. We don’t know the same about condensates.

When we looked across 24 different organisms, across the Tree of Life, we were surprised to see no correlation with disordered content. It seems to be that condensates are ubiquitous, and in archaea, there is even a higher condensate content predicted than, let’s say, in mammals. So, if a condensate seems to be an essential and ubiquitous across species, then also their dysregulation may lead to disease.

This is what we started collecting. Led by Allysa and Diana at Dewpoint, we now have a list of different diseases that are known to be associated with some condensate dysregulation. Basically, we collect which condensate is impacted in what disease, and what dysregulation type that actually links to. And this can be either a direct link to the pathophysiology of the disease, or these dysregulations could be treated as a biomarker of the disease, a kind of marking the disease. When we look at the most common stage-related diseases, they are actually either infectious, or cancer, or neuronal. And the most likely condensate to be affected is stress granules, but there is also a long list of other condensates that have been already implicated in diseases.

If we can observe that under disease conditions, there are differences in condensates, then this could be potentially a way to target these condensates. And this is what these condensate modulator drugs are doing. According to Dewpoint’s review, landmark review, we classify them into these four categories. So a c-mod could dissolve a condensate, or induce the formation of the condensate, changes localization, or changes properties, like material properties or shapes, which we call morphers.

We collect now all these small molecules in the database and cross-reference them again with known other chemical databases. And when we looked at these over 350 c-mods that we right now have in the database, then interestingly, most of them either dissolve or induce the condensates, but some can have different effects depending on which condensates they impact. And also, they kept different modes of action. And when Diego from Dewpoint looked at the chemical space of these different molecules, then he couldn’t see any clustering based on the mode of action. So, there are a wide variety of different chemistries that can give rise to the different actions, depending on which condensate we are also looking at.

What we are thinking about the future is that, so far, we collected information on condensates in CD-CODE about is there a condensate or not, and that was kind of the evidence which we were looking for. But we should also think about describing these condensate phenotypes, and also a bit more quantitatively, actually, by looking at the images and classifying them into these different categories.

And that’s what we started working on now, so developing a tool that can take hundreds or thousands of images and actually describe the condensate morphology, spatial distributions, or phase behaviors, let’s say concentration dependence, and give us a way to actually quantify the different condensate phenotypes. So, looking forward in the next version of CD-CODE that we call iCD-CODE, we will collect also these phenotypes and the images that are actually the raw data to say to why we think these condensates and these different proteins are actually there. And we also would like to collect information on mutations in the protein sequence, that would impact the condensate behaviors.

If you have any other ideas for what we should include, and what type of developments we should do, what’s new data we should collect, then please reach out and write to us. And, with that, I would like to thank you for your attention and am happy to take questions.

Diana Mitrea: Thank you, Agnes. Do we have any questions online? Please raise your hand or type your question in the chat.

We have a question here from Italo.

Italo do Valle: Hi, how is context dependency included in the database, and also in the tools that your lab has been developing? For example, the condensate formation which contacts tissue type and even condition?

Agnes Toth-Petroczy: At the moment, we don’t have it actually. We started collecting this information, but it’s not visible from the front end. We started collecting the cell types, so this is a very good direction forward, and I don’t know what other context we should record, so please suggest what you think we should add as information, because for now, we basically have only the experimental evidence, and if it’s in cells, or in vivo, or in vitro, we could add the cell types, so what else should we collect?

Italo do Valle: Yeah, I think cell type is a good start, but also the experimental evidence, the type of experimental evidence, because of the biases of each experimental approach to detect them.

Agnes Toth-Petroczy: So the experimental evidence we have, like, roughly, we have if it’s microscopy, or FREP has been done, or mass spec. And the cell type we can add, and I think if we have the images, so if we actually start collecting the raw images, which would be actually great if the community would also think this is useful. So similarly, how PDB collects the 3D coordinates, right? This is basically what you have to deposit. If we would deposit the raw images, then from that, it’s very easy to classify what was the cell type, what were the conditions, what was the concentration, because that’s something that we could actually also measure from these images in a standardized way. Maybe it’s even easier than digging it out from publications.

Giancarlo Franzese: Nice talk, thank you, and I’m just curious to know if your database can provide information about the structural changes of the proteins within the biomolecular condensates, when we are referring to structured proteins.

Agnes Toth-Petroczy: No, not at this point. So it’s basically only presence, absence of a given protein in a condensate, and it’s also kind of the union of all potential experiments that have been done about this protein. You would have all the links to the corresponding publications, but we don’t actually record the confirmations. And I think there’s also not that much data yet in the literature about this. It would be actually great to have that information.

Giancarlo Franzese: Indeed, indeed, that would be definitely something very important to add to your database. Good to know, thank you.

Diana Mitrea: We have another question in the room from Jian-Guo.

Jian-Guo Ren: This looks very interesting; however, I understand it’s quite complicated. So my question is that, do you have any plans to build a condensate map across different cell types or different diseases and use it to guide the drug screening?

Agnes Toth-Petroczy: Yes, the cell types we should add. That’s still missing, to have this full map of information. And I think also the canonical condensates would be the one that I named, let’s say nucleolus and nuclear speckles, but actually there are a lot of condensates that just have a few components that are named after the proteins, and these could be actually the same condensates, just we don’t know that these proteins potentially colocalize. So I think we don’t even know how many condensates there are still in the cell. That’s why it actually didn’t say that number, because I think this could change.

Jian-Guo Ren: That’s true. My concern is that I don’t care about single proteins, single pathways. For example, if two countries are fighting, the DoD building, they may be made up of different people, so what even the information, but you know the image, then you can hit it, then you can hit your enemies. This is my question. With that being said, if you can build a condensate map of different buildings having different functions, then in the future, you screen drugs for let’s say breast cancer or ovarian cancer, then you know we can hit this building, then we can screen drugs. The real world is not as simple as this one, but however, if you have a very cool map, it may be useful.

Agnes Toth-Petroczy: Yes, maybe in Francis’s talk, there will be more information of this map, and how we can aid that.

Diana Mitrea: We’re going in that direction with the condensateopathies recording. But it’s building. It’s not there yet. But I think that that’s on path with what Jian-Guo was talking about.

If there are no other questions, let’s thank Agnes, and I will introduce our next speaker, which is Francis Carpenter, the Head of Data Science and Engineering at Dewpoint Therapeutics, and he will continue the story, telling us how tools, like the ones that Agnes introduced, could be utilized and integrated into the drug discovery process.

Francis has a PhD in neuroscience from the University College of London, where he combined in vivo electrophysiology, chemogenetics, and computational modeling to study the mechanisms underlying the brain’s representation of space, and how this supports long-term memory. Prior to joining Dewpoint, Francis was an associate partner at McKinsey & Company and at Quantum Black, where he led cross-functional teams, supporting biopharma clients to leverage state-of-the-art AI approaches to enhance their drug discovery and development efforts. At Dewpoint, Francis leads development and scaling of innovative, data-led approaches to enhance the company’s drug discovery and clinical translation efforts. Francis, the floor is yours.

Francis Carpenter: Thank you for the introduction, Diana, and for the great intro to condensates, Agnes. What I want to do is give a perspective from the kind of Dewpoint lens. Of course, at Dewpoint, we’re interested in tackling diseases through a condensate lens and finding drugs that modulate condensates to tackle hard-to-treat diseases, and therefore we come to address both our drug discovery and our research into condensates using datasets like CD-CODE, which have been a great resource, and so I want to give two examples today, one of how we use CD-CODE to support target discovery, and the second one about how we use it really to understand more about condensate modulating drugs.

But before I go into that, and before I forget, I want to say a big thank you and shout out to everyone at Dewpoint, present and past, who’s been involved in this work. I haven’t done as good a job as Agnes highlighting everyone for their individual contributions, but everyone who has been involved knows who they are, and it’s been a huge collaborative effort, so I want to say thank you up front.

So, to start off, what I’d like to do is talk through an example of how we use CD-CODE among other datasets to look for novel condensate targets. And I’d like to start off by saying a few words about what a condensate target means to us at Dewpoint. We think a little bit differently about how we define a target at the beginning of a drug discovery program compared to other companies that might have a more kind of traditional or canonical definition of a target, where they are trying to find a specific protein, which maybe has a deep pocket, which can be well bound by an individual small molecule, or an antibody, or something like that.

What we’re thinking about in terms of a condensate target at Dewpoint is a condensate whose phenotype reads out a disease state. And what we mean is a condensate where the phenotype is different between the healthy condition and the disease condition we’re interested in, in a way where that phenotype tells us about the underlying dysfunction of the disease. The point being, therefore, that we can use these condensate readouts to find small molecule drugs that can reverse the disease state back towards the healthy state. What we do a huge amount of work on is using high-content confocal imaging of condensate phenotypes, together with associated assays, like functional assays or counter screens, to try and find small molecules that can revert the disease state. And that is quite differentiating, because that can help us find novel condensate modulator mechanisms, where we don’t have to define a pirori the specific target that we want to bind to. We’re looking for small molecules that actually have a network effect in reversing the condensate dysfunction of the disease back to its healthy state.

We also, of course, come to our discovery programs and to our target discovery efforts from the condensate lens in terms of the data that we integrate. I’ll spend most of this section of the talk talking about how CD-CODE fits in, but of course it represents condensate-specific data and knowledge that we integrate into trying to find new targets that read out on these diseases we’re interested in, alongside the more general data sets and knowledge bases that pretty much everyone else in the space uses in terms of protein-protein interaction networks, etc. And we believe this helps us, integrating knowledge like CD-CODE generate hypotheses that others wouldn’t come to without looking through this angle and this lens.

And lastly, we also validate the targets, the condensate targets that we’re interested in, in a different and, I think, important way compared to how a lot of other folks think about validating targets. So we combine high-content imaging, for example immunofluorescent staining of specific condensates with antibodies, with omics, to look for these differences between a healthy and a diseased state, which we can use as a readout.

And that’s important because, for example, if we just look at transcriptomics, or we just look at general morphological changes, for example, with cell painting, we might miss differences in the condensate, which is key to the disease, but maybe it’s not picked up in transcriptomics because it’s a change in the tertiary structure of the involved proteins or the condensate, or maybe a change in its localization.

We look at targets in a different way. We bring different data sets and a condensate-specific lens into how we look for them, and then we actually validate them in a sort of condensate-appropriate way. And this is a sort of 30,000-foot view of what makes our thinking about target discovery unique at Dewpoint.

What does it actually look like in practice? We’ve applied this similar approach across many programs with great success in multiple disease areas now. It always looks slightly different depending on the disease biology. This slide gives a high-level representation of the broad brushstrokes. And broadly, what we do is we combine omics knowledge with other datasets and integrate them together through knowledge graphs into networks, which we can then apply algorithms to and machine learning to prioritize potential condensate hypotheses.

So just to step through this in a little bit more detail In the first step, we start with gene disease or genes, which are nodes which seed our knowledge graph network. And we identify these genes based on omics associations with the disease of interest. So that might be a GWAS study, or some transcriptomics, or phase-specific proteomics, for example where we’re looking for genes who have a statistical association with the disease. And what we then do is use other datasets to expand those disease-associated genes into a network of multiple different node types, which represent the biology underlying the disease.

And this is where we bring in the condensate information from CD-CODE and other sources, and that’s what I’ll talk about more in the next few slides. We also then make predictions about which genes are likely to form condensate, or to phase separate, using models like PICNIC, which Agnes talked about. And we combine those predictions with network algorithms to basically look for nodes that we call integrators, so that well-connected nodes at the heart of the disease biology and the disease network which are likely to form condensates, which we think are potential regulatory hubs. If we can modulate those hubs, those condensates, we have a good chance of redressing the disease state back to the healthy state.

From these algorithms, we then nominate shortlists or long lists, depending on the capacity of our wet lab colleagues, for them to go and validate, using imaging, for example, as I said, immunofluorescent imaging or omics, and also using things like tool compounds or gene perturbations to validate that modulating that condensate actually will have the effect we’re interested in. What we then do is prioritize the ones we’re most interested in, and we go and use these as the starting point of drug discovery campaigns. So, they go into high-throughput screening, where we look for novel small molecules, novel chemical matter that will modulate the condensate target of interest and could act as potential drugs for those diseases.

I don’t have time to go into all of those steps in detail, but I want to talk a little bit more about building these disease-specific knowledge graphs and where CD-CODE fits in.

This slide talks a little bit about how we go from individual genes, or sets of genes that are associated with the disease of interest, through to the kind of graph model of the disease biology.

These disease-associated genes usually come from population-level omic studies, for example, GWAS, or transcriptomic studies. And sometimes we add on algorithms that look at the regulatory networks or predict how these genes interact with each other. We then use these as a starting point to expand, step by step into this model. So we start by adding in other connections from literature and public datasets that we know each of the disease-associated genes have. So, for example, what other genes or what other proteins do they interact with? Do they interact with RNA? Are they expressed in specific tissues? Do we know that there are certain drugs that interact with these, etc.? And then we also build in the condensate information, and that’s where CD-CODE comes in, where we add in, for each of the genes, information of the type Agnes illustrated. For example, does the protein or the gene phase separate into a known canonical condensate. Is that condensate associated with the disease? And are there any c-mod tool compound drugs out there which interact with those condensates.

And by combining all this information together, we also use these machine learning models to predict which of the genes are likely to phase separate. And so building these steps together, we end up with this network representation of the disease of interest. It’s effectively trying to capture in silico a model of all the relevant interconnectivity and pathways that underlie the dysfunction of the disease, so that we can try and look for hubs to go and drug, and condensates to go and modulate.

And that was probably a bit of an abstract hand-wavy representation of what we’re trying to do, so I wanted to add this slide to make it a little bit more concrete. This slide is just an example small slice or section of one of these knowledge graphs that I’m talking about. And this is focused on a graph that some colleagues built for dilated cardiomyopathy, so a cardiovascular disorder, genetically underpinned. And it has at the center a condensate of interest called BAG3, or a gene called BAG3 that forms a condensate.

And you can see on the right here the sort of simplified schema that was baked into this knowledge graph, where we have genes that can interact with each other, that can be part of condensates are known to participate in biological pathways, can be targeted by compounds, and are associated with diseases. And we build into the node size a sort of representation of how likely each gene is to form a condensate.

And so these networks, you can see, this is just a small slice of the network of nodes that are one hop or one connection away from BAG3. But you can see they very quickly get very complicated and hard to follow visually, so that’s why we start using algorithms to analyze them and try and find out, what are the algorithmic hubs at the center of these networks. Hopefully that gives you a bit of an idea of where CD-CODE fits in, and I just want to end this section with one case study of how powerful this approach can be, and how important, you know, it can underlie drug discovery programs.

This is an example of work from many colleagues that we have been doing over the last few years in collaboration with Novo Nordisk where we’ve been particularly interested in type 2 diabetes and insulin resistance. And we used a very similar approach to validate more than 60 novel condensate targets. We started prominently with genes associated with type 2 diabetes from GWAS analyses of UK Biobank. And we then, from those genes, built a graph representation of the disease biology of type 2 diabetes, including condensate information.

We used the algorithms that I mentioned to then nominate 200 condensates, which we predicted to phase separate, or predicted condensates, which also were sort of regulatory integrator hubs within the network. And those were then tested in both healthy states and disease states by colleagues with high-content imaging. And you can see some examples at the bottom of specific condensates that had a statistically significant difference in their phenotype between the healthy state and the disease state. So, for example, in terms of the shape, or the intensity, or even the localization of the condensate that we could potentially use as a phenotype for a high-throughput screening campaign.

And a subset of those were then further validated with things like functional assays or perturbations and tool compounds. We prioritized five, and we’ve since been following up into full high-throughput screening programs, and we have some interesting, you know, early candidate drugs coming out of those programs, which is very exciting. So that kind of wraps up a whistle-stop tour example of how we feed CD-CODE into our drug discovery efforts.

And in the interest of time, I’ll move on to the second example before I can jump and take questions at the end. Less directly feeding into our drug discovery efforts but relates more to how do we do more basic research into condensates and condensate-modulating drugs, c-mods through CD-CODE. One way that we’ve used CD-CODE is combining it with some internal data from an experiment called D.paint to ask questions about how the physicochemical properties of c-mods differ from non-c-mods. From CD-CODE, we took the c-mods that have been curated in the second iteration of CD-CODE. And we filtered it down to small molecule c-mods that pass some filters, they’re not flagged by PAINS filters, and they’re less than a kilodalton. So we have about 250 of those. We combined it from data from D.paint, which is an internal assay developed by colleagues of mine at Dewpoint whereby they ran high throughput, high-content imaging of tens of condensates in parallel. In this specific data set, it was about 12 condensates with 14 different antibody-based markers. And you can see some of the example condensates. There are canonical condensates, like stress granules and nucleolus, for example.

And they ran around 4,000 compounds from a diversity library of small molecules, so these are all compounds. Some of them are approved drugs, some of them have been tested in clinical trials but were not approved, or some of them are sort of drug-like preclinical compounds. And colleagues ran analyses to basically split this library into c-mods, which were compounds, which modulated the phenotype of at least one condensate in the panel and non-c-mods, which had no effect on any of the condensates.

It’s a bit a flawed definition, because some of these non-c-mods, you know, might be c-mods in other contexts. For example, they might modulate other condensates that we haven’t looked at, or they might modulate the same condensates in different cellular contexts, for example. So this was just in a single cell line.

But regardless of some of the limitations, this gives us a c-mod and a non-c-mod set of data, or multiple sets of data, where we can start asking questions like: Are there chemical distinction or physicochemical properties that differentiate c-mods and non-c-mods and how do they differ? Do we see differences across seamless of different mechanisms, et cetera?

These are the sort of examples that I want to spend the next few minutes giving you illustrative answers to of the kind of interesting insights we find. On this slide, which is a little bit dense, but, hopefully I can walk through it without too much confusion, is looking at the different physicochemical properties between c-mods and non-c-mods, whereby the analysis that we’ve done on both datasets suggests that c-mod have a distinctive, sort of, fingerprint of chemical properties compared to non-c-mods. And in particular, the c-mods are characterized by higher valency, by the presence of more rings, particularly more aromatic rings. They tend to be less soluble, and they’re often larger in terms of molecular weight than non-c-mods.

The left-hand side is 16 different box and whisker plots. They’re probably a bit hard to see, but they represent different chemical descriptors, with the c-mods in this pinkish color and the non- c-mods in blue. And the purple stars are labeling the 13 out of the 16 of these properties where we see a statistically significant difference between the c-mods and the non- c-mods. And you can’t see from these data, but when you look into the actual metrics and the tests, the properties that have the greatest difference are these ones I mentioned to do with valency, aromaticity, solubility, etc. This makes a lot of sense, given what we know and what Agnes described about condensates, that they’re these dynamic assemblies of biomolecules held together by often, many weak interactions between disordered regions of the constituent biomolecules.

And therefore, it seems intuitive that c-mods or small molecules that interact with condensates, that is the c-mods, have these properties that seemingly would make them maybe more able to interact with the many weak interactions that govern the condensates. For example, that they’re larger potentially gives them a greater surface area to interact with more weak interactions. Higher valency and more aromatic rings potentially enables more interaction with the kind of weak electrostatic forces that govern the condensates.

And we also know that condensates often have a distinct and often more hydrophobic solvation environment internally, so maybe the hydrophobicity of the c-mods helps them partition into the condensate. What’s really interesting is that even with these relatively simplistic and flawed data, where there’s a lot of caveats, we can see this quite striking difference between the c-mods and the non- c-mods.

And I think what’s important for us is that these kinds of findings are both interesting, but they also have implications for how we think about drug discovery. For example, if you’re thinking about designing a library for a screen to try and find c-mods, you might look towards biasing the compounds towards having these kinds of properties. Or if you’re trying to optimize c-mods with a lead optimization campaign, you might look to change the heuristics from those that are used in traditional medicinal chemistry towards, some of the ways that c-mods differ.

I won’t dwell on this slide too much, but we also analyzed similar data from a machine learning perspective, rather than a traditional kind of statistical perspective. So we built a machine learning model, or many iterations of a similar model in a crossfold, tenfold cross-validation setup to predict whether a given small molecule was a c-mod or not. You can see on the left-hand side, in the dark grey, the performance of the model across validation folds, that is, on basically compounds the model had not seen during training. And the model, was the training set and the testing set was balanced. So there was the same number of c-mods and non- c-mods, so chance performance is basically 0.5 on all of these metrics.

And you can see that the model does, across the board, statistically significantly better than charts. It’s basically further evidence that, indeed there does seem to be some kind of fingerprint which the model can learn to distinguish the c-mods and the non-c- mods better than you would expect if no such fingerprint existed. And when we look at the features that the model used most to guide the prediction, again, it’s a similar kind of set of things that were highlighted by the traditional metrics, so things like valency, solubility, molecular weight, etc. Intrigued by this fingerprint, we wanted to also see what other kind of questions we could ask about CD-CODE. A natural place that we looked next was into mechanisms and the mechanisms of action of the drugs.  CD-CODE includes information about the targets of the small molecules. And a lot of them are quite granular, and therefore, you don’t have that many of different drugs for each target. It was hard to quantitatively analyze the very granular annotations, but what we did was aggregate them up into kinase and non-kinase inhibitor categories. And when we did this for the CD-CODE dataset, and actually the same for our D.paint dataset, we see that the c-mods were enriched in kinase inhibitors. You can see on the left, about 34% of the compounds in CD-CODE, once we applied the filtering, were kinase inhibitors, compared to about 12% in the non-c-mod dataset, which is a, you know, statistically significant enrichment. And again, it’s something that makes kind of intuitive sense, given what we know about condensates and kinases and kinase inhibitors. So, of course phosphorylation adds a negative charge, and so changes the electrostatics and the valency of the condensate that is phosphorylated or the target that’s phosphorylated.

We also know that IDRs often have regions that are heavily targeted by kinases and heavily phosphorylated. So again, it makes intuitive sense that c-mods might be commonly, or enriched in or might frequently modulate kinases, which then either directly or upstream affect the properties of the condensate, and therefore change its behavior.

Just to say again, there’s always a few caveats with these things. I think it’s important to bear in mind that the CD-CODE and also the D.paint data is limited to a single cellular context, as a couple of folks have pointed out. And therefore, there are potential biases we need to keep in mind, like kinases maybe our relatively well expressed across different cell contexts and model systems, maybe more than some other targets, and maybe they’re therefore more identifiable as other targets. But with those caveats in mind, the degree of the enrichment is pretty striking.

And maybe just one last result before I wrap up, which I think is particularly important and promising for us at Dewpoint and those of us trying to find therapeutically promising c-mod drugs is the various evidence that we see that c-mods can be specific to individual condensates and to have structural characteristics or structural activity relationships that map them to specific common sites.

What we did here in the top left is, take the CD-CODE c-mods and convert their structure into a vector representation, so embedding or fingerprint using an AI model. And then cluster the c-mods based on their structure. So this was a hierarchical clustering algorithm, which gave us 20 different clusters of c-mods just based on the c-mod structure. And what we then did was basically check whether any of the clusters were enriched in c-mods of specific condensates.

What you can see, for example, in this cluster 3 up here is a cluster of c-mods with distinct structural properties that are enriched in c-mods for a condensate called INVAVA. And we see similar kind of examples within the D.paint data for different condensates. This is an important indication that part of the expanding evidence that c-mods can be selective or specific to individual condensates, but also that they can have structural activity relationships that map them to a given condensate. And that might reflect that some condensates are preferentially modulated by specific mechanisms of action upstream. Or it could reflect that they have certain biophysical or physicochemical properties that favors interaction with c-mods of a given structure and properties.

What’s important for us, looking for therapies it suggests that we can learn whether in silico, or just based on experimentation, the structure-activity relationships that govern this selectivity. And therefore, we can bake in selectivity to our drug discovery campaigns to try and build therapies that are more selective and therefore have fewer off-target effects over time.

I didn’t show the internal D.paint results. What’s remarkable is that the D.paint data and the CD-CODE code data are very closely aligned, so we see, again, a very similar physicochemical fingerprint for c-mods for CD-CODE and for D.paint, we see similar kinds of enrichment, for example. To wrap up then, it’s probably been a bit of a whistle-stop tour, but, happy to take questions. I think what I’ve hopefully done, or at least tried to do, was to give you two examples that gave you a bit of flavor of how we use CD-CODE code at Dewpoint. Firstly, in terms of an input, along with other data sets into some of our primary drug discovery efforts, for example, for condensate target discovery. And secondly, as a resource or a tool for better understanding condensates and c-mods that then downstream give us insights into their potential as therapies and hopefully help us optimize how we discover c-mods in the future. With that, I’ll wrap up, and happy to take questions.

Diana Mitrea: Great. Thank you so much, Francis, for this great talk. We have a question from Rishav. Would you like to unmute and ask your question?

Rishav Mitra: Hi, thank you for the great talk. I was just wondering if you can tell us a bit more about D.paint. I was very intrigued by the independent sort of validation and so wanted to know a little bit more.

Francis Carpenter: In short, this is an imaging assay which stains for multiple condensates in parallel through antibody markers. These are not stained necessarily in a single well. For example, we might multiplex 4 or 5 condensate markers per well, but then across multiple wells, we can effectively stain for tens of condensates in parallel. And given the high throughput nature of it, we can still then screen thousands of potential small molecules or other kinds of perturbations, like, you know, gene knockouts, for example. We use this assay in a number of ways. At Dewpoint, one example is to look at selectivity or specificity. So, for example, if we found a drug that we’re interested in for one of our drug discovery campaigns, you know, we might run it through the D.paint assay to check whether it interacts with any other condensates, or how it does so to help us understand selectivity.

John Manteiga: Hi, everyone. Similar to how people have used cell painting to try to predict MOAs, where you take known MOA compounds and see if your compound of interest clusters, with them. We can basically do the same thing, just using all the condensate markers instead of the cell morphology markers. That’s another kind of impactful way we’ve used it here.

Lila Ghamsari: I was wondering if you have used TCGA data integration into structural analysis of condensates and generating UMAPs, considering that there are some hottest spots in target proteins.

Francis Carpenter: Yeah, that’s a good question. So, the UMAPs that I was showing, and the analysis that I was showing was based on the structure only of the c-mods, the small molecules, and not of the condensates themselves. So actually, that was just looking at the small molecule structure and not the condensate structure. But I think your point’s a really good one. There’s a lot more that can be done with this, and relating the small molecule structure, the c-mod structure, to the condensate structure in both, with and without mutations, I think is, is a very good suggestion.

Diana Mitrea: Let’s give one more round of applause to our speakers, Agnes and Francis. Thank you very much for your really exciting talks.

And thank you, all of you in the audience, in the room, and online. And I would also like to give a shout out to the team here who makes the KTT possible behind the scenes – Jennifer Tallman, who’s supporting editing and marketing, and Curtis Boskee for IT support.

Lastly, the recording will be posted on Condensates.com and YouTube within a week. Take care and see you all at the next event.

 

  • Introduction to Condensates
  • Publications & Events
  • Resources
  • Expert Voice
  • Subscribe
  • Log In
  • My Account
  • Sign Out
  • About Condensates.com
  • About Dewpoint

Contribute

News, events, resources, and insights by and for the condensate community.

Contribute

© 2026 Dewpoint Therapeutics

  • Privacy
  • Terms
wpDiscuz