VIDEO: Rohit Pappu on Stickers and Spacers
| Author | |
|---|---|
| Type | Kitchen Table Talk |
| Topics | |
| Tags | |
| Share |
Professor Rohit Pappu joined Dewpoint on January 24, 2020, at our Boston office to participate in our series of “Kitchen Table Talks” in which we invite prominent researchers in the field of biomolecular condensates to share their thinking and recent work with the entire community. (Full disclosure, this time we weren’t actually at the kitchen table.) In his talk, Rohit filled us in on his stickers-and-spacers model for phase separation of multivalent proteins, and fielded questions from scientists in our Boston and Dresden sites.
We were honored to have Rohit join us. He has made seminal contributions to the field of biomolecular condensates, in particular the drivers of phase transitions that lead to the formation of protein and RNA condensates, and the role that disordered regions play in these cellular processes.
Rohit is the Edwin H. Murty Professor of Engineering and the Director of the Center for Science and Engineering of Living Systems at Washington University in St. Louis. He is also a member of Dewpoint’s Scientific Advisory Board and a wonderful advisor, collaborator, and friend. We hope you enjoy his talk as much as we did.
Click here or on the video to view the engaging discussion and read the transcript below.
Transcript (machine generated and hilariously uncorrected)
Rohit Pappu
All right, well, I’ll get started today. And basically what I’m going to do is catch everyone up on what we’ve been working on most recently. Which to my mind effectively is anchoring this stickers and spaces model for face separation. I say multivalent proteins here but I will also bring up RNA in the back, third of this talk. So just as a way of providing context and get getting through the background quickly.
It’s now well appreciated that multivalent proteins and RNA molecules are what Tony Hyman and Mike Rosen have taken to referring to a scaffolds that drive Intracellular phase transitions. Now, the way we will sort of think about multi Balan molecules is that from a synthetic polymer standpoint. These really are best taught off as associative polymers. Which in the language of seminar and Rubinstein is basically described as polymers or macro molecules that have attractive groups interspersed along the linear chain. Which will fit this typical sort of stickers and spaces architecture. So I’ll walk you through what I mean by that. So here’s let’s say a linear polymer the sort of bluish colored ovals are intended to represent the stickers. These are the groups that are capable of making attractive interactions. They could be motifs in intrinsically disordered regions that happened to have let’s say aromatic groups or charged residues or specific types of polar amino acids. But the key distinction between a sticker and a spacer is as follows. In terms of attractive interactions, the stickers will always be stronger in terms of sticker sticker interactions. So anything that sort of is inferior in terms of attractive interactions. To the sticker will end up being classified as a spacer, so much so that the spacer can actually sort of in code predominantly repulsive interactions, because it really likes the solvent Or it could have sort of week attractions when compared to the actual stickers themselves and This is a much more relevant a framework for thinking about the driving forces for face separation and it turns out it’s not just for intrinsically disordered proteins. But it actually can be adapted to, you know, folded molecules folded molecules connected by disordered lingers. And as I will demonstrate even for RNA molecules as well.
So the obvious question that arises is, If we are to let’s say leverage the power of this particular theoretical formalism, then one of the things that we’d like to be able to do is to look at a sequence and figure out What sort of stickers and spaces are encoded within the sequence or what sort of stickers and spaces might become emergent as a consequence of post translational modifications or legalization or things of that nature. Now in systems such as this popularized by Mike Rosen identifying the stickers in the Space Service actually becomes fairly straightforward. Because what typically one has our interaction sites on a folder domain. So in this in this Pac Man representation of the SH three domains, essentially the mouth of the Pac Man. Is intended to denote the binding site for the protein rich module. So effectively together these makeup complimentary stickers and as Mike showed in this beautiful paper now roughly eight years ago is that The valence of the SH three domains, ie the number and also the PRS will basically determine the driving forces for phase separation. And subsequent to that work. Work that Tyler Harmon did in my lab demonstrated that specific properties of the blinkers, which will service phasers will actually determine whether what you get are sort of spherical convent sets fully networked or whether you just make sort of Test Tube spanning gels. Right. But so in this case there is no fundamental challenge to sort of identifying stickers and spacers.
The particular challenge that might be in play is working out the properties of the spaces and or whether there are auxiliary stickers Resident inside of the disordered lingers but of interest to many of us is the problem of low complexity intrinsically disordered regions. And here is one archetypal sequence that I show you drawn from hnRNP A1. These, these molecules that are of interest to many of us have this typical architecture of, you know, multi valence of RNA recognition modules are our EMS Also in some places, depending on what your biases. They’re also known as RNA finding domains and often they have these low complexity domains. There upended, or interspersed between. Now, one thing we know for systems like this is that you can lop off the LCD and it can actually drive phase separation on its own so it either is the main driver of phase separation or it’s modulating the overall driving force for face separation. But the key question then becomes, if you look at a sequence like this that is sort of has basically a very parsimonious alphabet. What are the stickers. I mean, I have actually colored them to leave you with some breadcrumbs as clues and notice my cursor is hovering around what I think are the stickers But you know, that’s come from sort of. So, what we want is not to just use, you know, human intuition, but sort of arrive at some rigorous ways of identifying stickers. So what we’ll do is basically ask the following question, which is What are the stickers and archetype of low complexity domains. And I think based on sort of the extent theories, what we can propose is that mutations to stickers will directly altered the driving forces for face separation in sort of measurably significant ways Whereas mutations to space, sirs, may not necessarily impact the driving forces such as the saturation concentrations, which we will talk about, but they will Impact the opportunity of phase transitions and also the material properties of content sets. Yeah. So in the prototypical protein, where he had already biting domains and low complexity domains and then you got to bring in later on RNA is on it in condensates do all the RNA protein interactions happen between just the RNA binding domain and the RNA or some of the stickers and spaces in the low quality domain also interacting with the RNA. I think that is Really the burning question in the field, right, because if you work to use Simple measures like, you know, and these are not very easy to perform at RNA, but affinity measurements and very dilute solutions you will conclude that, you know, these are fairly weak interacts with typical RNA molecules. It turns out that, of course, you can also demonstrate that these low complexity domains and Geraldine say do has some very nice data and they’re really Parker has data. And what you can demonstrate that these low complexity domains definitely house, the ability to interact with Darren.
In fact, What I’m going to focus on is actually when I get to the RNA portion purely a low complexity domain. Right. And so we’ll, we’ll try to get that that grammar as sort of the first step.
So The story of trying to identify stickers in a sort of systematic way started with the collaboration with our colleagues in Dresden. Tony Hyman and zoom and I’ll guarantee driven almost exclusively by to long a postdoc in Tony’s lab and jungle che who was a postdoc in my lab and the the targets were, you know, essentially what we refer to as fast family proteins, but these are really fat family proteins. And all of these proteins have this very interesting by part tight architecture. So there’s this pre unlike domain over here in the end terminus. There are the RNA binding domains that basically encompass Arginine rich ID ours shown where my cursor is hovering and then bona fide a folded domains that are in a recognition modules lot of these proteins have this architecture. And one of the things that you would do sort of As a convention, alas, A is basically to ask as you crank up for some fixed solution conditions as you crank up the protein concentration At what is the threshold concentration, above which you see the onset of Droplets or condensates, and you can also do the spectrum automatically. So it turns out that you can measure the saturation concentrations two or three different ways. He actually had four separate ways of doing this. They all generate reproducible results and the number here is circa five micro molar for full length us What’s interesting is that GM went on to then measure the saturation concentrations at 150 millimolar potassium chloride for a series of different proteins that have very similar architectures. And the blue bars basically make the point that these Sequences clearly have sequence specific driving forces for face separation, as evidenced by the fact that the saturation. Concentrations Vary by up to two orders of magnitude. Just as a reference the here are the typical sort of average concentrations of these proteins inside cells. I put this up, mainly to make the point that you know there is the cell is also sort of perhaps cares a bit about you know the expression levels of these these proteins.
But having said that, These are sort of crude numbers that you know are kind of context, independent, but the key question is the following, which is What are the determinants of these sequences specific saturation concentration values. So when we were thinking about sort of what might be the underlying molecular grammar Alex Holehouse. Than a postdoc in the live and john Lowe sort of thought to interrogate the proteome of intrinsically disordered regions, in particular, and the entire human protein naturally And noticed a very distinctive signature that popped up which is that these proteins that we had sort of focused our attention on sort of listed here. We’re quite This to doubt for having a combination of a high tire scene and high Argentine content. Compared to this red blob and the bottom left corner here which is your garden variety protein really tends not to have a very high are Janine are very high tired of seeing content in terms of ID ours. So the thinking was thatEven though these sequences are highly dissimilar to one another from an alignment perspective. Perhaps this high frequency of piracy and arginine, which is uncommon might have something to do with the driving forces. So that led to a series of experiments. 65 Where in to went on to basically delete the pre on like domain under conditions where one can observe condensate formation for the full length one basically does not observe contents information for the pre unlike domain. Or the RNA binding domain on its own. But if you put them together in trans and solution, you actually can sort of get back a saturation concentration, that’s Definitely closer to a higher order of magnitude higher than when they’re in sort of tethered to one another. And that’s essentially an effective concentration argument that you can make up. So in the transfer You all to the concentrations of those to You okay, you know what I mean, like, see. Yeah. So, so those will definitely lead to sort of this closed loop type of phase diagram which would they look exactly like what would observe and Mike Rosen’s P Polly PRN Polly SH three and it and The concentration regimes, where you will define the phase boundary will be determined entirely by the valence of the arginine and and in fact We have that prediction of what that phase dining room would look like. And for candy bars recent paper in biology. So It would appear that these are stickers. So you can sort of fine grained this a little more and think about the amino acid chemistry. So, I mean, You look at this and you say, oh, well, you know, we know all know about cacti and pie interactions. So that’s what must be what’s important. So we can drill down a little more. So we’ve got a pi system here a pie ish system. So this is sort of a why aromatic system if you want to call it that, essentially a planar arrangement of charge. John Lowe basically then said, well, okay, we can adapt the sort of published stickers and spaces model from seminar and Rubenstein And then come up with a framework that basically generates a prediction for how the saturation concentration should depend on the valence of this numbers of stickers and spaces and, you know, Some sophisticated mathematics later actually what you end up with a very compact formula that says that the saturation concentration Should be essentially inversely correlated to the product of the numbers of Tyra scenes and the numbers of Argentines and the correlation. When this is a fairly crude theory. Because it sort of ignore some very specific spacer effects and things like that. And the correlation is actually pretty good. So it suggests that ok so the primary stickers in the sequences. Are probably in fact the tire scenes and Argentines together, um, Now we started to think a bit more about these stickers. What would make a good sticker right so clearly Argentines have charge. They have a D localized charge distribution. So that would lead to a Sort of a dipole moment just by, sort of, you know, the way the electron cloud would be distributed, but they also have a planar arrangement as to the pie systems and that could give you sort of these big quadruple moments. The reason that becomes important, is that a sticker than can encode essentially a hierarchy of interactions. Right. So charge charge interactions will be long range goal is one over our charge dipole will go as one over our square charge quadruple as one over r cubed and so on. So, what we were building up to Was actually a prediction that’s based on these intrinsic multiple moments that you can measure the gas face for the systems. So Tyra scene. Has a finite dipole moment just to calibrate you the dipole moment of water is about 2.6 Dubai so you know it’s essentially Very much water like in that regard. It’s got a substantial quadruple moment fennel alanine, because it doesn’t have the O H group basically has zero dipole moment and roughly the same quadruple moment as as Tyra seen Our Janine has a charge, just as lysine would but the the charge distribution is essentially I saw tropic For the amine, suggesting that and in contrast to the charge distribution for the Guan ido group which has this sort of planar arrangement and the localization. Giving it a substantial quadruple moment but so the key hypothesis that emerged was that if you were to substitute Tyrus scenes with tunnel alanine or Argentines with life scenes you make for inferior stickers right and so Let’s test that. So in fact here is, let’s say the intrinsic saturation concentration for full length us in 75 milli molar potassium chloride. You make the tire substitute all of the tire scenes defend Allah. Allah means you clearly increase the saturation concentration. And in fact, that seems to be can coordinate with the Sort of decreased sticker strength you make substitutions of the argentine’s to license. Again, you get sort of a weakening of the driving forces as seen by the increased saturation concentration And you get a mildly non additive effect when you substitute all of the tire scenes to funnel alanine and all of the Arginines delay scenes. This becomes fully additive and you account for the fact that there are some electrostatic differences. So a time and $50 million debt is perfectly additive So essentially what this says is that in intrinsically disordered regions which we tend to think of as not encoding any obvious specificity. Right, because it doesn’t have well defined structure. So, therefore, you shouldn’t get You know well defined specificity there indeed are these stickers and spaces and the stickers are delineate a ball because the encode this hierarchy of interaction ranges and interactions strengths And this, by the way, is these are findings that are entirely resonant with, you know, Results that Julie Forman-Kay has has identified and even converted into sort of a bio informatics predictor all that they don’t use the particular languages to persons pacers Um, so now you know to sort of really get at this interplay between stickers and spacers As opposed to just sort of predicting the effects of stickers you really need to get at sort of this using simulations. We’ve developed to simulation engines. I will talk about one of them. So let’s see, is actually this is published.
Now, Is essentially a way of sort of instantiate in protein architectures onto a lattice. And then you can sort of you know do simulations of full blown phase behavior. What I’ll do is talk about an unpublished piece of Work, which is based on essentially and and an engine that Alex whole house developed whereby each ID arc and essentially be written out as a single beat per lattice. And you can either learned the interactions between I DRS abuse or you can sort of come up with phenomenal logical models and I’ll sort of demonstrate this using a particular set of examples so First is that you know there are some standard things we want to be able to do, which is it’s all well and good to be able to predict Saturation concentrations. Once you know the identity of the stickers. What we want to be able to do is actually predict stickers de novo. And also calculate phase diagrams and so I’ll walk you through what a calculation of a FaceTime from looks like. Here. What I mean by this is, is coexistence curves. And typically what will happen is you start up a series of simulations at some particular protein concentration defined in terms of volume fraction, if this concentration. It happens to live in the inside of the two phase regime, what you should get in your Appropriately converged simulation is the formation of to co existing phases. The timeline in this case will be horizontal, because the temperature in the dilute and the dense phase should be exactly the same. But if the order network salt concentration and you let’s say got Preferential accumulation or exclusion of certain ions or let’s say small molecules, then what you will actually get our timelines that have slopes to them, right. Which is actually going to be really important for thinking about, you know, how small molecules interact with condensates, for example, um, There is this well known cemetery. For example, in homo polymers, whereby in dilute phases. The chain will compacted on itself. Basically indicative of the fact that it doesn’t like the interactions with the surrounding solvent But when you crank up the concentration it swaps out those in trauma molecular interactions for entire molecular interactions. Such that the concentration of the dense phase would essentially be the same as the concentration of these beads inside of the block you. We can actually calculate the analytics for homework. And you can show the pimps can reproduce. That is basically what that is saying that, as you can see, you essentially the radius of generation will increase as we increase temperature And then concordance with that in the regime where it is, you know, not expand that you start to see this two phase behavior, depending on the protein concentration So, but, of course, none of these low complexity domains are homo polymers. They actually have stickers and spaces intersperse so here’s kind of a generic question. That we set out to answer in collaboration with my colleague and good friend, Tanya me time from St. Jude. This is the two postdocs who sort of CO drove this with with Alex or Erik Martin and Ivan parent Who worked very hard on sort of some really elegant set of experiments. And so we went to h&r the the low complexity domain of h&r NPA one I showed the sequence earlier but zero in on the LCD, which as a way of reminding you, is sort of the same as the pre on like domain from an compositional bias standpoint. So let’s start by interrogating just a single chain behavior. So now orient you, in terms of what to expect. If so, what we’ve established effectively as a byproduct of this work is a well defined protocol that you can use to identify stickers and ID ours.
The reason this becomes really relevant is that those stickers are also possible targets that you can sort of manipulate using small molecules, etc, etc. So If I do a simulation of a homo polymer the stickers will be drawn toward one another. So I should pretty much be able to see that in a movie. In fact, I won’t Unfortunately, the title of the slide gives it away, but I could have easily sort of not have put the title up and you would have seen that, you know, essentially what’s happening. Is that the chain is making some transients sort of secondary structure largely sampling an assortment of confirmations. But any compaction of the chain is largely driven by these aromatic stickers or aromatic recipes that are interspersed along the chain. Despite the ways to be published in the next couple of weeks, um, The other thing you can do is analyze things like the radius of generation in terms of, you know, conventional polymer scaling theories were n will be the number of residues, there will be a scaling exponent which, by the way, you can extract from Sort of looking at radius of gyrations distributions, for example, and here is simply a calibration. So if I were to make all of the interactions be repulsive. So meaning. The only thing that I have here static exclusion and what I get is the green distribution. If I say that there there are repercussions and attractions that perfectly counter balance one another. I get the so called Gaussian chain, which is the black distribution. So what we have for the full blown interaction model is the pink curve, which basically says that, yeah, the chain is trying to be well salivated but what’s happening is that the sticker interactions are trying to compact the chain on itself. You can go and do small angle X ray scattering measurements and by you go and do Eric Martin goes up to are gone and the advanced photon source. Is become quite the master doing this. What is a particularly useful technological innovation that allows us to do these experiments on Intrinsically disorder domains is the coupling of size exclusion chromatography to the sacks beam line because historically, this was a real challenge with sacks measurements, because you would always be confounded by the aggregation pro nature of these molecules sacks is a Sample greedy technique needs high concentrations and and high concentrations are always fighting the problems with aggregation. But if you have a size exclusion column you can effectively elude out Predominantly modern American species. And in fact, you can even detect the presence of any kind of legal memorization anomalies in your sample. So here is a typical Scattering form factor, shown here, and black dots Erik is masterfully careful about signal to noise issues. And so, you know, effectively, he’s got several independent measurements. Sometimes he gets even more sort of persnickety and does measurements in two separate beam lines. Just to be absolutely careful and then from the simulated ensembles you can actually calculate the form factor know fitting parameters involved and you pretty much overlay on top. The SAX measurements you get pretty much sort of can coordinate values for the radius gyrations What you can do is analyze the SAX data using what is called the molecular form factor that Josh reback Developed when he was in Tobin sauce next lab. A couple of years ago, which gets at this sort of a parent scaling exponent. And when you fit the data you can extract the scaling exponent, essentially, you get something that looks like this. So just as a calibration. If the chain were to make a perfect globule this value would be one third. If it were classic self avoiding walk it would be three fifths. For a Gaussian chain where you exactly counterbalance the repetitions and attractions this number would be point five. So it’s clearly in this sort of crossover between the globulin the Gaussian The other thing you can do is go back and ask, well, hey, maybe you know these stickers are encoding some secondary structure. For example, so classic and Mr spectroscopy will allow you to do that here is sort of a proton. Nitrogen ages QC spectrum and this stands out, mainly because you get this very poor dispersion along the proton access and very sharp peaks completely can coordinate with the idea that these are intrinsically disordered regions, you can deploy. Julie Forman-Kay’s SSP score profile that analyzes these types of data and turn them into secondary structure propensities and what you find is a very, very weak bias, all the way through. This is pretty much in the noise. You can go back and analyze the simulation results and you get roughly similar types of patterns quantitatively. They’re not exactly the same. But again, there’s to go from here to a calculation. You have to actually convert to secondary structure sort of predictors based on Assignments and so then there’s always some mismatch there. So then, having established that the simulations and the experiments are roughly on the same page, you can actually analyze sort of a normalized sort of pattern of inter rescue distances. So when you see blue. The idea is that you effectively have sort of an expansion when compared to a self avoiding walk. But if you start to see lots of red blobs, essentially what you’re saying is that the chain is sort of more compacted in Those regions and those invariably involved. These aromatic recipes And so that led to the idea that okay the prediction based on the single chain studies would be that the aromatic residues in the pre on like domains would be the stickers If so, a zero dollar prediction would be that a tight ration of the valence I eat the number of those aromatic stickers should have a direct impact on let’s say chain compaction. So we designed three variants, we call them era minus era, minus, minus, and ERROR. PLUS effectively that sort of And there’s a reason why we designed them the way we did and that will become very clear in about five minutes time Um, so effectively are tight trading the valence here and the prediction would have been that visa v. The wild type era minus should become more expanded era, minus, minus, should be even more expanded our applause should become more compact.
Okay so beautiful systematic trend. These are from simulations. And by the way, those are corroborated by experiments as well. I’ll get to that in just a second. So then this led us to, well, okay, it looks like the aromatic recipes or the stickers Let’s just come up with a super coarse grained model using PIMMS where now what we do is we add color to our homework polymer right so what we’ll do. : Is wherever we have an aromatic rested, you will turn that into an orange bead that basically is a sticker. And then in between. We’re just going to have blue beads, which are going to be spacers We parameter is the interactions between the stickers. Stickers and spacer spaces and phasers to be such that, and so this by surveys in units of thermal energy 12 kT, this is about three kT, this is about one kT And you essentially parameters, this model to make sure that you get back sort of the radius of generation converting from lattices to offer lattice that sort of correlates well with the experimental value. And that’s the only parameterization involved.
Now we do simulations. So at this particular concentration in milli molar. So effectively what we see are, you know, sort of dispersed phases. You see very few molecules running around in this gigantic simulation volume. Occasionally you’ll see them running around, I can play this again. At that point, you essentially see the separation into two coexisting phases. You see the dilute phase concentration would be about, you know, slightly less than point oh one milli molar, and the dense phase coexisting concentration You can actually fit this whole the points here come from the simulation getting at the critical point is a real bit of a challenge in these simulations, because The fluctuation become enormous and so we just switched to a flurry hugging style mean field model to be able to predict the critical temperature. That’s the two body that’s the three body interaction. And that’s where this sort of curve is coming from. Right. Okay, let’s go to measurements on the one LCD. So in blue dots are the measurements from Tanya’s lab that I’ve done, and Eric and brima actually performed These are really painful experiments right and so Sorry in blue dots are the stickers and similar space or simulations in black triangles are the actual measurements. So, this particular measurement uses this manner drop methodology where effectively what you do. Is you spin down the convent sets you spit out our pipe it out is the technical term on the The stuff that is in the pilot and then you dissolve it in urea. So then you essentially sort of your dissolving the convent set and then use the specter of automatic as a to actually measure these numbers. Very easy to describe. They are you need a truckload of material that just takes is good. The requirements are quite problematic. So given the intrinsic noise in these measurements we reached out to a colleague country or Serrano Who’s at Wash U and he used FCS and this actually is the first demonstration that you can in certain types of contents. It’s used for essence correlation spectroscopy as a way to Measure the coexisting dilute and dense phase concentrations. And actually, I won’t go into this these data also beautifully illustrate that the chain in the condensate, at least for these types of sequences essentially is freely diffusing Is modern American so effectively it’s effectively diffusers like it’s in a very, very viscous medium. Right. So having then fit to the Florida Huggins we could Estimate the critical temperature and then you can do a cloud point measurements essentially sit that this concentration go up and temperature go down and temperature and ask what is the concentration Or temperature at which sorry you the system becomes cloudy versus clear and that becomes the estimate that the critical temperature. So this in effect becomes one of the first sort of almost fully measured by notables where you also get at the critical point. Now let’s go and sort of think about our simulation design. So again, this is showing you a movie at that particular concentration, you start to see sort of, you know, the binomial has squished in and it has gotten shorter because we have cranked down the valence of the stickers You go down to sort of this era, minus, minus, or IRA to and effectively you know at this temperature. It’s basically a fully dispersed system. And so now we can turn to experiments. Here are the facts measurements that actually make the point that indeed these the chain becomes less compact or becomes more expanded as you crank down the valence becomes more expanded as you crank up the valence Here are now all of the vinyls right so the circles are the points from the stickers in space or simulations. The triangles are from the experiment. You can see that this particular sequence is inaccessible experiment, unless you find a way to go into the supercooled regime. And the Fifth are the the solid curves are not joining the dots there actually fits to the Florida Huggins theory and you can clearly see you tight rate the valence you tight trade the driving force for confidence information right technology is probably quite real But The thing that sort of baffled us originally was Decent hetero polymers and we spend all our time bellyaching that, you know, homo polymer theory really shouldn’t describe you know header upon America systems and yet they do. So the thing that sort of on a complete lark, I sort of wondered if The uniform distribution of the aromatic group. So along the linear chain was sort of responsible for this and you know when you have
A collaborator like Alex, that becomes very easy to sort of muse something and then an hour later, you have an answer right and so Alex then said, Well, okay. You go get yourself a coffee. I’ll think about this. And so, what he said was that, oh, he can come up with the sort of binary patterning parameter sort of been inspired by things that we’ve done in the past. But he basically asked the following question. If you treat each of the aromatic rest of us as the stickers. Everything else is a spacer. You can basically compute A parameter omega, that is going to be one. If all of the aromatic groups are clustered together in the linear sequence. Or approaches zero. If you essentially disperse them along the linear sequence, this becomes a meaningful number if and only if you have At least 20% or 15% of your residues in the sequence being stickers. This is turning out to be quite robust, by the way. And so then, so this is a number. What does it actually mean. And this is where you know if you have Alex’s bioinformatics skill to actually sort of are able to make sense of it. But then what you can do is generate a gigantic library of random sequences and ask the following question. From an evolutionary perspective, this patterning makes sense if This is something that you would not stumble upon at random. Right, so if any garden variety sequence that you pick from a bag that has this composition has this pattern and it’s not terribly meaningful. And in fact, what it fine. What we find is that well over 99.99% of the sequences that you would generate a random would never have this type of well mixed patterns. So it starts to lead down as down the idea that maybe this is an evolutionary fingerprint. So if you have a design principal, you can come up with a Query to ask if that design principal has has legs so does the patterning of aromatic stickers matter. And so we came up with. Two types of variance. So we’re keeping the amino acid composition exactly the same, which means the valence the intrinsic valence of stickers is exactly the same. We have one shuffled the variant where we lowered the omega and another we essentially collect them up. And in fact, we actually designed even more aggressive variant. Those are just impossible to even get : Expressed in cells right and so experimental constraints, sort of, we actually from a competition perspective have like 100 different variants of this And so here are just to orient view. Here’s arrow. Perfect. Here’s wild type. Here’s arrow patchy and this is simply showing you sort of the patterning here. So when you do simulations, you get beautiful droplets with our perfect. Same with the wild type. With the patchy variants, effectively, you start to make these very my seller looking structures. That essentially are a computer manifestation of something that’s just going to fall out a solution. And make precipitates right because essentially when you make these myself looking systems. That’s a manifestation of what you would refer to as micro face separation. And if you just essentially what you end up with are making solid like species and the saturation concentrations, but they’re not microphone sorry at four in the morning when you do this, I was writing micro face separation of microphones will fix that. And so when you do experiments you see exactly that’s right i mean the the thresholds. The solubility limit just goes to the floor. These molecules just fall out a solution and become amorphous precipitates and, you know, we can reproduce this with very number of sequences. So essentially, the next question you go back and ask is, well, okay, you can design these things. But there are lots of sequences that have These pre and like domains and shortly enough in all of them. They’re the patterning of aromatic residues is uniformly distributed right So we, we were able to find a lot of well known, guys. And then we’d found this one sort of as a prediction and then like, you know, a week later, we saw this work from Jennifer lip and contorts talking about these Acts seven proteins that essentially are forming condensates that basically help walk life those arms along axons right and so These are, by the way, involved in the secular trafficking that synopsis and turns out that mutations in this that screw up the patterning basically create all kinds of Uber and short term memory problems right so they absolutely screw up there is actually there are other here predictions. Which means that in all of these cases. If you screw up the patterning. You’re going to start seeing interesting effects. And a really beautiful outlier is actually work that comes from Alpha Boca from who when she was working with to Mitchison and and Tony Hyman She been studying X below this protein that makes that’s essentially a scaffold for bauby antibodies And it turns out that the prion like domain of X below has a very strongly clustered patterning of the aromatic residues and Nothing that they could do could turn this into a liquid, it always makes these amorphous solid like Bodies right and with. So in fact I’m Alex and I are actually working to sort of test this hypothesis right now so There are two messages in this part of the talk, which basically says, turns out that in IB RS Amino just individual amino acids can service stickers their valence is clearly of central importance than And the patterning of these stickers will essentially altered the intrinsic sticker strength so as to sort of essentially control this interplay between phase separation and precipitation and the obvious place to start thinking about this, as you say, Well, okay. In the context of certain types of diseases associated mutations are you effectively cranking up the patterning of the effective valence through this Linear patterning. Right. I mean, you may not necessarily have mutations that are just increasing the aromatic content, but you could sort of have emergence stickers In the context of the convent said that along the lines of what I’ve been anxious shown over de about 40 years ago could actually lead to sort of sprouting out the solid like species.
And the last bit. What I will do is sort of make this point that you know if you increase the valence of the stickers, you’re basically pushing yourself. To word, sort of, you know, sticker driven precipitation or aggregation. If you crank up the valence of spacers then effectively you’re cranking up you know essentially just Network formation without condensates, and so I think the space of condensates is actually in a sweet spot, right, that sort of optimize this Three things the valence of stickers, the patterning of stickers and the properties of sponsors. Right. And so I think to say that I just pick a random ID are out of the hat and assume that it’s going to make condensates is is is a bit of a misimpression that I think has unfortunately taken hold in the literature and show we’re not in a place where I think we can start to sort of provide some guidance to how to think about these at least these ideas. So the obvious question is, of course, you know, there are a truckload of all of these proteins all of these low complexity domains invariably are excised from RNA binding. And we zero in on the RNA recognition modules as the players for army binding, but the obvious question is, what would the local taxi domains do. And so that led us to a collaboration actually with Stephen Barnum’s And I have to actually point out that I think throughout the collaboration. I don’t think we had one word with Aaron. It was entirely Stephen Alex and myself sort of going back and forth with one another and the Naira actually did some beautiful work, which I’ll talk about briefly as well. And the system, we decided to sort of interrogate where an archetypal low complexity. The main sort of the sea nine or 797 sort of, you know, die peptide repeat system that makes this really important and relevant And we thought let’s take a simple Homer Paul America RNA. And I should point out here that, you know, whenever we protein Allah just basically say RNA molecules we add RNA. I was at a beat biophysical Society meeting last year and somebody walked up to me and said, I don’t say when I’m studying RNA molecules I add protein. I tell you which protein. I’m adding. So RNA molecules also have just the same degree of sophistication. So what the hell do you mean you know you add RNA. Right. And so I’m learning. So, so, of course, now we’re talking about sort of a mutuality right so effectively in the ancient literature now dating back 115 years There’s this face separation used to be known as complex conservation because you could bring together two opposite really charged molecules. And if they could find a way to neutralize their charge But realize multivalent interactions, what you would get is essentially about some threshold concentration. A coalface separation due to the complexity of the Complimentary ions, not in in sort of binary interactions, but in sort of, you know, a network of interactions, giving rise to face separation, but complex classification And indeed, in fact, you can tell I’m learning RNA ology side. I’m Tim lohman has taught me that To do the way you distinguish RNA from DNA is to make sure you have the little are so right bows versus DLC rivals. So, you know, I’m learning. Tim has also taught me that you always list all of the solution conditions because you have incredible Dependence as Tom record has taught us over the years on solution conditions. So, under these conditions, you basically get these PR 30 species making spherical common sets with at name you get these very irregular architectures with The guanine the poly guanine sequences. Right. 2 So this is a beautiful methodology which is soft extra tomography. It’s taking advantage of the fact that essentially They’re the absorption of X rays. Actually, the transmission rather of X rays is different. Soft X rays is different for water been compared to carbon and oxygen. And that differential transmission can be used to actually construct a full blown image reconstruct the full blown image. And this is done. Pretty much on a regular basis at Lawrence Livermore, and so veneta exact was. It was a staff scientists there and you can actually see the spherical condensates being formed. I’ll just sort of play this again. And you see this very irregular morphology is with the guanine tracks. What is so special about guanine. So it turns out that one of the things that you can get through the apology sequences are these G quadruple axes. Right. And so it would appear that the ability of the RNA. To molecules to make specific types of stable structures can lead to different types of morphology, but it leads to also very important question, which is the impact of RNA structure on condensate morphology. So effectively, if you added this PR 32 a mixture of non base pairing RNAs, you always get The spherical content sets. And so essentially here, what we’re doing is tight trading the ratio of yourself to cytosine. And effectively, what you see is pretty much across the entire range of ratios. What you get a card spherical condensates Same thing with admin inside of seeing essentially non based pairing and you get back spherical conferences. So that’s all good. But once you start looking at base pairing Condon sets this directly points to the possibility that structure is somehow impacting The nature of the content sets that you form in the limits you get back basically spherical condensates but then when you start looking at sort of ratio metric mixtures, you start to see these irregular morphology is forming And the obvious question arises is a structure somehow fundamentally altering face separation and giving you irregular morphology. So you get structured content sets. And so this dates back to sort of ideas in the 90s that came out of indifference Francesco Cirillo Dino and gene Stanley, making the point that If you had molecules that have strong cross linking ability strong base pairing abilities. You can connect typically arrest face separation, right, because they’re so busy making these structural interactions or networking interactions. That you can sort of essentially get gel like states as kinetic traps. And so one way to ask the question, was to ask whether The lack of spherical morphology was essentially kinetic traps and this is not to say that these won’t be long live these could be eternal right But you can if they’re kinetic traps.
What you can do is give them thermal kicks and try to see if you in your system will kneel back into spherical condensates, and with a simulation engine in hand, this is what Alex actually did. So here are basically you know PR 30 Polly aren’t a con to diffuse fairly slowly, it turns out, and this by the way we could potentially have experimentally as well. But when you start Essentially, adding base bearing abilities you actually get these connected arrested things essentially they’re frozen in right and then But then if you essentially sort of do a thermal kick you basically get back to the spherical concepts and that, by the way, is exactly what you get even experimentally. So if you do subject. These Arrested content sets to some level of heating and it’s really close to boiling. In this case, and then you kneel THEM BACK, YOU GET BACK, YOU KNOW, ESSENTIALLY spherical looking contents. So this filaments network. May well serve as RNA mediated kinetic traps, which is something to be thinking about. So as we think about the impact of RNA structure versus long non-coding RNAs. We get to start realizing that there are three effects of different types of RNA dirty structured RNA that could be sort of valence limiting or valence altering So effectively what needs to happen is the synergistic change in the RNA confirmation that changes the valence that allows the kneeling back to spherical content sets. Simone and teach us semen Alberta indigenous friends men have actually made this observation with Jordan as well that aren’t entanglement actually can be a very important player in terms of sort of arresting condensate and kneeling Another yeah yeah I just missed. Is that Holly RA and are you mix together or is that compositionally 60 and 40% compositionally 16 for games. Yes, thanks. Yeah. So, so The other thing that, of course, comes up is, you know, now we have stickers in space or so we can start to think about, you know, the sort of the nuclear base versus the amino acid, which is of course the captain. And just as a way to orient ourselves. We’re now going to have the periods and the pyramid deans and of course we’re not thinking of timing. We’re thinking if you want to sell in the context of the pyramids. So you can make Condon sets with periods and you start to have these fusion dynamics really being sluggish when compared to the pyramid Dean’s right And in fact, I’ll quantify this in terms of these inverse capillary velocities. So you actually see lower inverse capital learning philosophies indicative of sort of considerably more fluid, the droplets that also fuse quite readily when you have pyramid Dean’s as opposed to purines. So effectively, the poly Puritans, you know, are slowing both fusion and dynamics. Now you can go back and ask, well, of course, we get to choose the time so Argentineans versus life scenes. 2And this builds on are actually adds to rather a recent story that we have contributed to not just from the electronic grammar work, but also in terms of the preference of the nucleus versus speckles for Argentine rich versus lysine rich protein. So there’s incredible specificity there. So here for example is with polypurines, you can clearly see sort of two orders at least two orders of magnitude difference in the inverse capillary velocities, these are essentially quantifying the dynamics of Droplet fusion. So you change all the arginines to license kept on his cat Diane, but the nature of the captain really matters. You basically altered the fusion dynamics. That persists even when you go to changing the periods to pyramid Dean’s And of course it also depends on what type of nuclear base we have, I mean, you can see that there are some the actual differences changed quantitatively, there are differences. Here are internal dynamics basically measuring the recovery dynamics off PR 30 in the context of these different types of condensates You can actually see the yellow, the yellow is under you can kind of see that sort of poking out it’s effectively underneath the red. So the Puritans are Fundamentally different and feelings down here are fundamentally different from the pyrimidines That difference in the internal dynamics effectively is almost abrogated when you change the Arginines to the license right so clearly what this is starting to say is that There is also an intrinsic valence difference in the way these cations are pointing Act. So the way I like to think about Arginine is really having this forked tongue right and this, why are an activity is enabling these sort of identity or multi talented interactions. But one of the things that we were really interested in, of course, is that every RFP grand new all is of course a multi component system. In fact, you have multiple types of RNA molecules. And so if you take multiple components. What, of course, you can have is if I have n polymers, plus a solvent and I fixed the temperature and pressure. I can get n plus one coexisting phases, what does, what does coexisting phases look like Here would be let’s have two polymers p one, p two and a solvent a homogeneous mixture would just be sort of monochromatic A Condon set that is enriched in one protein that coexists with the dispersed phase. That’s basically enriched with the solvent and the other polymer would look something like this. You can flip it, of course. You can get a condensate that’s enriched in the two polymer that’s coexisting with the dilute face that’s basically enriched insolvent or deficient in these polymers. You can get like Amy glad filter has shown with wheat, the wheat three system.
You can have, for example, let’s say, two RNAs, they will actually sit in two different content sets for reasons that are slowly starting to become clear. Structure is but one component of this, the valence is a pure ingredients those actually mattered a lot or you can get this wedding behavior, right. The nucleus is an example of this employer speckles are an example of this, I rather suspect that every RFP granular as an example of this type of spatial organization behavior, right. There where you get this in homogeneous distribution. And so this and of course you can flip this around as well. So we decided to basically ask the following question. So you you take mixtures of The Polly. Polly R Us with our PR 30 system and then you basically tied trade the ratios. These are compositionally different so they’re not in the same polymer So these are truly three plus one component system, namely the, the, the, the plus one here is the solvent And here are you actually start to see that you know here. Basically you have essentially a binary mixture. We go now to sort of different types of ternary mixtures. Then you come out again to binary mixture and in these ternary mixtures. You’re actually starting to see Effectively these multi layer droplets, you can actually go and image this using soft x ray tomography. I love this technique come and you can label free right Get this beautiful sort of density organization and the obvious question is, you know, I haven’t told me which one is which layer is which But we asked a simple question of can we be produced the experiments using our pins based simulations. And so in effect. We have here the PR are sort of set to be repulsive for one another. They attract they’re attractive for the admin sequence less attractive for the situs enrich one this is based on our purity in versus pyramid and observations. So Make a prediction. The prediction should be that we get Polly a course policy shells and our protein is effectively defusing freely between the two. I didn’t show you the data, but in the paper we actually make this point that is directly relevant to some recent observations about making measurements of the intrinsic my abilities of proteins. And then, arguing that these might not be content sets because you know the diffuse devotees might be very similar inside versus outside what I failed to mention and show data for but it’s in the paper is that Our protein molecules actually are freely diffusing inside of this condensates And so if you were to do frat measurements only by looking at the labeled protein and photo bleaching the label protein and studying recovery, you would convince yourself. Oh, this is a liquid like condensate. The RNA is pretty much a mobile Right. And there you would say, oh, well, this is an amorphous or solid like common sense. So this is where I think both the client scaffold relationship sort of seems to have a lot of legs and so therefore, in a multi component system. You could have Molecules that are equally mobile across a phase boundary and which is exactly what all those times correctional hubs are right, they’re all massively multi component systems. But you’re basically measuring the few cities have one species out of what n minus one, right. So that’s an important point to keep in mind. So here you reproduce the core style architecture beautifully essentially the core is basically Polly a core That is effectively wetted by upon the sea shell and the protein is pretty much uniformly distributed across slightly non uniformly because, of course, the affinities tilting it toward The palm. The a core. But if we made the interactions equivalence. In other words, we said that RNA doesn’t bring any stickers to this to the to the to the dance, so to speak. Puritans and pyramid Ian’s are created equal. So then we equalize the interactions now. We basically don’t get any spatial work. Right, so overall summary then is that there are decipherable rules for actually low complexity domain and RNA condensates. This is again tied to the multi valence of stickers And an emerging sort of in emerging work. What we’ve demonstrated is that You know, you don’t need a whopping big advantage and interaction affinities it what you really need is this combination of multi valence patterning. And just sufficient differences in the sticker sticker interactions versus sticker spacer start seeing these common sets versus right and so RNA structure will matter. Probably in terms of determining the overall timescales, because even if the thermodynamic ground state is the formation of a nice vertical condensate. All the RNA entanglement and structure formation abilities can essentially arrest these these molecules in, you know, sort of amorphous or filament structures. The Argentine versus lysine composition really matters, it’s turning out that in a lot of our, our EMS and you know folded arms. We are starting to realize that there is this what I call Janice like architecture, because there’s a sort of partitioning of The argentine’s and aromatics to different faces of these folded domains. And there’s an incredible specificity. If you know the valence the surface valence of our genes versus last scenes. : And so clearly. These will also contribute not just to the driving forces to the morphology the dynamics overall dynamics overall radiology and internal dynamics, but also to the spatial organization and so I rather feel like, you know, I think if we can sort of take a battalion of sort of low complexity domains and sort of model RNA A’s and I recognize that there are differences among RNA molecules. And we should be able to work out the underlying rules and then you really zeroed order, you would go back and ask for a particular condensate. You know, effectively, I have some combination of stickers that I’m borrowing from certain types of low complexity sequences certain combination of space search and similarly from the RNA side. I think where RNA is a bit more wimpy in comparison to proteins is that It’s spacer architectures are not going to be as interesting, right, because I think you can modulate continents and properties because you have a richer alphabet with proteins than you do with RNA, but That other than that, I think the protein RNA sort of synergy becomes really, really important. But though the one hope is that I leave you with a message that yes, these are unstructured RNA molecules are intrinsically disorder protein molecules. Entirely target trouble in terms of specific interactions. Just have to sort of find the right lens to identify them.
So anyway, and there And I think I’ve acknowledged people as I’ve gone along and there’s lots of work going on from other people in the lab and other collaborations as well. But I do want to point out that I didn’t call out for cons work quite significantly. But he is pretty much taken on the mother of all challenges of, you know, working out. How to think about multi component content sets and not doing three components. But, you know, hundreds of components and stay tuned because I think these results are going to be really fantastic. So anyway, um, I think we can stop there and take questions.
Mark Murcko
Think we have time for maybe a few more questions. I guess I open it to Dresden if there are questions. Yeah. From the Dresden site.
Edgar Boczek
I was wondering if you consider it to add affects that would increase the dynamics and the systems like helicases that would increase the structures of these RNAs and cuz you in insert this into your dynamics simulations.
Rohit Pappu
Well, absolutely. Yeah. So, so, you know, you can start to think about sort of fluidizers, emulsifiers. I mean, there are a lot of these cool factors. Right. And so I think Or things that sort of bring either. And the way I think about it is that you will have these cool factors. This this grammar kind of helps us in thinking about the following way. The co factors zero order, you can think about as modulators of sticker valence or modulators of sticker strengths Are modulators of the space or excluded volumes and so on and so having a way to think about how Mike my Contents at scaffold components be modified. And what are each of the modifier is doing sort of gives us a framework for thinking about this. And so, absolutely yes. Because the thing that you allude to, of course, is that a lot of these aren’t a binding domains also have RNA Healy cases. Right. And another increment and organizing systems you have DNA helix cases and various other things and so They’re active there and you can also start to think about sort of these energy dependent or energy independent Molecules that are sort of enabling unfolding and things like that. And so the other thing direction in which we’re going as you can start to append chaperones. For example, Right, that sort of then facilitate the unfolding of proteins that are RNA chaperones and so on. So, absolutely, yes. I think the framework is really robust in the regard of being able to sort of interrogate what co factors will actually do.
Question
General yeah simple question, which is, can you buy it by thinking about it this way. Can you say that the RNA is usually playing a scaffolding role and in the states, or is it too early to say or not quite but
Rohit Pappu
oh yeah so I think actually your last comment, pretty much answers the question because. So I think what we are coming to is a place where We’re thinking about or not thinking. We’ve actually made measurements. Where we’ve got RNA concentration and in fact here there’s enough specificity in the RNA that we’re choosing different Darren and molecules fix the protein RNA concentration on one axis protein concentration on the other, what you get is a closed loop shape the phase diagram. That closed loop gets kind of locked off and different regimes, depending on the interplay between the RNA. RNA protein, protein RNA protein and so header of traffic versus public interactions. The domination of the RNA interactions often occurs at low very low protein concentrations and, you know, high Concentrations and and here we’re not speaking about generic come upon America and is actually talking about very specific types of fungal RNAs that we’re working on with me, Glenn Salter so Conversely, and then of course you can have, you know, protein. Protein regimes as well. So the ellipse, then what for Khan has actually realized is that the Yes, you get a closed loop, but the lips kind of has dimples get shaved off in different regimes, the timelines change slow but cetera depending on for a given stoichiometry to what extent is the header typically interact or the header of different contracts is dominating the home with typical patterns. Right. And so what is nice about this is that We are establishing how to read all this, just from looking at the shapes of the ellipse. And then the next thing but and you know you cannot start making buckets and buckets of protein and RNA to sort of generate ellipses for every system, but there are two things that we’re actually trying to solve. And I think we should be fairly close. One is that To be able to reconstruct the full closed loop based on up parsimonious set of measurements and then use the the underlying computational engines to then drive this forward. And the second thing is that, you know, if you have a bona fide a grand new will you probably have on the order of, you know, 700 to 1000 separate components to think None of us are going to actually plan on you know tagging 700 molecule simultaneous I don’t have trouble finding to channel so you know forget seven. So let’s say we’re probably going to track two or three molecules at a time. So you’re working in this low dimensional space, but you actually have a multi dimensional phase diagram that’s impacting what you observe and the movement space. How does the shape of the ellipse. Tell you what all those hidden variables are actually doing this actually turns out to have a perfect mapping to tomography. And so in much the same way that you would use data, you know, sort of low dimensional data to sort of reconstruct high dimensional ellipse sites or Projections of high dimensional websites, um, you can do the exact same thing here as well. And so that’s the other thing that we’re currently working on.
Mark Murcko
So yeah, I think we have to wrap it up. Absolutely presentation of some really seminal work in the field.
Rohit Pappu
Thank you. Thank you.