there's a lot more data, a lot more genomic data that's not annotated. Actually, most of it, pretty much in many ways, almost all of it is not annotated. And so being able to learn from an unsupervised manner, hugely desirable
Ensemble comparisons require more than one favorable result
Eric Nguyen describes comparing his team’s work with ensemble methods that combine several approaches. He says the team did not want to select one model and simply claim to be better than it.
they'll take the best methods and kind of do Do an ensemble, right? So they'll take up another, even if the best method is another previous model, they'll mix it with like an SVM and just like throw the kitchen sink at it. And so you can see why it would be the best, right? And so that was the bar for us. We're like, if they're going to throw the kitchen sink at it, like we're not going to cherry pick one model and say we're better than that.
Promising computational results still need wet-lab validation
Eric Nguyen says wet-lab validation was in progress and had not been shown in the discussion. He regards the approach as potentially valuable if it works in the lab.
We were actually in the process of validating the wet lab right now. So we didn't get to show it here, but we wanted to know, right? Actually, can it not just do this in silico, which it can, it showcased that it was able to continue And now we think this is a, you know, obviously, if this works in the lab, we think this is a hugely, hugely valuable paradigm
John Schulman says the field depends heavily on generalization and that it is difficult to predict when generalization, including out-of-distribution generalization, will occur.
The whole field relies a lot on generalization and it’s very hard to predict when you’re going to get generalization, or when you’re going to get some kind of out-of-distribution generalization
Research methods may have substantial room to improve
John Schulman suggests that clever small-scale experiments could support theories that generalize to larger experiments. He expects research methods to be far from their ceiling.
If you think hard enough, you probably could have expected some of these things beforehand. There is probably some very clever way to do a small-scale experiment that’ll let you build the theory that then will generalize to the large-scale experiment. So I would expect that we’re nowhere near the ceiling of how well you can do research
Distillation depends on a realistic prompt distribution
John Schulman says a realistic prompt distribution can help a distilled model match a larger model. He contrasts this with easily verifiable tasks, which may produce strong benchmark results but weaker performance across a broader distribution.
If you have a really good realistic prompt distribution for distillation, you can match the big model really well. But if you only have this distribution of easily verifiable tasks, then you can match the big model on all the benchmarks, but you do worse on this broader distribution