Quality in, quality out:
How biophysical data powers
AI-driven formulation optimization for next-gen biologics
by Lea Valla — August 13, 2026
Engineered antibody formats (bispecifics, scfv fragments, nanobodies) are among the most exciting therapeutic modalities in biopharma today. But excitement rarely survives first contact with a developability screen. Unlike the monoclonal antibodies that have been refined over decades, these next-generation formats didn't evolve naturally, and their stability profiles can be unpredictable, variable, and deeply molecule-specific.
The result: formulation scientists are being asked to do more, faster, with less historical data to lean on. Bayesian optimization offers a smarter path forward — but it only delivers when the biophysical data feeding it is precise enough to trust.
In a recent NanoTemper webinar, Isabelle Waibel, a PhD candidate in the Arosio Group at ETH Zurich, shared how her team combined rigorous biophysical characterization with a multi-objective Bayesian framework to identify optimal formulations in just 33 experiments. The central lesson? The quality of the data you put in determines everything that comes out.
Beyond Humira: The promise – and the problem – of next-gen antibody formats
For the better part of three decades, monoclonal antibodies (mAbs) like Adalimumab (Humira) have been the backbone of biologics development. We know how to make them, stabilize them, and formulate them. The field has accumulated a rich body of historical data that makes formulation design, while never trivial, at least navigable.
-
Improved therapeutic efficiency: Smaller fragments diffuse faster, achieve better tissue penetration, and are more likely to cross the blood-brain barrier.
-
Multi-specific targeting: Bispecific and multi-specific formats can engage several molecular targets simultaneously, opening entirely new therapeutic approaches.
That foundation is now being tested. Engineered antibody formats including bispecifics, scFv fragments, and multi-domain constructs offer genuine therapeutic advantages that have the field's attention:
-
Lower production costs: Many engineered formats can be produced in prokaryotic or yeast systems rather than the expensive mammalian cell lines required for full-length mAbs.
But these formats come with a critical liability: they didn't evolve. Their structures are engineered, not optimized by natural selection, and the mutations that give them their therapeutic properties often come at a cost to intrinsic stability. The field also lacks the systematic developability datasets that exist for canonical mAbs, making it far harder to predict where problems will emerge, or how to fix them. That gap is exactly what Isabel and her collaborators at Novo Nordisk set out to address.
15 assays, 73 constructs: Here's what the data told us
Melting temperature (Tm) and onset of aggregation upon heating were measured using
Prometheus Panta. Most formats exceeded the 60°C threshold that characterizes most
marketed antibodies. But variation within format families was significant, showing that the
specific sequence matters as much as the format category.
Thermal stability
The diffusion interaction parameter (kD) was also measured by Prometheus Panta and told a
more differentiated story. Small fragments, particularly scFv and scFv-scFv constructs,
showed the most pronounced colloidal stability challenges. Critically, engineered formats
showed far greater within-group variation than canonical mAbs, underscoring how
sensitively these formats respond to individual mutations.
Colloidal stability
At 30°C, most variants remained stable, though scFv-scFv fragments showed some aggregation. At 40°C, full-length mabs outperformed fragments and bispecifics across the board, and significant fragmentation was observed and attributed largely to GS linkers, which can be optimized further.
Long-term stability (12 months)
To make the complexity of 15 assays actionable, the team adopted a "flag" system: each time a molecule exceeded a risk threshold for any measured property, it received a flag. The more flags, the higher the overall developability risk.
The results were instructive. Canonical mAbs accumulated the fewest flags, confirming their relative robustness. But the data also showed that format family is not destiny: some bispecifics and scFv variants performed nearly as well as traditional mAbs, while scFv-scFv
tandems showed the highest risk across the board.
The flag-based risk ranking
Engineered antibody formats carry higher developability risk than traditional mabs, but the specific liabilities vary widely by format and by molecule. That variability is precisely why tailored, data-driven formulation strategies are essential.
The key finding
As a final step, the team tested whether the in silico tool TAP (Therapeutic Antibody Profiler) could predict these developability profiles from sequence alone. It performed well for canonical formats but struggled with the novel constructs, reinforcing that the field needs more experimental data to train and validate in silico tools for engineered modalities.
What are the current challenges with traditional Design of Experiments (DoE)?
Conventional DoE approaches such as central composite design or Box-Behnken design, fix the experimental plan at the outset. You define your variables and levels upfront, run the full matrix, and analyze the results. This works reasonably well for simple, well-understood systems. But for complex biologics in large design spaces, it has two significant weaknesses:
The method cannot learn from early results and adapt. Every experiment is equally uninformed.
-
Rigidity
With a pre-fixed plan, DoE has no mechanism to redirect experiments toward areas that look promising. You can easily converge on a locally good formulation while missing the global optimum entirely.
-
Local optima
With the developability liabilities of their model protein, bocutizumab, clearly mapped, the next question was: how do you fix them, without running hundreds of experiments you can't afford?
The answer Isabel’s team developed: a multi-objective Bayesian optimization framework.
From screening hundreds of formulations to finding the optimum in 33 experiments
How does Bayesian optimization work?
Bayesian optimization is a sequential, adaptive approach to searching complex design spaces. Here's how it works in practice:
The optimization simultaneously improved three properties: thermal stability (Tm), colloidal stability (kD), and resistance to agitation-induced aggregation (remaining monomer after shaking). Osmolality was incorporated as a hard constraint, keeping all generated formulations within the physiologically tolerable range of 100–600 mosm/kg.
The Results
After 33 total experiments (13 for initialization, then four rounds of five) the algorithm identified near-Pareto optimal formulations. That's less than half the experiment count of comparable doe methods, and crucially, it comes with a guarantee of global search coverage that static doe methods cannot offer.
In terms of speed: Tm and kD measurements on Prometheus Panta take under one hour per round for up to 48 samples in parallel. The rate-limiting step in this study was the agitation assay, which requires one week of shaking at 1 mg/mL. With alternative interfacial stress assays, even that bottleneck could be removed, making the full optimization achievable in a single day.
Run a small set of experiments (13 in this study) covering the design space broadly.
-
Initialize
Use the results to train surrogate models: in this case, Gaussian processes, that learn the relationship between formulation variables and stability outcomes.
-
Train
-
Exploitation (75%): Focus on regions of the design space that have already shown promise, guided by the Pareto front of the surrogate models.
-
Exploration (25%): Send experiments into novel, uncharted regions to ensure global coverage and avoid local optima.
-
Suggest
The model recommends the next experiment by balancing two strategies:
Measure the new formulations, retrain the models, and iterate. To make this experimentally practical, the team used the Kriging Believer approach, retraining the model four times with simulated results to generate five new formulation suggestions per round, all of which can be prepared and measured in parallel.
-
Repeat
"Garbage in, garbage out": Why measurement precision makes or breaks the model
Bayesian optimization is only as powerful as the data that feeds it. The surrogate models learn from experimental results, which means measurement noise must be substantially smaller than the effect of the excipients being tested. If your instrument can't reliably distinguish a 0.5°C shift in Tm, neither can your model.
This is where instrument precision becomes a strategic asset, not just a technical specification. In this study, Prometheus Panta delivered:
Confidence intervals were wider, as expected for interaction parameter measurements, but remained within acceptable range. The most reliable results were observed at clearly high or clearly low kd values, where the DLS curve slope is
most distinct.
-
kD precision
An average 95% confidence interval below 0.1°C, one of the tightest reported for nanoDSFTM-based thermal unfolding measurements.
-
Tm precision
“What was especially useful for me is that I could measure the melting temperature and also the Kd with the same sample in the same instrument. So this saved me a lot of time and also material.”
Isabel Waibel
Scientist, ETH Zürich
This is the practical meaning of "quality in, quality out." The machine learning model doesn't know your instrument. It only knows the numbers you give it. Prometheus Panta's measurement precision is what makes those numbers trustworthy, and that trustworthiness is what makes the entire optimization framework function.
The surprising story the data told about Arginine
Beyond identifying optimal formulations, the Bayesian approach generated a rich dataset that shed light on why the algorithm made the choices it did, and some of the findings challenged conventional formulation wisdom
pH converged at ~6. This reflects a fundamental trade-off: higher pH tends to improve thermal stability (Tm), while lower ph drives better colloidal stability (kD) by moving the molecule away from its isoelectric point and increasing intermolecular repulsion. The algorithm balanced both by consistently converging near pH 6, closely aligned with the pH range (~5.8) seen in most commercially approved mAb formulations.
Sorbitol settled at ~300 mm. Intermediate-to-high sorbitol concentrations consistently improved Tm, a well-characterized effect since the 1980s, explained by the preferential exclusion of polyols from the protein surface, which promotes a more compact and thermodynamically stable protein structure.
Arginine was consistently excluded. This is the headline finding. Arginine is one of the most widely used excipients in biopharmaceutical formulation, and yet, for this molecule, the algorithm avoided it in nearly every optimization round. Across all three target properties, arginine had a negative effect.
The likely mechanisms: at formulation-relevant pH, arginine is positively charged, significantly increasing ionic strength and reducing the repulsive intermolecular interactions that maintain colloidal stability. Its guanidinium group also has known chaotropic properties that can destabilize protein structure, contributing to reduced Tm.
This is precisely the kind of system-specific insight that prior-knowledge-based formulation design misses. Arginine works well for many molecules. For Bocutizumab, it doesn't, and only a data-driven approach that interrogates the actual behavior of the actual molecule would have revealed that.
The broader lesson: Optimal formulation is not generic. It's molecule-specific, and it requires the kind of high-quality, systematic experimental data that biophysical tools like Prometheus Panta are built to generate.
What this means for your formulation program
The ETH Zurich and Novo Nordisk collaboration demonstrates something important: the convergence of biophysical measurement and machine learning isn't a future ambition. It's a working methodology today.
Here are the key takeaways:
-
Engineered antibody formats need formulation strategies as novel as the formats themselves. Historical mab formulation knowledge is a starting point, not a solution.
-
Bayesian optimization finds global formulation optima with a fraction of the experimental effort of traditional DoE, without the risk of converging on a local optimum.
-
High-precision biophysical data is the essential input. Instruments like Prometheus Panta that deliver low measurement noise and multi-parameter output from a single sample are direct enablers of better ML models.
-
Multi-property optimization is achievable. Thermal stability, colloidal stability, and interfacial stability can all be improved simultaneously, and trade-offs can be navigated algorithmically rather than by intuition.
-
The dataset you generate has value beyond the optimization itself. It becomes a mechanistic resource for understanding how individual excipients behave with your specific molecule which is an asset for future programs.
Want to learn more about Isabel’s work: read the publications.