
Ginkgo Datapoints and Apheris Launch Antibody Developability Consortium to Advance AI-Driven Drug Discovery
Ginkgo Datapoints, an offering of Ginkgo Bioworks, and Apheris GmbH have announced the launch of the Antibody Developability Consortium, a new industry collaboration focused on improving the prediction of manufacturability and developability risks during antibody drug discovery. The initiative brings together founding members AbbVie, argenx, Lundbeck and Takeda to create what the organizations describe as the largest standardized antibody developability dataset designed specifically to support machine learning and artificial intelligence applications.
The consortium is intended to address a persistent challenge in biopharmaceutical development: promising antibody candidates can encounter significant problems related to their physical and chemical properties after substantial discovery work has already been completed. Identifying those risks earlier could allow researchers to make more informed candidate-selection decisions, reduce avoidable development efforts and potentially accelerate the progression of viable therapeutic candidates toward clinical testing.
The Antibody Developability Consortium remains open to additional pharmaceutical and biotechnology companies. Its model is based on combining proprietary antibody data from participating organizations with publicly available sequences while using federated data infrastructure to protect confidential information.
Addressing Developability Challenges Earlier
Antibody developability refers to the collection of properties that influence whether an antibody candidate can ultimately be manufactured, formulated and developed successfully as a therapeutic product.
While an antibody may demonstrate desirable biological activity against a particular target, that alone does not guarantee that it will be suitable for development. Candidates can encounter challenges involving stability, aggregation, solubility, viscosity, expression, purification and other biophysical characteristics. Such issues can complicate manufacturing or formulation and may affect the feasibility of advancing a molecule through development.
The ability to identify these characteristics earlier in discovery could help research teams prioritize candidates with a greater likelihood of meeting development requirements.
However, developing predictive models for these properties requires large and diverse datasets. Existing datasets are often fragmented across organizations, generated using different experimental methods or limited in sequence diversity. Even companies with substantial internal datasets may face challenges because their collections are naturally restricted to the antibody sequences generated through their own research programs.
The consortium has been established to address these limitations by bringing together standardized data generation, diverse antibody sequences and AI-based modeling in a collaborative environment.
Building a 10,000-Antibody Dataset
Each founding member of the consortium will contribute proprietary antibody sequences to the initiative. Ginkgo Datapoints will supplement the member-contributed sequences with publicly available antibody sequences to reach a target dataset of 10,000 antibodies.
The resulting dataset is intended to provide greater sequence diversity than an individual company could typically generate from its own internal programs. This diversity is important for machine-learning applications because models trained on narrow datasets can have difficulty generalizing to molecules that differ substantially from those represented in the training data.
Ginkgo Datapoints will oversee the scientific design and execution of the consortium. Its responsibilities include developing the sequence-selection strategy, overseeing antibody production and conducting high-throughput laboratory characterization across core developability endpoints.
The organization will use the resulting experimental data to create a foundation antibody developability model. The model will be trained within Apheris’ secure infrastructure, enabling consortium participants to access the benefits of the collective dataset while maintaining protections around proprietary information.
The consortium’s initial dataset is expected to be delivered to members by early 2027.
Federated Infrastructure Protects Proprietary Data
A central component of the collaboration is Apheris’ federated infrastructure, which is designed to allow organizations to collaborate on AI models without directly exposing their underlying proprietary datasets.
For pharmaceutical and biotechnology companies, data confidentiality is particularly important because antibody sequences can represent valuable intellectual property and may be connected to active or future drug-development programs.
Under the consortium model, participating companies retain ownership of the proprietary sequences and assay data they contribute. Members can use the broader consortium dataset to train, benchmark and refine models without exposing their raw proprietary sequences to other participants.
Apheris’ infrastructure will also enable the foundation model developed from the consortium dataset to be delivered into each member’s environment. Companies can then fine-tune the model using their own proprietary data within their respective environments.
This approach is intended to combine the benefits of collaborative machine learning with the need for individual organizations to maintain control over sensitive drug discovery information.
Robin Röhm, CEO and co-founder of Apheris, emphasized the importance of enabling AI models to perform effectively on each company’s own molecules.
“For AI to impact developability decisions in a drug program, it has to perform on a pharma’s own molecules,” Röhm said. “The Antibody Developability Consortium delivers the largest standardized antibody dataset and the foundation model trained on it. Apheris’ federated infrastructure brings that model to each member and lets them fine-tune it on their proprietary molecules inside their own environment.”
Combining Laboratory Data With AI
Ginkgo Datapoints will play a central role in generating the experimental information needed to build and validate predictive models.
The organization will use its laboratory capabilities to produce antibodies and characterize them across multiple developability-related endpoints. The resulting measurements will provide standardized data intended to support machine-learning applications.
The consortium will also use Ginkgo’s diversity algorithm as part of the sequence-selection process. The objective is to ensure that the dataset represents a broad range of antibody sequences rather than relying on a limited or convenient collection of molecules.
Rich Cohen, Senior Director at Ginkgo Datapoints, said the initiative is designed to address limitations associated with existing datasets.
“We are building the largest, most standardized antibody developability dataset the industry has ever seen, along with the predictive models trained on it,” Cohen said. “Ginkgo Datapoints brings the lab data generation scale, the diverse sequence selection expertise, and the published modelling track record needed to lead this initiative.”
The combination of experimental data generation and AI modeling is intended to create a foundation for earlier and more systematic evaluation of antibody candidates.
Scientific Oversight From Academic Experts
The consortium will receive independent scientific oversight from Charlotte Deane, Professor of Structural Bioinformatics at the University of Oxford, and Peter Tessier, Professor of Pharmaceutical Sciences and Chemical Engineering at the University of Michigan.
Their involvement is intended to provide additional scientific guidance as the consortium develops its dataset, experimental strategy and predictive models.
The initiative is designed to give participating pharmaceutical companies access to a broader portfolio-level developability capability without requiring each organization to independently create a dataset of the same scale or build the necessary infrastructure from scratch.
AbbVie Highlights Machine-Learning Dataset Design
Athena Hadjixenofontos, Director of Data Science and Head of AI in Biotherapeutics and Genetic Medicine at AbbVie, said the consortium could help address limitations in the datasets commonly used to develop predictive models.
“This consortium represents an important step forward in building predictive models for antibody developability by creating datasets that are designed for machine learning, addressing limitations associated with convenience datasets,” Hadjixenofontos said.
She also highlighted the role of federated infrastructure in allowing companies to participate in collaborative data development while keeping proprietary sequences private.
“These capabilities could meaningfully accelerate antibody discovery and help advance new medicines for patients,” she added.
The emphasis on machine-learning-ready datasets reflects a growing focus across drug discovery on generating experimental data specifically structured for computational modeling rather than relying solely on historical datasets assembled for other purposes.
Industry Collaboration Across Multiple Therapeutic Areas
The consortium’s founding members bring experience across different areas of drug discovery, creating an opportunity to evaluate developability models across a broad range of antibody programs.
At argenx, collaboration is an important component of the company’s approach to antibody-based drug discovery. Erwin Pannecoucke, Principal Scientist Discovery at argenx, said predictive developability models could support more efficient discovery and development.
“Predictive developability models can significantly accelerate discovery and development,” Pannecoucke said. “The Consortium’s extensive antibody dataset and federated design enable every partner to learn together and ultimately bring better medicines to patients faster.”
Lundbeck is also participating in the initiative with a focus on the potential role of developability predictions in complex therapeutic areas, including central nervous system diseases.
Allan Jensen, Vice President, Biotherapeutic Discovery, at Lundbeck, said selecting candidates with favorable developability characteristics is important for advancing therapeutic programs.
“In complex therapeutic areas such as CNS, the ability to select well-behaved candidates with superior developability properties is essential,” Jensen said. “By bringing together diverse antibody datasets, this collaboration has the potential to strengthen predictive approaches.”
For Takeda, the initiative aligns with the company’s efforts to incorporate artificial intelligence and machine learning into drug discovery.
Yves Fomekong Nanfack, Head of AI/ML Research at Takeda, said combining standardized data across companies could produce predictive models that would be difficult for a single organization to build independently.
“Pooling standardized developability data across the industry can create stronger predictive models than any one company could build alone,” Nanfack said.
Expanding the Scope of Antibody Developability
The initial consortium dataset will focus on establishing a standardized foundation for antibody developability modeling. Over time, participants plan to explore adding more complex antibody formats and additional properties.
Expanding the dataset could eventually support new classes of antibody-based medicines and allow researchers to evaluate a broader range of characteristics earlier in the discovery process.
The longer-term objective is to improve the ability of computational models to distinguish candidates that are more likely to progress successfully from those that could encounter manufacturing or formulation challenges later in development.
By identifying potential issues earlier, drug discovery teams could potentially reduce the number of candidates that progress into resource-intensive development activities only to encounter avoidable developability barriers.
The Antibody Developability Consortium therefore combines three major components: large-scale standardized laboratory data generation, AI-based predictive modeling and privacy-preserving infrastructure for industry collaboration.
With AbbVie, argenx, Lundbeck and Takeda as founding members, and additional pharmaceutical and biotechnology companies invited to participate, the initiative is designed to establish a shared foundation for antibody developability research.
As the initial 10,000-antibody dataset moves toward delivery in early 2027, Ginkgo Datapoints and Apheris will continue developing the experimental and computational infrastructure intended to support the consortium. The effort reflects a broader shift toward data-intensive and AI-enabled approaches to biopharmaceutical discovery, with the potential to help researchers make developability assessments earlier and incorporate those insights into antibody candidate selection and development planning.
About Ginkgo Bioworks
Ginkgo Bioworks builds the tools that make biology easier to engineer for everyone. The company offers autonomous laboratories that replace manual laboratory work with robotics in the lab, greatly improving the productivity of scientists. Ginkgo’s in-house autonomous lab is also available as a “Cloud Lab” through our Datapoints and Solutions contract research services. For more information, visit ginkgobioworks.com, read our blog, or follow us on social media channels such as X (@Ginkgo), Instagram (@GinkgoBioworks), Threads (@GinkgoBioworks), or LinkedIn.
About Apheris
Apheris powers the largest federated data networks in drug discovery, used by leading pharmaceutical and biotech companies across domains including co-folding, binding and structure prediction, ADMET, in vivo PK, antibody developability, and virtual cell modeling. Members provide privacy-preserving access to their proprietary data and, in return, get models with higher performance and a broader applicability domain. Apheris’ expertise lies in driving the adoption of AI models within pharmaceutical drug programs, customizing and fine-tuning them on each company’s proprietary chemistry and biology so they perform where public data cannot reach.

