OpenAI's GPT-4b micro redesigns Yamanaka factors with 50-fold gains
OpenAI's GPT-4b micro, built with Retro Bio, redesigned Yamanaka factors to achieve over 50-fold higher reprogramming marker expression and better DNA repair than wild-type proteins.

Updated
Why it matters
- GPT-4b micro, a miniature GPT-4o variant trained for protein engineering, designed Yamanaka factor variants achieving greater than 50-fold higher expression of stem cell reprogramming markers than wild-type controls in vitro.
- Over 30% of the model's proposed SOX2 variants outperformed wild-type despite differing by more than 100 amino acids on average; conventional mutation screens typically see hit rates below 10%.
- Validated in mesenchymal stromal cells from donors over 50 years old delivered via mRNA, over 30% of cells expressed pluripotency markers within 7 days, and derived iPSC lines showed normal karyotypes and genomic stability.
OpenAI says a custom miniature version of GPT-4o, trained for protein engineering and deployed with longevity biotech startup Retro Bio's Applied AI team, has produced redesigned Yamanaka factor proteins that achieved greater than a 50-fold higher expression of stem cell reprogramming markers than wild-type controls in vitro. The redesigned proteins also showed enhanced DNA damage repair capabilities, which points to higher rejuvenation potential than the baseline factors. OpenAI made the finding in early 2025 and has now validated it through replication across multiple donors, cell types, and delivery methods, with confirmation of full pluripotency and genomic stability in derived iPSC lines.
The stakes here are considerable. The Yamanaka factors — OCT4, SOX2, KLF4, and MYC, together known as OSKM — are among the most important protein sets in regenerative biology. Their discoverer, Shinya Yamanaka, received the Nobel Prize in Physiology or Medicine in 2012 for showing that they can reprogram adult cells into induced pluripotent stem cells (iPSCs). iPSCs have since been used to develop therapeutics to combat blindness, reverse diabetes, treat infertility, and address organ shortages. But the factors suffer from poor efficiency: typically less than 0.1% of cells convert during treatment, and the process can take three weeks or more. Efficiency drops further in cells from aged or diseased donors, which is precisely the population a longevity-focused company like Retro Bio cares about.
The collaboration produced a model called GPT-4b micro. OpenAI initialized it from a scaled-down version of GPT-4o, then further trained it on a dataset composed mostly of protein sequences, along with biological text and tokenized 3D structure data — elements most protein language models omit. A large portion of the training data was enriched with additional contextual information about the proteins: textual descriptions, co-evolutionary homologous sequences, and groups of proteins known to interact. That context lets users prompt the model to generate sequences with specific desired properties.
Because most of the data is structure-free, GPT-4b micro handles proteins with intrinsically disordered regions just as well as structured proteins. That matters for this particular target. The activity of the Yamanaka factors depends on forming numerous transient interactions with a diverse array of binding partners rather than adopting a single stable structure, as AlphaFold visualizations of KLF4 and SOX2 show — the majority of these proteins are unstructured, with flexible arms that attach to other proteins.
The enriched training data also had an architectural payoff. By training on proteins with additional evolutionary and functional context, OpenAI substantially increased the effective context length of its training examples beyond standalone sequences. During inference, the team ran prompts as large as 64,000 tokens and continued to observe gains in controllability and output quality. While routine for text LLMs, that context size is unprecedented in protein sequence models, according to OpenAI.
The development process itself mirrored the playbook of language model research. OpenAI observed the emergence of scaling laws similar to those seen in text models — larger models trained on larger datasets yielded predictable gains in perplexity and downstream protein benchmarks. That allowed the team to iterate at small scale before training the final GPT-4b micro model. But OpenAI acknowledges that in silico evals for protein AI models are often of limited value, since it is unclear whether benchmark improvements translate to real-world utility. So Retro's scientists put the model to work on their cell-reprogramming research program.
The scale of the search problem explains why the results are notable. SOX2 contains 317 amino acids and KLF4 contains 513. The number of possible variants is on the order of 10^1000, meaning traditional directed-evolution screens that mutate a handful of residues at a time explore only a miniscule fraction of the design space. A leading academic effort tested a few thousand SOX2 mutants and found a handful of triple-mutants with a modest gain. Fifteen years of work on chimeric SOX proteins has yielded variants that differ from natural SOX constituents by only five residues.
GPT-4b micro operated on a different scale of change. Retro built a wet lab screening platform using human fibroblast cells, initially validating it with baseline OSKM and SOX2 variants manually designed by Retro's scientists. Then the team asked GPT-4b micro to propose a diverse set of "RetroSOX" sequences. Over 30% of the model's suggestions outperformed wild-type SOX2 at expressing key pluripotency markers — even though they differed by more than 100 amino acids on average. In traditional screens, hit rates below 10% are typical.
The team next tackled KLF4, the largest of the Yamanaka factors. KLF4 is known to be replaceable with other KLF-family factors, but without an increase in reprogramming efficacy. Prior attempts to improve KLF4 through expert-guided single amino acid substitutions produced a single hit out of 19. Prompted to generate enhanced RetroKLF variants, the model delivered: 14 model-generated variants were superior to the best cocktails from the RetroSOX screen, a hit rate of nearly 50%.
Combining the top RetroSOX and RetroKLF variants produced the largest gains. Across three independent experiments, fibroblasts showed a dramatic rise in both early markers (SSEA-4) and late markers (TRA-1-60, NANOG), with late markers appearing several days sooner than under the wild-type OSKM cocktail. At day 10, the two late-stage markers were strongly enriched, while no expression could be detected in cells reprogrammed with wild-type OSKM at the same timepoint. Alkaline phosphatase staining at day 10 confirmed the colonies exhibit robust AP activity indicative of pluripotency.
The validation work went beyond a single cell type and delivery method. OpenAI and Retro tested mRNA delivery instead of viral vectors, using mesenchymal stromal cells derived from three middle-aged human donors over 50 years old. Within 7 days, more than 30% of the cells began expressing key pluripotency markers (SSEA4 and TRA-1-60). By day 12, numerous colonies appeared with morphology similar to typical iPSCs. Over 85% of these cells activated endogenous expression of critical stem cell markers, including OCT4, NANOG, SOX2, and TRA-1-60. The researchers verified that the RetroFactor-derived iPSCs could differentiate into all three primary germ layers — endoderm, ectoderm, and mesoderm — and expanded multiple monoclonal iPSC lines over several passages, confirming healthy karyotypes and genomic stability suitable for cell therapies. These results consistently surpassed benchmarks from conventional iPSC lines generated by contract research organizations using standard factors.
The team then tested rejuvenation potential directly, focusing on DNA damage, a canonical hallmark of aging. Earlier work demonstrated that Yamanaka factors can erase DNA damage-related aging markers in cells derived from mice without fully reverting cell identity. In OpenAI's assay, human fibroblasts were treated with doxorubicin to induce double-strand breaks, then reprogrammed with either a fluorescent control, wild-type OSKM, or the engineered RetroSOX + RetroKLF variants. γ-H2AX immunostaining, which quantifies DNA damage, showed cells expressing the Retro variants had a marked drop in signal relative to both controls (GFP vs RS4: p=0.03; GFP vs RS5: p=0.01; OSKM vs RS4: p=0.04). The results suggest the RetroSOX/KLF cocktail reduces DNA damage more effectively than the original Yamanaka factors, offering a potential path toward improved cell rejuvenation and future therapies.
For OpenAI, the collaboration doubles as an argument about method. The company frames the work as an illustration of how quickly a domain-specific model can deliver results on a focused scientific problem. "When researchers bring deep domain insight to our language-model tooling, problems that once took years can shift in days," says Boris Power, who leads research partnerships at OpenAI. "We look forward to seeing what other advances emerge as more teams pair their expertise with the models we're building." OpenAI says it is sharing insights into the research and development of GPT-4b micro to make the findings discoverable and replicable for the life sciences industry.
The high hit rates, deep sequence edits, accelerated marker onset, and AP-positive colony formation provide early evidence that AI-guided protein design can substantially accelerate stem cell reprogramming research — and, if the pattern holds across other protein targets, compress discovery timelines that have historically been measured in decades.
Original: retro.bio
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- OpenAI Launches GPT-Rosalind, a Reasoning Model for Life Sciences
- AI Agents Proposed Over Half the Ideas, Humans Made 85 Percent of Calls
- AI Was Supposed to Hit New Grads Hard. Unemployment Data Says Otherwise
- AI Experts Underestimated the Field's Speed, Study Finds
- Anthropic Says Claude Found a New Enzyme System; CRISPR Researchers Call It Routine