The design of new proteins or the discovery of small molecules with therapeutic activity often hits a critical bottleneck: the scarcity of labeled data to train predictive models. In this context, the scientific community has explored approaches based on in-context learning, a technique in which a model trained on synthetic tables —generated from random causal graphs— is capable of solving tasks with very few examples. Surprisingly, these foundational tabular models, such as TabPFN or TabICL, show competitive performance even in biomolecular domains, where the structure of the data (protein sequences or molecular graphs) seems unrelated to the causal prior with which they were trained.
The key lies in pairing the predictor with an appropriate representation. When evaluating these models on benchmark sets such as ProteinGym for enzymatic fitness regression or TDC ADMET for small molecule classification, it is observed that the choice of molecular descriptor (for example, ESMC, ECFP, or RDKit) strongly conditions the result. The tabular predictor, lacking a specific inductive bias toward biological matter, delegates that responsibility to the encoder that converts the sequence or graph into a feature vector. Thus, the success of the approach depends not only on the learning model, but on the synergy between the representer and the in-context classifier.
This flexibility opens practical opportunities for companies seeking to accelerate their R&D pipelines. Implementing a system that combines pre-trained tabular models with state-of-the-art biomolecular representations can drastically reduce the number of experiments needed. At Q2BSTUDIO we develop custom applications that integrate artificial intelligence for companies, allowing laboratories and pharmaceutical companies to deploy AI agents capable of predicting properties with little data. Our teams design architectures that leverage AWS and Azure cloud services to scale training and inference, while ensuring cybersecurity in the handling of sensitive information.
Additionally, to monitor and analyze the results of these models, we offer business intelligence services based on Power BI, which transform predictions into actionable dashboards. In this way, research teams can make informed decisions without relying on costly laboratory trials. The combination of in-context tabular models with appropriate representations represents a promising path to overcome data limitations in biomedicine, and at Q2BSTUDIO we are prepared to help organizations adopt this technology efficiently and securely.

.jpg)
