SEFORA: Essay corpus and LLM feedback evaluation framework

SEFORA: public corpus of essays with teacher feedback. UniMatch evaluates whether LLMs generate aligned feedback. Result: no model exceeds 0.4 F1.

jueves, 2 de julio de 2026 • 2 min read • Q2BSTUDIO Team

UniMatch: measuring the precision of AI feedback

Effective feedback in academic writing is one of the most decisive factors for learning, but providing it at scale remains a major logistical challenge. Recently, the scientific community has introduced SEFORA, a public corpus that collects instructor annotations on real essays, and UniMatch, an evaluation framework designed to measure the quality of feedback generated by large language models. This advancement not only has implications for the educational field but also opens the door to more precise and contextual artificial intelligence applications in corporate environments, where the automatic generation of reports, corrections, and feedback is increasingly in demand.

Evaluating feedback generated by artificial intelligence requires robust metrics that capture the instructor's intention. UniMatch segments feedback into units, compares them semantically according to pedagogical criteria, and aligns them through optimal matches to obtain precision, recall, and F1 scores. Experimental results show that no model exceeds an F1 of 0.4, indicating that machines still struggle to prioritize the comments a human teacher would consider essential. This gap is especially relevant for companies looking to integrate AI agents capable of reviewing complex documents, drafting technical reports, or assisting in internal training processes.

At Q2BSTUDIO, we understand that implementing artificial intelligence-based solutions is not limited to deploying a pre-trained model. A comprehensive approach is necessary, combining custom software with robust cloud infrastructure and cybersecurity measures. For example, for an automatic feedback system to work in an organization, a platform is required that manages data securely, processes large volumes of text, and offers interfaces adapted to existing workflows. Our AWS and Azure cloud services allow these systems to scale flexibly, while our business intelligence and Power BI capabilities facilitate the visualization of model performance metrics.

The challenges identified in SEFORA and UniMatch —imperfect semantic alignment, degradation with extensive generations— are analogous to those faced by companies when implementing AI for businesses. It is not enough to generate content; it must be relevant, coherent, and safe. That is why at Q2BSTUDIO we work on developing custom applications that integrate personalized criteria, similar to the principles of UniMatch, but adapted to business processes. Furthermore, incorporating AI agents capable of making decisions based on rules and machine learning is one of our most promising lines of innovation.

In conclusion, research on LLM feedback reminds us that artificial intelligence still requires careful design and rigorous evaluation to be truly useful. From custom software development to process automation, at Q2BSTUDIO we offer solutions that combine cutting-edge technology with a deep understanding of each client's real needs. If your organization is looking to implement intelligent content review or writing assistance systems, we can help you build the right infrastructure, ensuring both efficiency and security.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.