In the current landscape of digital education, automated essay evaluation has become a growing necessity. WrAFT (Writing Assessment and Feedback Tool) emerges as a modular solution that combines artificial intelligence with advanced architectural design to deliver accurate scores and comprehensive feedback on argumentative essays. Unlike monolithic systems, WrAFT divides the evaluation process into three independent modules: scoring, surface-level feedback, and deep-level feedback. This separation not only facilitates maintenance and updates of each component but also allows customization for different educational or business contexts.
The scoring module uses large language models (LLMs) such as LLaMA-3.3-70B-Instruct, GPT-4o, and Claude 3.7, evaluated through direct prompting and supervised fine-tuning. Using a proprietary dataset of 480 TOEFL Independent Writing essays, WrAFT achieves state-of-the-art performance: a quadratic weighted kappa (QWK) of 0.84 and a root mean square error (RMSE) of 0.44 on a 0–5 scale. These numbers reflect a high correlation with official scores, making it reliable for academic and corporate use.
Surface-level feedback addresses grammar, spelling, and style issues, while deep-level feedback is divided into macro (argumentative structure, global coherence) and micro (sentence-level cohesion, connector usage). In a human evaluation, approval rates reached 96.14% for surface-level feedback, 93.03% for deep-level macro, and 94.69% for deep-level micro. This demonstrates that WrAFT is not only accurate but also useful and acceptable to end users.
From a technical perspective, WrAFT's modular architecture allows different AI models to be swapped out. For example, a company could replace the base LLM with a lighter one for resource-constrained environments, or add a specific artificial intelligence module to detect bias. Additionally, the system can be deployed on the cloud using services like AWS or Azure, ensuring scalability and availability. The generated feedback can feed Business Intelligence dashboards, such as Power BI, to analyze student performance trends and improve teaching processes.
For organizations looking to implement similar solutions, the development of custom software is key. WrAFT is an example of how combining independent modules, AI agents, and user-centered design can transform educational assessment. At Q2BSTUDIO, as a software and technology development company, we understand that each client has unique needs. That is why we offer services ranging from modular AI system creation to cloud platform integration and BI dashboard implementation. Cybersecurity is also a fundamental pillar: protecting student data and evaluations is critical, so we apply pentesting and encryption practices in all our solutions.
Looking ahead, the trend points toward increasingly autonomous systems. AI agents can personalize feedback according to the student's profile, while process automation reduces the administrative burden on teachers. WrAFT, with its public and free interface, is already paving the way. However, to scale these innovations to an enterprise level, it is necessary to have technology partners who master both theory and practice. Integrating language models with cloud infrastructure and data analytics is not trivial; it requires expertise in software engineering, machine learning, and information security.
In short, WrAFT shows that automated essay evaluation can be accurate, useful, and modular. Companies that want to adopt similar technologies should consider a strategy that combines AI, cloud, and BI, always with a focus on customization. At Q2BSTUDIO, we help organizations design and implement these solutions, ensuring each component fits their specific goals. Whether building a system from scratch or integrating existing modules, the key lies in modular architecture and intelligent use of data.





