PHITSBench: Execution-scored benchmark for AI radiation transport simulation

Evaluate AI-assisted PHITS radiation transport simulation generation with PHITSBench. Results show success rates, failure analysis, and future directions.

martes, 28 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Evaluación de IA en generación de simulaciones PHITS

In the field of scientific simulation, artificial intelligence is beginning to transform how complex phenomena such as radiation transport are modeled. However, for generative models to be truly useful in production environments, rigorous benchmarks are needed that evaluate not only the ability to generate syntactically correct code, but also the physical accuracy of the results. In this context, PHITSBench was born—a benchmark specifically designed to measure the performance of AI models in generating and modifying simulations using the PHITS code (Particle and Heavy Ion Transport code System). Throughout this article, we explore what PHITSBench is, what its first results reveal, and how software development companies like Q2BSTUDIO can help integrate these capabilities into real solutions, combining custom software, cloud computing, cybersecurity, Business Intelligence, and AI agents.

PHITSBench consists of 282 tasks covering three main workflow categories: parameter editing (Edit), syntax repair (Repair), and complete simulation generation from natural language descriptions (Reproduce). Each task is evaluated using a composite metric that combines execution success with agreement between generated and reference transport observables. This approach measures both syntactic correctness and physical validity of the simulations produced by the model. Tests were conducted with GPT-5.4-based configurations, ranging from zero-shot prompting to knowledge-augmented and autonomous agent workflows.

The initial results are revealing: in the absence of domain-specific knowledge, the model achieves 95% success on editing tasks and 70% on repair tasks, but fails completely at generation from scratch (0% on Reproduce). This shows that although large language models can manipulate existing code with some fluency, they lack the deep knowledge needed to build complete and physically coherent simulations. However, when a structured machine-readable PHITS knowledge catalog is provided—alongside the user manual—Reproduce success jumps to 57%. Agentic execution raises it up to 73%, albeit at higher computational cost.

Failure analysis shows that remaining errors are not mainly syntax-related, but rather due to incorrect selection and configuration of physical observables. That is, the model knows how to write the code, but not which physical parameters to choose for the simulation to make sense. This observation is key: the next frontier in AI-assisted simulation is not just about improving foundation models, but about building machine-readable knowledge bases, curated training datasets, and execution-grounded evaluation environments.

For companies working with scientific or industrial simulations, these findings have direct implications. Integrating AI into simulation processes requires more than a language model: it needs a solid software infrastructure capable of managing data, running simulations in the cloud, ensuring system security, and extracting insights from results. This is where the capabilities of AI and cloud AWS/Azure offered by Q2BSTUDIO come into play.

Q2BSTUDIO is a software and technology development company specialized in creating custom software tailored to the specific needs of each organization. In the context of radiation transport simulation, for example, we could develop a platform that combines an intuitive frontend for describing scenarios in natural language, a backend that invokes AI models like those evaluated in PHITSBench, and a cloud execution engine that launches simulations on AWS or Azure. All integrated with cybersecurity systems to protect sensitive data and BI/Power BI dashboards to visualize results interactively.

Creating AI agents is another area where Q2BSTUDIO can make a difference. Autonomous agents, like those tested in PHITSBench, can iterate over simulations, correct errors, and optimize parameters, but they require careful orchestration and domain knowledge that must be encoded in rules or knowledge bases. Our team can design and implement these agents, integrating them with structured knowledge catalogs and simulation APIs.

Furthermore, Q2BSTUDIO's expertise in custom software allows us to build complete solutions that not only run simulations but also manage the entire lifecycle: from input data ingestion to result publication, including script version control, cloud cost monitoring, and security auditing. All following best software development practices and agile methodologies.

Returning to PHITSBench, the fact that errors concentrate on the choice of physical observables underscores the importance of specialized knowledge bases. At Q2BSTUDIO, we can help organizations build these bases, extracting knowledge from technical manuals, scientific articles, and historical data, and structuring them in formats that AI models can consume (such as knowledge graphs or embeddings). This, combined with the power of foundation models and cloud scalability, can bring AI-assisted simulation to a professional level.

Finally, we cannot forget cybersecurity. Radiation transport simulations may involve critical or regulated data, especially in nuclear or medical environments. Q2BSTUDIO implements security measures at all layers: from data encryption at rest and in transit to multi-factor authentication and network segmentation in the cloud. Our cybersecurity service includes audits and penetration testing to ensure the infrastructure is robust against threats.

In conclusion, PHITSBench represents a significant advance in evaluating AI models for scientific simulation. Its results show that while language models are powerful, they need a support ecosystem that includes knowledge bases, cloud execution, and intelligent agents. Q2BSTUDIO is ready to provide that ecosystem, combining its expertise in custom software, cloud AWS/Azure, cybersecurity, BI/Power BI, and AI agents to help companies harness the full potential of AI-assisted simulation, whether in radiation transport or any other technical field. The future of simulation is not just more artificial intelligence, but an intelligent integration of tools, data, and knowledge.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.