The promise of AI agents applied to data science has seduced many organizations. A professional spends hours on recurring tasks that combine data manipulation, database queries, methodological decisions and presentation of results; all seem ideal for delegating to automated assistants. However, accumulated experience in real environments suggests that automation is not as simple as saving a generic script. The question is not whether AI can help, but what kind of help is truly effective.
In the current ecosystem, the idea of 'skills' or reusable skill files has become widespread. Instead of writing a prompt from scratch every time, the team stores a document with recommendations, examples and constraints for a family of tasks. These skills can be written by experts, with the advantage of condensing best practices, but keeping them up to date across multiple domains becomes a bottleneck. That is why low-curation alternatives have appeared: skills automatically generated by a language model. The issue is whether they offer a real advantage over the basic prompt.
A recent study evaluated this scenario across different phases of the data lifecycle, from initial preparation to final report generation. The results invite caution. The authors found no reliable improvement when using full LLM-generated skills compared with a prompt without any skill. Moreover, when removing components from those skills, performance stayed practically the same. No variant was able to significantly beat the baseline. Even a skill with irrelevant content produced results similar to the full versions. This suggests that, at least in single-use configurations, generated content does not provide the expected guidance.
Why does this happen? An automatic skill tends to be generic. It may include correct definitions, but it lacks the specific context of the organization: how data is modeled, what naming conventions are used, what level of detail the business requires. The result is a text that consumes tokens, portrays a plausible methodology and yet does not guide real execution. In contrast, a well-designed task-specific prompt usually contains the essential information. The conclusion is not that skills are useless, but that automatic generation without criteria does not automatically turn them into valuable tools.
For a company that wants to leverage AI, the lesson is clear: technology must be integrated critically, not as a fad. At Q2BSTUDIO we address this challenge from a technical and business perspective. We help design artificial intelligence solutions that fit real processes, something very different from installing a generic skill. Our experience in custom software development allows us to understand that each organization has different business rules, data sources and quality requirements. Automation based on AI agents must be built on that knowledge, not on reused text.
Infrastructure also matters. Many companies support their pipelines on cloud AWS/Azure, where data resides in distributed environments. An agent that generates a SQL query needs to know the data catalog, access policies and table partitions. That information is usually not in a generic skill, but in a semantic and governance layer. Without it, the model can produce technically correct but unusable answers for the business. That is why at Q2BSTUDIO we consider the cloud to be an enabler, not an end in itself; the combination of cloud, data and AI requires custom design and constant supervision.
Statistical analysis and reporting are another critical front. Once the agent prepares the data, the results are usually displayed in BI/Power BI dashboards. If the language model generates incorrect calculations or biased interpretations, the whole dashboard loses credibility. Organizations must implement validation layers and automated tests on generated SQL and summaries. In our BI/Power BI projects we work so that natural conversation with data does not replace human verification, but complements it. The skill is not in the model, but in the system around it: orchestration, version control, monitoring and acceptance criteria.
The use of AI agents in production also requires an engineering mindset. Publishing a skill and waiting for results is not enough. You have to measure impact, compare with a baseline and decide when the agent can operate without supervision. This rigorous approach is what we apply at Q2BSTUDIO when building AI agents: we define bounded use cases, train the model with company examples and create metrics to evaluate each version. Artificial intelligence thus becomes a real competitive advantage, instead of a technological promise.
It is impossible to talk about AI agents without mentioning cybersecurity. Assistants that access databases, APIs and internal documents handle sensitive information. A poorly curated skill could induce the model to ignore permissions or leak data. Therefore, any automation strategy must include pentesting, access control and continuous auditing. At Q2BSTUDIO we integrate cybersecurity as a transversal layer in projects: from prompt design to production deployment. Trust is not improvised.
So, should companies abandon LLM-generated skills? No. They should treat them as a first draft, not as a final solution. A generated skill can serve to document the initial knowledge of a process, but it needs human review, real examples and alignment with data strategy. It is also worth testing variants: instead of one large skill per functional area, a short prompt specific to each sub-task may be more effective. The study data show that there is no magic formula; improvement comes from iteration.
The future of data scientists and AI agents is not about replacing human judgment, but about creating tools that amplify their analytical capacity. Language-model-generated skills can be useful if they are curated and adapted to context. But when they are used as a shortcut without validation, performance looks suspiciously like an empty prompt. At Q2BSTUDIO we have been applying this philosophy for years: pragmatic technology, results-oriented and backed by solid engineering.
In short, the recent evidence is a valuable reminder. Labeling a content as a 'skill' is not enough for it to work. It takes well-governed data, a prepared cloud AWS/Azure infrastructure, reliable BI/Power BI dashboards, information protection and custom software integration that respects business logic. That is the path for AI agents to improve productivity without compromising quality. And along that path, having a technology partner like Q2BSTUDIO makes the difference.




