Knowledge Engineering in the Age of Generative Artificial Intelligence: Towards a New Discipline of Knowledge Engineering
Abstract
The rapid adoption of Large Language Models (LLMs) has profoundly changed how organizations produce, manage, and use their knowledge. To date, research and industrialization efforts have focused primarily on model performance, agent architectures, prompt engineering techniques, and Retrieval-Augmented Generation (RAG) mechanisms. Relatively little work, however, has examined how knowledge itself should be designed so that it can be reliably reused by artificial intelligence.
This article proposes the hypothesis that a new stage of maturity is emerging: Knowledge Engineering applied to generative systems. This discipline seeks not only to structure knowledge for humans, but also to design knowledge that conversational agents, digital assistants, and multi-agent systems can use directly.
The objective of this article is to define the research framework that will guide MuBrain’s future work.
1. Context
An organization’s knowledge has long been produced according to a document-centric model.
Procedures.
Reports.
Standards.
User guides.
Document repositories.
This content had one primary objective: to be read and interpreted by a human.
The arrival of language models profoundly changes this assumption.
For the first time, a significant share of an organization’s knowledge is no longer consulted directly by an employee, but by artificial intelligence acting as an intermediary.
This evolution changes the very nature of knowledge.
The question is no longer simply:
How should information be documented properly?
It becomes:
How should knowledge be designed so that it is interpreted correctly by artificial intelligence while remaining understandable to a human?
2. Limitations of Current Approaches
The approaches currently observed in organizations rely primarily on four strategies.
- improving prompts;
- adding conversational memory;
- enriching context through RAG;
- specializing agents.
These approaches improve system performance.
However, they do not answer a fundamental question.
Does response quality depend primarily on the model—or on the quality of the knowledge provided to it?
This question remains largely open.
3. Research Hypothesis
MuBrain’s proposed work is based on the following hypothesis.
Given models of equivalent quality, explicitly structured knowledge has a measurable influence on the quality, reproducibility, and stability of responses produced by generative artificial intelligence.
This hypothesis naturally shifts the object of study.
The goal is no longer only to optimize the model.
It is to optimize knowledge.
4. Proposed Definition
For the purposes of this work, we propose the following definition.
Knowledge Engineering is the discipline of designing, structuring, maintaining, and governing knowledge to optimize its simultaneous use by humans and artificial intelligence.
This definition introduces several new characteristics.
Knowledge becomes:
- version-controlled;
- modular;
- composable;
- traceable;
- reusable;
- portable across different assistants.
These properties are now commonplace in software engineering.
They remain largely absent from traditional document management systems.
5. From Documentation to Skills
Organizations today primarily document their processes.
Artificial intelligence, however, requires more than documentation alone.
It requires methods.
A procedure describes a sequence of actions.
A skill describes a way of acting.
This distinction is fundamental.
For example, a document explaining how to write a report constitutes a procedure.
By contrast, a description of quality criteria, assumptions to verify, common errors, and checkpoints constitutes a skill.
Knowledge can then be applied directly when performing a task.
6. A Multilayered Knowledge Architecture
MuBrain’s initial work has led to a distinction between several layers of knowledge.
Layer 1 — Personal Memory
Relatively stable information about a user.
Preferences.
Writing style.
Professional context.
Layer 2 — Standing Instructions
General operating rules.
For example:
- distinguish facts from hypotheses;
- flag missing information;
- cite sources.
Layer 3 — Skills
Specialized methods associated with a family of tasks.
Examples:
- writing a report;
- analyzing a budget;
- preparing a training session;
- creating a presentation.
Layer 4 — Assignment
Description of the requested work.
This separation avoids systematically repeating the same knowledge in every interaction.
7. Skills as Fundamental Units
We propose introducing the concept of a Knowledge Skill.
A Knowledge Skill is a self-contained unit of knowledge that describes a method for completing a task rather than an expected result.
A skill may include:
- its objective;
- its prerequisites;
- its steps;
- its quality criteria;
- its checkpoints;
- its common errors;
- its dependencies on other skills.
This structure deliberately echoes the concepts of software modules and reusable components.
8. Knowledge Governance
One direct consequence of this approach concerns governance.
Skills become information assets.
They require:
- version control;
- scientific or subject-matter validation;
- traceability of changes;
- peer review;
- controlled distribution.
This approach progressively brings knowledge closer to established software development practices.
9. Research Directions
MuBrain’s future work will focus in particular on the following questions.
-
How can the effect of knowledge structure on the quality of responses produced by an LLM be measured objectively?
-
What criteria can be used to evaluate a skill independently of the model being used?
-
Can the same skill be used identically by different assistants (Microsoft 365 Copilot, ChatGPT, Claude, or other agentic systems)?
-
What mechanisms can ensure the governance, versioning, and distribution of skills across an organization?
-
Is there an optimal level of granularity for a skill that balances reusability, readability, and conversational efficiency?
-
To what extent does explicitly structuring skills improve the reproducibility of results across multiple users?
These questions will form the foundation of the experiments published in future work.
Conclusion
The emergence of generative artificial intelligence challenges more than human-computer interfaces.
It calls into question the very nature of the knowledge we produce.
Just as software progressively led to the emergence of disciplines such as Software Engineering and Data Engineering, it is reasonable to consider that conversational systems now require their own discipline for structuring knowledge.
Knowledge Engineering, as proposed in this article, is not a departure from historical approaches to knowledge engineering.
It is an evolution of them.
Its object of study is no longer limited to knowledge representation, but extends to the capacity of knowledge to be reliably understood, reused, governed, and transmitted by artificial intelligence working with humans.
The work presented here is an initial proposal for a conceptual framework. Its validation will rely on reproducible experiments, objective measurements, and comparisons between different approaches to structuring knowledge. Only under these conditions can the hypotheses put forward be progressively confirmed, refined, or refuted, in keeping with the scientific method that will guide MuBrain’s research.