Explaining Knowledge Is Not the Same as Designing Understanding: Why We Built CEA
Some Background That Is Not Very Related (You Can Skip This)
Our team was originally just a very traditional content creation team. We have used AI tools heavily since the DALL·E period, but until then, we had never thought about doing AI research ourselves. In December 2025, we were writing a social mystery story. For some reason, an idea came up:"Can we build an LLM that we can talk to while creating a story?"
Writers with experience creating stories know this: a great human creative partner does not write the story directly, but they can be a great help to the process. Such a creative partner has clear qualities, which I will not list here. In late 2025, almost none of the LLMs we tested had these qualities.
We started exploring and learning with little knowledge and a lot of courage. We originally thought the study and research process would be very hard. Fortunately, around the same time, Google launched Antigravity. Later, during the 2026 Chinese Spring Festival, Antigravity made the excellent Claude Opus 4.6 model available. These external tools made our research and learning full of joy, even though it was still not easy.
Our early goal of "building an LLM that we could talk to while creating a story" changed along the way. The research direction eventually became Creative Quality Alignment (CQA). In May this year, our team completed one systematic round of CQA research.
Background of CEA Research
After finishing the initial exploration of CQA, I naturally became a bit overconfident.
What should we do next?
Our team has done science communication for 13 years. We have accumulated many original drafts whose copyrights we own, as well as mature professional judgments. At that time, an idea came to my mind: Can we turn these experiences into a kind of post-training data?
I shared this idea with a friend who also works in science communication. His first reaction was that this had no value. His reasoning was: AI is already very good at explaining science. If you do not understand a concept, you ask it to rephrase or make an analogy. After one or two rounds of questions, human users can usually understand most of it.
I thought about it and felt he was right.
Today's LLMs are extremely strong at explaining a single concept. Their knowledge coverage and instant rephrasing far exceed those of any single human.
But after one night, I suddenly realized: that is not right.
Translating knowledge is not the same as creating a work of science communication. LLMs can translate knowledge very well, but that does not mean they can write a truly good science communication article. At least in my professional judgment, top human writers are much better than LLMs.
So, there is still room for us to explore.
What Is Still Missing Between Excellent Knowledge Translation and Good Writing?
LLMs have “knowledge translation” abilities that humans cannot match. Why does this not fully turn into excellent writing?
In a one-on-one conversation, the user's initial question already defines what they care about. During the chat, praise or complaints from the user are real-time signals. The model can adjust and continue explaining based on this feedback.
Pre-written content has no real-time feedback at all. The writer must estimate the audience's possible reactions on their own. Then, the writer uses these estimates to guide the writing process, making the final piece engaging and clear to read.
Here are some of the things writers have to estimate: Why would the audience want to enter this topic? What do they already know? Which step will feel difficult to them? After which sentence might they have a question? When should we give an easy answer first, and when should we deliver a more abstract principle? Is a joke just for fun, or does it also help understanding?
Professional communicators do not just translate knowledge. They constantly simulate the audience's mental state. Then they turn these judgments into decisions about entry points, order, level of detail, transitions, and endings.
Research in educational psychology on prior knowledge, working memory, and cognitive load has long shown that new information is not understood in a vacuum. How information is organized directly affects whether learners can process it. For example, the 20-year review of Cognitive Load Theory by Sweller, van Merriënboer, and Paas focused on the relationship between limited working memory and instructional design.
A piece of work is not just a collection of correct sentences. It is a specific cognitive path designed for the audience to walk through. Doing research from this perspective gives us clear room to work. This is why we built CEA (Cognitive Engagement Alignment).
However, CEA is not about turning existing educational theories into data labels. What we care about is this: In a complete work, how do real professional creators turn their audience judgments into one concrete choice after another?
The Final Text Keeps the Words, but Loses the Judgments
The problem is that once a piece of work is finished, the design process behind this path usually disappears. Neither human readers nor LLMs can recover the author's earlier creative choices from the final work alone.
Common formats of training data save many important signals, but they usually do not fully preserve this mapping:
- Finished-text corpora preserve "what was written in the end";
- Instruction data saves "how to answer when given a question";
- Preference data saves "which output humans prefer among several candidates".
For example, the InstructGPT study used human demonstrations and model output rankings. This kind of data has been shown to improve model behavior. CEA is not trying to replace it. Instead, CEA adds another relationship: how experts judge the audience's cognitive state at a specific point in a full piece of work, and how that judgment leads to concrete content decisions.
Selecting an answer does not mean the conditions, reasons, and trade-offs behind each local choice are saved. A good article entering the pre-training corpus does not mean the model can reliably infer the author's audience model from the final text.
Turning Hidden Audience Judgments into Data
CEA is a dataset of expert decisions on how professional creators design the audience's cognitive engagement process.
It preserves a complete piece of work and uses expert annotations that point directly to the specific locations in the text that they apply to. The data does not just record "this sentence is well written." It records what state the audience might be in at this moment, what role this writing choice plays, and why this arrangement is needed right here.
The first two public records we released both come from real science communication practices at Bread Studio.
In Long March 7: The Heavy Hauler of Space, the article does not list rocket parameters at the start. Instead, it begins with a question:
What is a launch vehicle?
Put simply, it is a vehicle that transports things into space.
The creator's annotation points out that this very simple question and answer gives general audiences an early feeling of "I get it." The article then fills in basic background on payload, engines, and historical comparisons. Only after the audience has the necessary background does it officially introduce Long March 7. Once it reaches the main subject, it does not aim for an exhaustive technical checklist. Instead, it uses "three major advantages" to build a bounded, clear, and memorable impression.
In Why Do Some People Always Feel Like Peeing in the Shower?, the first obstacle was not knowledge, but embarrassment. Therefore, we designed a separate questioner to ask the awkward question first on behalf of the audience. This allows the audience to step back into the role of an observer. After explaining the first reason, the phrase "Besides conditioned reflexes" sends a clear structural signal: the previous attention task has ended, and the next task begins. The alien joke later is also not just decorative. By suddenly creating tension, it reinforces from the opposite side why relaxation affects the urge to pee.
Each of these choices looks small on its own. Put together, they form an understanding path that the audience can enter, follow, and finish. What CEA does is make these hidden judgments readable, traceable, and verifiable.
The current Chinese passage-level annotations were written by the original creators and reviewed by the team's editor-in-chief. AI tools helped with text formatting, internal position indexing, and span mapping, but did not write first-person expert judgments in place of the creators. In the dataset, we also clearly distinguish between original expert annotations, AI-assisted formatting, and expert-verified derived summaries.
CEA Is More Than a Science Writing Corpus
Science communication is the first area where we can provide real sources, clear rights, and traceable creative rationale, but it does not define the full scope of CEA.
Education, journalism, policy communication, medical information, legal and financial explanations, and other tasks that organize complex content for specific audiences all share similar questions: How does the audience enter a topic? Where do they face cognitive load? How do they build or revise their understanding? In broader creative writing, if expert judgments deal with how the audience enters the content, forms expectations, and connects ideas, that part can also belong to CEA. Character, language style, narrative aesthetics, and artistic quality, on the other hand, are mainly covered by our earlier CQA framework.
This expansion does not mean that every design affecting attention and emotion belongs to CEA. Optimizing purely for clicks, watch time, conversions, or emotional arousal is not CEA. Manipulative persuasion aimed at bypassing audience judgment is also not CEA. Only expert judgments that directly address how the audience forms or revises its understanding of a subject fall within the research scope of CEA.
What Do Two Samples Prove?
The current public preview contains two full Chinese articles, 16 passage-level expert annotations, and their corresponding whole-article summaries. Bread Studio has about 450 other organized original drafts whose copyrights we own, but they have not all been fully annotated by experts, and they are not included in this open release.
These two public records directly prove something small but concrete: Retrospective judgments by professional science communicators about audience attention, understanding, cognitive load, and expectations can be collected, grounded, and structured into real, verifiable data artifacts.
They do not directly measure whether real audiences understood the text exactly as experts expected. Nor do they prove that these judgments have causal effects or represent all creators. Larger-scale data, audience studies, training experiments, and systematic evaluations are still needed to determine whether CEA can provide training gains and transfer to education, medical information, policy communication, law, finance, or other fields.
Therefore, CEA is not yet an answer verified by empirical results. It is a research hypothesis backed by concrete artifacts: If we want models not only to answer knowledge questions, but also to plan, generate, and evaluate complete content for specific audiences, then training data should probably also preserve how experts design the audience's understanding process.
You can check the CEA: Cognitive Engagement Alignment Public Preview on Hugging Face. The Chinese records have been verified by experts. The English records share the same IDs as the Chinese records and correspond to them, but currently serve only as AI-assisted reading translations.
For research, data licensing, or custom projects, contact: bo@BreadMedia.com