Picture an eighth-grade history class. A student hands in a clear, confident paragraph about the Dust Bowl. One date is wrong. When her teacher asks where it came from, she shrugs. “The AI said so.”
She isn’t being lazy. She was handed a tool that writes like a person, and nobody showed her what it actually does with words.
Students are using chatbots. Do they know what they are using?
Pew Research Center surveyed 1,458 U.S. teens ages 13 to 17 in fall 2025. It found that 64% use AI chatbots, and 54% say they have used them to get help with schoolwork. That is a finding about use. It says nothing about understanding, and that gap is our concern here.(Pew Research Center)

Using AI and understanding AI are not the same educational objective. The first can be learned in an afternoon. The second needs a few ideas that most chatbot screens hide. The first of them is that, to a language model, language is data: before it can work with a sentence, the sentence has to become numbers.
What “language as data” actually means
A computer has no ears and no sense of meaning. So the first step is to cut text into small pieces called tokens, usually whole words or fragments of words, and give each piece an ID number. The second step is to turn each token into a longer list of numbers, called an embedding. These lists are built from patterns of use, so words that appear in similar contexts, like “teacher” and “professor,” end up with similar numbers.
Nothing in that process knows what a teacher is. The model has a map of how words are used, not a picture of the world. Students can see this for themselves by splitting a sentence into tokens by hand, then comparing how close two words sit on a simple map. A short hands-on session can make the idea concrete.

Three ideas worth assessing
Context is calculated. The word “bank” means different things next to “river” and “loan.” Modern language models handle this with attention, a mechanism that weighs which other words in the sentence matter for reading each word. Students don’t need the math at first. They need to know that context is something the model computes, not something it simply grasps.
Fluent is not the same as true. A language model generates text by choosing a likely next token, over and over. Likely and correct often overlap, but not always, and that is how a smooth paragraph ends up with a wrong date. If a student can get a fluent answer but cannot explain why it might be confidently wrong, what exactly have they learned?
The data has people in it. The text a model learns from was written by people, so it carries their habits, including skewed ones. Sentiment analysis is a good small-scale demonstration. Have students run a sentiment tool on real texts and hunt for mistakes, such as sarcasm read as praise. We think finding a failure first-hand is a better starting point for a talk about bias than being told “AI can be biased.”
Where this belongs in a curriculum
Not in a single prompt-writing lesson. Language as data leans on earlier ideas: models learn from examples, examples can be skewed, and a classifier can be wrong. A school that has already covered neural networks through small models will find this a natural next step, because large language models are neural networks working on tokens. Our view is that middle school is a workable home for tokenization and embeddings, with the ideas returning in high school at greater depth. I’m not aware of a U.S. standard that fixes one grade for this, so leaders have to make that call locally.
When reviewing any unit, ask three things. Do students watch text become numbers themselves? Do they test a model and record its failures? Is responsible use taught beside the mechanism, not in a separate policy handout? These checks separate generative AI lessons that go beyond prompting from tool tutorials. They also prepare students for questioning plausible AI-generated claims and for the wider picture in what students should understand about generative AI.
FAQs
Q: What does “language as data” mean in AI education?
A: It means an AI system does not read words the way a person does. It splits text into tokens, turns them into numbers, and does math on those numbers. Students who understand this can explain how chatbots, translators and sentiment tools work, and why they make mistakes.
Q: What grade should students start learning how language models work?
A: There is no single answer I can point to in U.S. standards. In our view, Grades 6 to 8 work well for tokenization and embeddings, once students have met training data and classification. Younger students can start with the simpler idea that AI learns from examples.
Q: Do students need to code to learn natural language processing?
A: Not at the start. Splitting sentences into tokens by hand, comparing word similarity, and testing a sentiment tool on real texts teach the core ideas without code. Coding can follow once the concepts are clear.
Q: How is teaching language as data different from teaching prompt writing?
A: Prompt writing teaches how to get better output from a tool. Language as data teaches why the tool behaves the way it does, including why it can sound sure and still be wrong. Schools need both, and the second makes the first safer.