Why “hard to read” is hard to pin down
A passage can feel “hard” even when every sentence is grammatical and every word is common. A middle-schooler may stumble on pronouns that refer to ideas two paragraphs back, while an adult gets stuck because the topic assumes unfamiliar science or civic knowledge. Both readers report the same problem, but the cause is different.
That’s why a single readability number is slippery: it compresses vocabulary, sentence structure, cohesion, concept load, and text purpose into one label. Two texts can share a score yet fail different readers for different reasons.
Pinning it down takes observation and time—collecting where readers hesitate, what they misinterpret, and what they skip. That effort is the cost most teams try to avoid, but it’s also where the actionable diagnosis lives.
What actually makes a text difficult for readers
Watch a reader for two minutes and you’ll usually see the culprit. They slow down at dense noun phrases (“the rapid postwar industrial consolidation”), reread embedded clauses, or lose track of “it” and “this” when the referent isn’t nearby. Even small choices—stacked prepositional phrases, long chains of modifiers, or sentences that delay the main verb—can push working memory past its limit.
Difficulty also comes from what the text assumes. Unfamiliar terms, unexplained acronyms, and “everybody knows” background knowledge create invisible gaps. Some passages are hard because their structure hides the point: key claims arrive late, transitions are thin, headings don’t match content, or examples don’t clearly connect to the rule. None of this shows up cleanly in a grade-level score, but it shows up fast in errors, wrong inferences, and skipped lines.
For teams, the constraint is triage: you can’t fix everything. The practical move is to separate problems you can revise (syntax, cohesion, signaling, glossing) from problems you must accommodate with supports (pre-teaching concepts, visuals, or alternative texts).
What language models can explain—and what they can’t

Language models are useful because they can turn a vague complaint (“this is hard”) into a list of plausible friction points. Given a passage, they can flag long modifier stacks, delayed main verbs, abstract nouns that hide actions, or pronouns with distant referents. They can also paraphrase the argument, outline the discourse moves, and suggest where a reader might need a definition, an example, or a clearer transition.
The catch is that these are best treated as hypotheses, not measurements. A model can’t see where your actual students pause, and it may invent “assumed knowledge” that isn’t really required—or miss the one prerequisite concept that matters. It can also sound certain while being wrong about what “this” refers to or what the author’s claim is, especially in technical or poorly edited text.
Someone still has to inspect the flagged spots and test revisions with real readers or tasks, otherwise you just trade one kind of guess for another.
A practical workflow: ask, inspect, then test changes
You can get value from a model without letting it “grade” your text. Start by asking targeted questions tied to a reader and task: “Where might a 6th grader lose the thread?” “Which terms need a quick gloss to answer the end-of-passage questions?” “What background knowledge is assumed but not stated?” The goal is a shortlist of concrete suspects (a pronoun chain, a dense sentence, an unstated concept), not a generic verdict.
Then inspect the flagged spots yourself, in context. Check referents (“this,” “they”), identify the real main clause, and look for places where the text switches ideas without a signpost. Do a quick “could I underline the claim and evidence?” test; if you can’t, a reader likely can’t either. This is also where team constraints matter: decide what you will revise (split a sentence, add a heading, define a term) versus what you will support (a sidebar, a diagram, pre-teach a concept).
Finally, test changes with something observable: a short comprehension question set, a retell, or timed rereading plus error checks. If the model predicted the wrong sticking point, treat that as signal about your prompt or about the reader profile, and iterate.
Pairing model explanations with readability metrics and rubrics
A readability metric is still useful, just not as a diagnosis. Use it to spot outliers and track whether revisions moved the text in the intended direction, then use the model to explain what might be driving the score. If Lexile or a grade-level formula says a passage is “hard,” ask the model to attribute that difficulty to specific features: long sentences, low-frequency academic vocabulary, heavy nominalizations, weak connective tissue, or high concept density.
Rubrics make this safer. Pair the model’s notes with a fixed checklist your team trusts (syntax load, cohesion/referents, discourse structure, vocabulary/definitions, assumed knowledge, task alignment). Score each category yourself, then compare to the model’s claims. When they disagree, default to the rubric and the text you can point to. The practical constraint is time: rubric scoring is slower than running a metric, so reserve it for high-impact passages, not every worksheet.
Instead of “8th-grade,” you get “two sentences exceed working-memory limits,” “three key terms lack glosses,” and “the claim is implicit,” which can be revised or supported directly.
Prompts that reliably surface the real sticking points

You’ll recognize the unhelpful prompt because it produces a tidy list that could fit any passage (“define terms, shorten sentences”). Better prompts force the model to point to specific locations and make falsifiable claims. Ask it to mark the exact words or sentence numbers where a reader would likely reread, then explain the mechanism: “What is the referent of ‘this’ in sentence 4?” “Rewrite sentence 6 as two sentences without losing meaning, and explain what load you removed.” Require a before/after comparison so you can judge whether the change actually targets the issue.
Task-anchored prompts surface deeper friction than “is this hard?” Try: “A student must answer: [your question]. What parts of the passage are required evidence, and where might they miss it?” Or: “List the background knowledge assumed in each paragraph; label each as optional vs required.” Add a constraint that the model must cite only what’s in the text (“quote the phrase that triggers the assumption”). The practical cost is setup time: these prompts take longer to write, but they reduce time wasted chasing generic advice.
Common failure modes: confident nonsense, bias, and oversimplifying
You’ll still see failure modes even with good prompts. The most common is confident nonsense: the model picks a “likely” referent for this, invents a missing definition, or misstates the claim—then justifies it smoothly. Treat any explanation that can’t point to a specific phrase as unverified, and spot-check by paraphrasing the sentence yourself and asking, “What would a reader have to hold in memory here?”
Bias shows up when the model assumes what “typical” students know, often mapping background knowledge to a narrow cultural or socioeconomic norm. Require it to list multiple plausible reader profiles and to label assumptions as “text-required” versus “curriculum-dependent.”
Oversimplifying is the quieter risk: splitting every long sentence can erase precision, tone, or necessary logical links. The constraint is editorial time—keep a short “meaning preserved?” checklist before accepting revisions.
Using explanations to improve clarity without dumbing down
When a model flags “hard” spots, treat each note as a design choice you can adjust, not a mandate to simplify ideas. Often the win is structural: move the main claim earlier, add a one-sentence roadmap, or make the evidence-to-claim link explicit with a connective (“therefore,” “for example,” “in contrast”). Precision can stay intact if you replace vague academic nouns with concrete actions (“evaluation of” → “we evaluate”) and keep key terms, but add a tight gloss the first time they appear.
Clarity has costs. Adding definitions, headings, and examples increases length, and longer isn’t always easier—especially on small screens or timed assessments. Use the model to propose two alternatives (shorter vs more supported), then choose based on the reader’s task: answering questions, learning a concept, or skimming for a decision. The operating principle is to reduce avoidable load (memory, referents, missing links) while protecting necessary difficulty (new concepts, domain terms, and nuance).