Friday, October 2, 2026
Issue 10 of this newsletter sat finished for seventeen days. Proofread, SEO fields filled in, thumbnail attached, every link opened and checked. It sent on Friday, September 18 at 6:00 AM, and within a few hours I found a defect that seventeen days of review had not caught: every section header rendered near-black on a dark indigo background, so no one could read them.
The defect was in the draft the whole time. Reviewing it was never going to surface it. Only sending it would.
Reader Question
No subscriber sent me this one. It is the objection I hear most often when a district is a few months past a launch, so I am answering it the way I would answer it in a meeting.
"We ran an AI pilot last spring. How do I tell whether it worked?"
Most of the time the answer is that you cannot, and the reason is not the pilot. It is that nobody wrote down the before.
This is the piece of my 30/60/90 AI Implementation Plan that gets skipped most often. The first thirty days are Foundation, and one of the two exit criteria is a documented baseline: how long the task takes now, how many people touch it, how often it gets redone. That number is boring to collect and it is the only thing that makes day 90 answerable. Without it, the evaluation turns into a survey of how people feel about the tool, and people feel good about new tools for about a quarter.
So if you are sitting on a spring pilot with no baseline, here is what I would do rather than write it off. Pick the single workflow the pilot touched most. Measure it now, the way you should have measured it in March: time one full cycle, count the hands it passes through, count the reworks in a two-week window. That is your baseline, taken late. Then run the same measurement in six weeks with the tool still in place, and you have a real before and after, just shifted.
You lose the spring semester as evidence. You do not lose the pilot.
And when a team tells me the tool is working, I ask what it is working compared to. If the answer is a feeling, we go find the number first.
If you do not have a baseline yet, the Invisible Work Audit is the thing I built to find one.
Worth Your Time
Guidance on Responsible Use of Education Technology in the Classroom by the U.S. Department of Education, Office of Elementary and Secondary Education (Dear Colleague Letter signed by Assistant Secretary Kirsten Baesler, August 20, 2026)
The letter puts three questions in front of every technology decision: does it have a clear instructional purpose, is it supported by evidence, and does it improve student outcomes. What I find useful is what it declines to do. It does not tell districts how much screen time is too much, and it says the decision belongs at the state and local level. That puts the burden of proof back on the people buying, which is where it belongs and where most districts have the least process. If you are heading into a purchase with those three questions and no way to answer them, the District AI Pilot Checklist is the sequence I use.
Seven Washington Districts Chosen to Pilot AI-Enabled Data Solutions by the Center on Reinventing Public Education, with Education First and StrategicEDU (August 26, 2026)
Seven districts, up to $45,000 each, and the target is not the classroom. It is the plumbing: siloed systems, interoperability, scattered data on attendance and behavior and achievement. CRPE says each district will document what works, what does not, and what made responsible adoption possible. That last part is the reason I am linking it. Most AI pilots in education produce a decision and no record of how it was reached, so the next district starts from zero. If you want to know where your own district sits before you commit to something like this, the AI Readiness Assessment takes about fifteen minutes.
Measure ed tech by more than screen time, American Psychological Association says in K-12 Dive, reporting on new guidance from the American Psychological Association
The APA's argument is that minutes on a screen is a weak measure on its own, and that what matters is what the student is doing, what the screen replaced, and whether it helped or hurt the learning. Its position on generative AI is more cautious: extra guardrails, age-appropriate ones. I keep coming back to the first half. Screen time is the metric districts reach for because it is the one that counts itself. The measure that would tell you something takes work to define, which is why it usually does not get defined. If you want a measure that is harder to collect and worth more, start with the Invisible Work Audit.
Try This
Send one thing to one real person this week.
Pick the AI-assisted output you have been drafting and not sending. Everyone has one. A parent-facing summary, a staff memo about the new tool, a one-page guidance sheet, a template you built in July and have been polishing since. Something that is finished by every measure except the one that counts.
Send it to exactly one real recipient. Not a colleague who will be kind about it. Someone who has to use it.
Then collect three things, and give yourself ten days to collect them:
What did they do with it? Not what they said about it. Did it get forwarded, filed, printed, acted on, or ignored? Behavior is the measurement.
What did they ask you about? Every clarifying question is a defect report. Write the question down in their words, not your summary of it.
What did they get wrong? This is the one that pays. If a reader misread a section, the section is wrong, and no amount of rereading it yourself will show you that. You wrote it, so you already know what it means.
Then fix exactly those three things and send it to five people.
Here is why I run it this way rather than circulating a draft for feedback. Feedback on a draft tells you what people predict they would do. A send tells you what they did. Those are different measurements, and only one of them is about your actual reader.
The cost of the one-recipient test is one recipient's mild inconvenience. The cost of skipping it is the version I shipped in September, where a color value nobody had looked at made every heading invisible for 38 people at once.
The thought to bring to your next meeting: if the only way to test the work is to ship it, what are we holding? To make that a habit rather than a resolution, the District AI Pilot Checklist builds the send date in before the build starts.
That's it for this one. If something here was useful, share it with someone who'd get something out of it. If it wasn't, that's part of the deal too. You can find more of what I'm working on at evalveconsulting.com, or book a call if you want to talk through what you're dealing with.
Talk soon,
Chris