District Pilots

How to run a six-week school AI pilot that produces a real decision

A pilot should not prove that an AI tool can answer questions. It should reveal whether the district can govern, teach with, support, and explain the workflow after the demo glow wears off.

By HonorlyAI Team · 2026-07-23 · 12 min read

Quick answer

A strong six-week school AI pilot begins with a narrow educational use case, approved users and data, baseline measures, trained teachers, assignment-level rules, family communication, and a named support and incident team. The pilot should include real classroom use, weekly evidence review, student and teacher feedback, documented problems, and a final decision with conditions for expansion, revision, or termination.

Define the decision before the pilot

Many pilots begin with access and end with anecdotes. The district distributes accounts, collects a few enthusiastic quotes, and calls the experiment successful because nothing caught fire. That proves only that the product opened.

Write the decision the pilot must support: whether to expand a specific use case to a defined group under stated conditions. Then identify evidence that would justify yes, no, or not yet. A pilot without a decision threshold is a product tour with homework.

  • Use case: what educational job is being tested?
  • Population: which grades, subjects, teachers, and students are included?
  • Scope: which tools, features, assignments, and data are permitted?
  • Decision: what could expand, change, pause, or stop after six weeks?
  • Owner: who signs the final recommendation?

Week 0: approve the conditions

Complete privacy, security, accessibility, legal, procurement, and instructional review before student access. Configure identity, roles, retention, data limits, teacher controls, support routing, and logging. Record the approved version and settings so later evaluation refers to the actual system used.

Select a small teacher cohort that represents more than early enthusiasts. Include different subjects, grade levels, confidence levels, and student needs. A pilot built only around the person who already runs an AI club will not predict ordinary implementation.

Week 1: teach the workflow

Teachers need a concrete launch session: what the tool is for, what it should not do, how to set assignment rules, what activity is visible, how to intervene, how to report a problem, and what students and families were told. Give them two or three ready-to-run activities rather than a catalog of features.

Students need the same clarity in age-appropriate language. Demonstrate permitted and prohibited requests, explain disclosure, identify private information they should not enter, show how to question an incorrect answer, and explain when to ask a human.

Weeks 2 and 3: observe normal use

Run the tool in real assignments where the learning objective is known. Avoid basing the entire pilot on optional exploration; voluntary use attracts the students and teachers already most comfortable with AI.

Collect evidence lightly but consistently. Weekly teacher notes can capture setup time, support requests, useful moments, harmful or incorrect output, workarounds, alert quality, and whether the workflow changed instruction. Student prompts and transcripts should be reviewed only within the approved visibility and privacy design.

1. Learning

Did students receive explanations, practice, feedback, or merely faster answers?

2. Teacher workload

What setup, review, intervention, and support time did the workflow create or save?

3. Integrity

Were assignment rules understood, and what misuse patterns appeared?

4. Access

Who could not use the workflow effectively because of disability, language, device, connectivity, or account barriers?

5. Reliability

What incorrect, unsafe, irrelevant, or inconsistent behavior occurred?

6. Governance

Could the district explain and respond to incidents using the documented process?

Week 4: test the edges

Do not wait for accidental failure to discover the limits. Use an approved red-team and scenario review: ambiguous assignment rules, attempts to obtain final answers, prompt injection in uploaded material, unsupported factual claims, harmful content, sensitive personal information, language variation, accessibility workflows, and account-sharing scenarios.

The purpose is not to attack the vendor theatrically. It is to learn how the configured service behaves, whether staff understand the response, and whether the district can document and correct problems.

Week 5: compare evidence to baseline

Compare pilot outcomes with the baseline defined before launch. That might include teacher response time, student help-seeking, completion of practice, revision quality, frequency of integrity concerns, support burden, or understanding on a common assessment. Do not invent a causal claim from six weeks of uncontrolled use.

Segment the evidence. Averages can hide that one grade or student group benefited while another faced access or error problems. Review teacher-level variation too; a workflow that succeeds only with extensive individual customization may not be ready for scale.

Week 6: decide in public language

The final report should state what was tested, who participated, settings, data limits, evidence, incidents, limitations, teacher and student feedback, unresolved risks, costs, and the recommended next step. Separate observed facts from interpretation.

Choose one of four outcomes: expand within the same scope, expand with conditions, revise and re-pilot, or stop. Expansion should have gates such as contract changes, feature fixes, training, accessibility remediation, or stronger metrics rather than a vague promise to monitor.

  • Name the approved scope and settings for any next phase.
  • Publish a plain-language summary for staff and families.
  • Assign owners and dates for every condition.
  • Preserve the evidence and incident record needed for renewal.
  • Schedule the next formal review before broader rollout.

The pilot is also a test of the district

A product may perform well while the district lacks training, support, communication, or decision ownership. The reverse can happen too: an organized team may reveal that the product cannot support the intended workflow.

The correct question is not only "Does the AI work?" It is "Can this school system use this service responsibly, instructionally, and sustainably under real conditions?"

Frequently asked questions

How long should a school AI pilot last?

Six weeks is often long enough to move beyond first-day novelty and observe repeated classroom use, while remaining short enough to limit exposure and make a timely decision. The right duration depends on the use case and school calendar.

How many teachers should participate in an AI pilot?

Use a cohort large enough to include varied subjects, grade levels, confidence, and student needs but small enough for close support and evidence review. Representativeness matters more than a vanity participant count.

What are the possible outcomes of an AI pilot?

Expand within scope, expand with conditions, revise and re-pilot, or stop. The final decision should name evidence, unresolved risks, owners, required changes, and the next review date.