AI Coaching at Scale: How to Build a Coaching Culture
Leadership DevelopmentL&DCoaching CultureArtificial Intelligence

AI Coaching at Scale: How to Build a Coaching Culture

Kontaim

Kontaim

@Argraide

Sep 4, 2026

At 9:07 on Monday, a newly promoted operations manager asks an AI coach to help with her first performance conversation. She provides the employee’s role, the missed handoff, and the outcome she wants. Five minutes later, she has three possible opening questions, a reminder to describe impact rather than label the person, and a plan for ending with one agreed action.

Useful? Yes. Coaching? Not yet.

A coaching culture is built when people routinely set a behavioral goal, try it in real work, get specific feedback, and adjust. AI can support the spaces between those moments. It cannot create the permission, trust, or managerial follow-through that make them happen.

That distinction matters as organizations try to provide personalized development to everyone, not only to senior leaders with a coaching budget. Corporate AI coaching can reduce the waiting time for practice and reflection. It can also make it cheap to produce bad advice, collect sensitive disclosures, and confuse activity with growth.

Here are five claims worth rejecting before an AI coaching program gets a bigger budget.

Myth 1: “Give everyone an executive coach for the price of software”

The expensive part of coaching is not access to words. It is diagnosis, context, challenge, and a relationship in which someone can notice patterns over time.

An AI model can respond to the context an employee supplies. It cannot know what was left out, whether the employee’s account is self-serving, or how a team has interpreted the same behavior for the past six months. Fluent language can create a false impression of informed judgment.

Research on human coaching is encouraging but narrower than many AI claims suggest. A 2014 review by Theeboom, Beersma, and van Vianen found positive effects of coaching on areas including performance, skills, well-being, coping, and work attitudes. That review supports coaching as a development practice; it does not establish that generative AI provides the same relationship or results.

The sensible use case is smaller: low-risk, repeatable support for a clearly defined behavior. An employee can rehearse a difficult opening, turn a broad ambition into a specific action, or reflect after a customer conversation. The system is much less suitable for deciding whether a conflict involves discrimination, illness, retaliation, or a serious power imbalance.

Build the development target before choosing the AI workflow. A useful development card contains three fields:

  • Situation: the next two one-to-ones with a direct report who brings vague problems.
  • Behavior: ask the direct report to propose two options before offering a solution.
  • Evidence: record whether the employee left with an owner and a next step.

That is far more workable than “become a more strategic leader.” AI can ask follow-up questions against the card, suggest a rehearsal, and prompt reflection after the meeting. A manager can then discuss what actually happened.

Can AI coach employees?

Yes, for structured reflection, rehearsal, and action planning. No, as the sole source of judgment, accountability, or support in high-stakes situations.

The unit that scales is not a personal coach in a box. It is a repeatable coaching loop with a human point of accountability.

Myth 2: “The more data AI sees, the more personal the coaching becomes”

Personalized development built from surveillance is usually development people will avoid.

A transcript of every message, meeting, and private reflection does not automatically create insight. It creates a large collection of material that may be incomplete, misinterpreted, or risky to retain. Employees also change what they say when they suspect a manager can inspect the record.

Self-determination theory, developed by Edward Deci and Richard Ryan, identifies autonomy, competence, and relatedness as important conditions for motivated behavior. A coaching system that removes autonomy by quietly monitoring employees may weaken the very motivation it claims to support.

Start with a minimum viable profile instead of a total employee record:

  • the employee’s declared development goal and role level;
  • the situations they have chosen to work on;
  • reflections they enter voluntarily;
  • the evidence they explicitly agree to discuss with a manager.

Consent should be specific, revocable, and meaningful. A box ticked because access to development depends on it is a weak form of consent. Employees should know what is collected, who can see it, how long it is retained, and whether it can affect promotion, pay, performance ratings, or redundancy decisions.

The counterintuitive design choice is to standardize more than expected. Use common rules for permissions, retention, escalation, and what counts as a development goal. Personalize the examples, sequence, language, and practice situations. Without shared guardrails, articulate employees often receive richer coaching because they give the system better prompts, while others are judged by thin or misleading data.

The NIST AI Risk Management Framework is useful here because it puts privacy, validity, reliability, fairness, and transparency in the same conversation. Do not allow an AI coach to infer personality, promotion readiness, protected characteristics, or mental-health status from conversational data. Those inferences are neither necessary for coaching nor safe foundations for employment decisions.

The evidence for AI coaching producing durable behavior change is still thin. Much of the public discussion concerns user reactions or short-term task completion, not whether a behavior survives six months later. More data is not a substitute for a better goal or a safer relationship.

This approach fails when employees believe their private reflections are visible to their boss, even if the formal policy says otherwise. Trust is part of the intervention. Treat it as an operating requirement, not a communications exercise.

Myth 3: “A chatbot and a prompt library create a coaching culture”

The distribution of a tool is not the creation of a social norm.

A coaching culture shows up in ordinary work: a manager asks what someone tried before giving advice; a project lead reviews a decision rather than assigning blame; an employee can request feedback on a specific behavior without making a grand declaration about personal growth.

The 70-20-10 model can be a helpful reminder that development happens through work and relationships as well as formal instruction. It is a prompt, not a scientific budget formula. The practical question is whether the job gives people a chance to practice and someone gives them useful feedback.

Baldwin and Ford’s 1988 model of training transfer made this point clearly: learning must generalize to the job and be maintained there. Manager support, opportunity to perform, and the work environment all affect whether a new behavior sticks. An AI coach can assist with preparation and reflection, but it cannot manufacture those conditions.

A simple operating loop can fit inside existing one-to-ones:

  1. The employee chooses one live situation and writes what they want to do differently.
  2. The AI coach asks for the relevant facts, challenges vague intentions, and helps prepare an if-then plan: if the conversation becomes defensive, then pause and ask for the other person’s view before responding.
  3. After the event, the employee records what they observed, what they tried, and what happened. No performance verdict is needed.
  4. The manager uses ten minutes of the next one-to-one to ask: “What did you try? What did you notice? What will you change on the next attempt?”

The GROW model—goal, reality, options, way forward—can keep this exchange from becoming motivational mush. AI is helpful when it turns each stage into a few precise prompts. It is less helpful when it generates fifty questions no one has time to answer.

The organization needs one shared protocol, not a hundred prompt packs. That protocol should define when coaching happens, what employees own, what managers review, and where a human specialist takes over. Once the routine is familiar, the AI can vary the practice to fit a new manager, engineer, recruiter, or project lead without changing the basic contract.

Myth 4: “Once AI is coaching employees, managers can stop doing it”

That is a category error. Coaching is partly a conversation, but it is also a leadership responsibility.

An AI system can generate questions, help someone rehearse a difficult message, and turn an employee’s reflection into a next experiment. A manager knows the team’s priorities, the history behind a disagreement, the consequences of a missed commitment, and the support the employee can realistically access.

Keep the division of labor explicit:

  • AI: reflection prompts, rehearsal, action planning, and reminders tied to an employee-owned goal.
  • Manager: agree priorities, observe work, challenge incomplete accounts, give feedback, and follow up on commitments.
  • Qualified human support: harassment, discrimination, self-harm, medical concerns, legal advice, whistleblowing, and disciplinary or termination decisions.

The evidence base for generative AI specifically is too young to support replacement claims. Even the positive research on coaching cited earlier concerns human-led coaching, not a chatbot operating without context or accountability.

The strongest early use may therefore be manager preparation rather than manager substitution. Before a one-to-one, a manager can review the agreed behavioral goal and prepare two observation-based questions. That is different from receiving a hidden sentiment score or an automated judgment about the employee.

Corporate AI coaching should make good managerial behavior easier to repeat. It should not give managers a reason to outsource difficult conversations or turn development into an automated assessment system.

Myth 5: “If people are using the coach, personalized development is working”

Usage is an exposure measure, not an outcome. A high number of sessions may mean people find the tool useful. It may also mean the system keeps producing generic advice that feels productive for ten minutes.

The Kirkpatrick Model offers a basic discipline: separate reaction, learning, behavior, and results. “Employees liked the coach” is a reaction measure. “Employees can describe the feedback principle” is a learning measure. “Managers observe a different behavior in live work” is closer to the outcome that matters.

What should a first pilot measure?

Run a contained six-week test with one role and one behavior, rather than launching a catalogue of leadership competencies. Twenty to thirty participants can be enough to expose workflow and trust problems, though not enough to prove broad business impact. Use a phased rollout or a comparison group only when withholding access is ethical and practical.

Track four things:

  • whether participation and drop-off differ by role, level, location, or access need;
  • whether development plans name an observable behavior, a real situation, and evidence of change;
  • whether the behavior is attempted and discussed in normal manager routines;
  • what the program costs in manager time, administration, escalations, and privacy concerns.

Capture a baseline before the first session. At the end, ask employees for one concrete example of changed behavior and ask managers what they observed. Review a sample of plans with a simple rubric. Do not treat chat volume, completion badges, or positive reactions as proof of transfer.

Peter Gollwitzer’s research on implementation intentions is relevant because specific if-then plans are more actionable than vague intentions. The same principle applies to measurement: “communicates better” is difficult to observe; “asks for the other person’s proposed solution before giving advice” can be checked in a real meeting.

Stop or redesign the pilot if employees cannot tell who sees their data, managers report extra work without better conversations, or the system gives inconsistent advice to comparable cases. There is no prize for scaling a bad coaching loop.

Start this week by choosing one role, one behavior, and five employees willing to test the routine. Give them a private, employee-owned development card; collect no transcripts; and check after two weeks whether anyone tried the behavior in real work. That small test will tell you whether you have the beginnings of a coaching culture—or merely a new place to ask for advice.