Skip to content

[New Mission] Operative NextGen: Obtain user feedback and evaluate #3154

Description

[New Mission] Operative NextGen: Obtain user feedback and evaluate

Summary

Author the continuous-improvement mission using the modern Copilot Studio experience. Learners enable built-in reactions, create representative test sets, run evaluations, compare agent versions, inspect Monitor, and turn evidence into a controlled instruction or capability improvement.

This replaces the classic topic-based feedback implementation with native feedback, Evaluate, version comparison, and Monitor workflows. Duplicate checks found no issue proposing this complete Operative NextGen mission.

Proposed mission

Field Value
Section Operative NextGen
Folder docs/operative-v2/09-user-feedback-evaluation/
Title Mission 09: Obtain User Feedback
Operation codename OPERATION ECHO
Difficulty 3
Time 60 minutes

Scenario

The hiring system is functional, but the team needs evidence before broader release. The learner collects reactions, creates test cases representing core hiring tasks and safety boundaries, compares a baseline with a revised version, and uses Monitor to investigate behavior rather than relying on anecdotal success.

Use a flight-recorder analogy: reactions reveal user sentiment, evaluations run controlled checks, and Monitor shows what actually happened during execution.

Mission objectives

  • Enable and validate built-in user reactions.
  • Build balanced evaluation test sets.
  • Run and interpret evaluations.
  • Compare baseline and revised agent versions.
  • Investigate conversations and tool use in Monitor.
  • Apply and document one evidence-based improvement.

Prerequisites to state in the mission

  • Completion of Missions 01–08.
  • Published Hiring and Interview Prep agents with sufficient test activity.
  • Access to Feedback, Evaluate, version history, and Monitor features available in the tenant.
  • Fictitious test data that contains no real candidate information.

Proposed lab outline

Enable and collect reactions

Verify User feedback under Safety & access, publish if required, generate controlled interactions, submit positive and negative reactions, and explain the limits of reaction data.

Build test sets

Create versioned test sets for instructions, routing, resume extraction, matching, record creation, document output, workflow behavior, MCP scheduling, safety, and failure handling. Define expected outcomes and evaluators without encoding brittle exact wording.

Run and compare evaluations

Run the baseline, make one documented change, create or publish the revised version as required, rerun the same tests, and compare aggregate and case-level results for regressions.

Investigate Monitor

Inspect relevant sessions, tool calls, connected-agent routing, errors, latency, and user reactions. Trace one failure to a concrete root cause and document the remediation and residual risk.

Acceptance criteria

  • Use native reactions; do not recreate classic feedback Topics.
  • State what reaction data captures and what it cannot prove.
  • Build test sets covering all nine mission capabilities plus safety and failure paths.
  • Use fictitious data and exclude personal information from reusable test sets.
  • Define expected behavior and appropriate evaluators for each case.
  • Preserve the same test set when comparing baseline and revised versions.
  • Report aggregate results, individual failures, regressions, and nondeterministic cases.
  • Use Monitor evidence to explain at least one failed or slow interaction.
  • Make one scoped improvement and demonstrate that it does not regress critical safety tests.
  • Preserve <mission-meta /> and the Operative Mission 09 analytics tag.
  • Follow repository writing and validation standards.

Out of scope

  • Custom Application Insights telemetry.
  • Model retraining from user feedback.
  • Production-scale analytics or Power BI dashboards.
  • Storing free-form candidate feedback in a new custom system.

Risks

  • Evaluate and Monitor availability can vary by tenant rollout.
  • Small reaction samples can produce misleading conclusions.
  • Exact-match tests are brittle for generative output.
  • Version comparison can be confounded by model updates.
  • Monitor records can contain sensitive candidate or calendar data.

References

Metadata

Metadata

Labels

feature-requestThere is a request for a new feature or an improvement to an existing feature

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions