The problem: measuring activity instead of outcomes
Managers ask for productivity analysis and receive hour counts, task-created vanity metrics, or green burndowns that hide stagnant work. Teams game the numbers — log time on the wrong items, close tiny tasks for dopamine, leave hard work "in progress" for weeks.
Productivity is not hours logged. Sustainable delivery combines:
- Throughput — items completed per cycle with quality intact
- Flow stability — aging and WIP patterns, not spikes and stalls
- Estimate calibration — planned vs actual for learning
- Blocker resolution time — how fast impediments clear
AI can synthesize these signals from live PM data — if you refuse to reduce people to a single score.
Why hour-count dashboards mislead
Common traps:
Confusing presence with output — Long hours often signal overload or bad estimates, not heroism.
Ignoring work type — Research and incident response look "slow" in raw velocity.
Snapshot bias — Done count this week ignores queue buildup in Review.
Surveillance backlash — Weaponized time data destroys honest logging.
Good analysis uses time for calibration, throughput for trends, and aging for bottlenecks.
Framework: outcomes → flow → calibration → action
Structure team productivity review in four layers:
1. Outcomes — Releases shipped, milestones met, approvals cleared — linked to tasks, not slides.
2. Flow — flow_aging for stuck items; flow_cfd for column band trends; WIP concentration.
3. Calibration — time_summary: estimated vs logged hours by work type — retro input, not ranking individuals in group channels.
4. Action — Rebalance load (workload_heatmap), fix process (approval SLA), adjust scope — logged in decision record.
AI via MCP or reports can draft the synthesis; managers interpret with context humans hold.
| Signal | Question it answers | Misuse to avoid |
|---|---|---|
| Done throughput | Are we finishing committed work? | Punishing "low" weeks during research |
| flow_aging | Where is work stuck? | Blaming assignee without checking blockers |
| time_summary | Are estimates learning? | Individual surveillance |
| workload_heatmap | Who carries too much open work? | Ignoring skill fit |
Where WKFGo fits: honest productivity lenses
Reports and project analytics
get_project_report and in-product reports surface delivery health, overdue items, and portfolio context — evidence for manager syncs, not leaderboard games.
time_summary for calibration
time_summary compares logged vs estimated time. Use in retros to adjust defaults by work type. Teams that never weaponize time data collect more honest numbers.
Important: High logged hours on a saturated person often indicates overload, not high productivity.
Flow metrics
flow_aging highlights cards sitting too long in a column — standup fodder before dates slip.
flow_cfd shows band trends — thick Review means approval bottleneck, not "lazy developers."
MCP synthesis
Prompt connected assistants: "Summarize last sprint: completed tasks, top three aging items, estimate variance from time_summary — cite IDs."
Weekly ritual (15 minutes)
- Scan
flow_agingred items — unblock or reassign - Glance at
flow_cfd— any widening band? - Retro snippet from
time_summary— one estimate pattern to fix workload_heatmap— rebalance before next sprint commit
No individual scorecards in shared channels.
Getting started this week
Before retro, export one-page view: completed tasks last sprint, top three flow_aging items, one pattern from time_summary (e.g., backend estimates 40% low). Discuss as team — no individual rankings. Pick one process fix (approval SLA, WIP cap). Revisit next sprint whether throughput improved. Repeat; productivity analysis is a rhythm, not a dashboard install.
Executives should never receive individual hour leaderboards derived from this framework — aggregate team patterns only.
Anti-patterns
- Productivity = hours logged leaderboard
- AI narrative without task citations
- Ignoring quality and rework in "velocity"
- Using productivity analysis for layoff justification without human review
Productivity metrics that reward hours logged encourage performative work. Focus on throughput: tasks completed per week, estimate accuracy trend, and aging percentiles. time_summary plus flow views give managers honest signal—pair with conversation when numbers dip, not surveillance dashboards.
Compare team throughput before and after process changes using the same two-week window—avoid comparing holiday weeks to crunch weeks. get_user_performance supports fair narratives when data exists.
FAQ
Does WKFGo score individual productivity with AI?
No dedicated "productivity score" product — use reports, flow, and time_summary for team-level patterns; keep HR decisions human.
Do we need mandatory time tracking?
Remaining estimates on open tasks suffice for heatmaps; logging improves calibration over time.
Can AI compare teams fairly?
Only with caution — different work types and contexts make cross-team ranking misleading.
How is this different from executive_brief?
executive_brief serves officer decision lenses; this framework serves engineering and PM leads on delivery rhythm.
Measure what ships, not who looks busy
Use AI-assisted reports and flow signals to see real productivity — throughput, aging, and calibrated estimates — without reducing people to hours alone.
Try it now
Put these patterns on live project data—not slide decks.