How Aidealy turns engineering activity into answers
Aidealy reads the signals your team already produces as they build - then connects them into something you can actually ask questions of.
A bug that breaks production this quarter was probably written weeks ago - by a person or an AI agent. You could dig through Git by hand and trace that one bug back to the pull request that introduced it, its author, and its approver.
What no one has time to do by hand is that - continuously, across every contributor and every correction, stored and rolled up into an overview you can actually act on.
That's what Aidealy is for. It starts with the work your team already does, in two places:
From there, Aidealy connects everything into a queryable knowledge graph of your R&D:
An editor session is tied through your Git history to the branch it touched and the pull request it became, so a question about what someone was working on and a question about what shipped are answered from the same connected picture.
The method
Here's how Aidealy measures R&D productivity: tickets are never fully up to date, so they can't be the measure of real work.
Code is the only honest record of what actually shipped - so that's what Aidealy measures.
(Aidealy complements your tracker; it doesn't replace it.)
Aidealy treats every commit merged into a repository's default branch as having reached production. It's worked out from your Git history - not by connecting to your CI/CD or production systems, and not by reading production logs.
Every shipped line is labeled AI or human. Every line changed in a pull request that merges to production is accounted for. For the code your people committed, the label is AI-written or human-written. Not a detector's guess - a record:
That per-line honesty is what lets the quality lens below compare AI-written and human-written code.
Every PR carries its real cost: hands-on hours. Cycle time measures how long a PR waited; hands-on hours measure how long people actually worked on it - recorded from the editor, never estimated, a floor, never a census.
Like tokens, hands-on hours never feed the rank - they're what a PR costs, not a measure of worth.
Aidealy turns that activity into one score per contributor - the Aidealy Productivity Rank - over rolling 7-, 30-, and 90-day windows. It's built from four signals:
work merged to production (not lines of code), weighed by its quality impact, with a penalty when it had to be corrected soon after merge. The faster something needs fixing, the larger the penalty - so code that lasts is rewarded. Tests and infrastructure-as-code count too.
how much you review and how carefully: substantive review is rewarded; sitting on a pull request, or approving code that then broke, counts against you.
working across more languages and more repositories.
how effectively you put AI to work in Cursor - for example, running several AI sessions in parallel rather than one at a time. A minor, deliberate signal: it surfaces a genuinely new way of working, but what ships and its quality always lead.
Analyzed your engineering data · 7/7 steps
Productivity Rank · last 90 days
Productivity score · by engineer
90-day window · signed, uncapped
A negative score means cleanup outweighed shipped work.
Signed, and uncapped. The rank is signed - it can go negative when a contributor's output is outweighed by the cleanup it caused. And it has no ceiling: there's no "100" to max out, so there's always headroom, and the rank never leans on a benchmark that means something different on every team.
Our formula, transparent inputs. Those four signals are on the table - we show you exactly what feeds the rank, and what never does. The way they're weighed and combined into a single signed score is our own formula. We show you what we measure; the formula that turns it into a score is ours.
Tokens and hours never feed the rank. Token usage is tracked by model, by engineer, and per pull request - and hands-on hours per contributor on that same pull request; you can ask about both, but neither ever enters the score. They're what a PR costs, not a measure of worth. A token count is a fuel bill, not an achievement.
Where measurement stops: business value is your call. Aidealy deliberately doesn't score the business value of a PR or a ticket. Value is a judgment - the same feature can be critical when one of your customers demands it and minor when another does - and judgments don't belong in a formula. You decide what's worth building when you prioritize; Aidealy measures how well the prioritized work shipped, its quality, and its cost. We measure the execution; you judge the worth.
Four lenses on R&D productivity and code quality
Each lens answers a real question on its own. Three of them also feed the Productivity Rank; the fourth stands deliberately apart from it.
The lens most teams have never had. How long your code survives in production, and where a bug actually came from.
A single 0-100 score for your actual source code. Measured against the full Clean Code discipline - every rubric, not just complexity or lint.
Every AI-assisted session, classified automatically. See what your team uses AI for and what you pay for it, without anyone writing a status report.
Reviewing is real work, and Aidealy measures it. On every merged pull request an engineer reviewed:
How the lenses relate to the rank: three of them - whether your code lasts, how good it is, and how your team reviews - feed the Productivity Rank. The fourth - what your team uses AI for, and what it costs - stays separate from it by design: it's there to inform you, not to rank anyone. Each lens earns its place on its own.
The score
The Aidealy Code Quality Score is a 0–100 built on the Clean Code discipline. The AI judges each open rubric; a fixed formula of ours computes the final score.
Every file is judged as what it is - production code by the Clean Code principles, test code by the discipline's own test rules.
Every rubric is on the table - naming, function size, complexity, error handling, duplication, and the rest. Nothing that feeds your score is hidden.
The final 0–100 comes from a fixed, weighted formula, never from an overall impression - repeatable, not a one-off opinion.
The full method, in plain Q&A, is on the FAQ.
Same principle as the rank: open rubrics, our formula. It's the Aidealy Code Quality Score - our scoring formula, built on Clean Code - and the why behind any number is one question away.
(We didn't make these principles up - they're the discipline Robert C. Martin set out in his book Clean Code. Aidealy isn't affiliated with or endorsed by him.)
We publish our rubrics and every signal that feeds a score, for transparency; the method is versioned and improves over time. Scores aren't guaranteed comparable across methodology versions.
Ask about engineering productivity in plain English
All of it - the signals, the scores, the lenses - turns into answers to questions you ask in plain English, for one person, a team, or your whole R&D organization.
The point of connecting all of this isn't another dashboard to interpret. It's that you can ask the questions you actually care about and get an investigated answer back - a ranking, a trend, a correlation - not a wall of charts.
Aidealy is built for the aggregate, altitude question - the one you couldn't answer by hand:
Rank my R&D workforce by real productivity this quarter, grouped by team.
Which engineers caused the most critical bugs in production last quarter - and who approved them?
Which repos improved or declined in code quality this quarter?
Who reviews the most - and who's letting pull requests sit unreviewed?
How many tokens did each PR cost us, by model - and did the priciest ones ship code that lasts?
A rank or a chart is where you start: the why is one more question away.
Aidealy is in early access - what you see here is how you'll work with it.
Trust
Three principles hold the whole method together:
The ranking is the start, not the end. A score is an invitation to look closer, never a final verdict on a person. The why is always one question away.
Production is the unit of truth - for better and for worse. What never ships doesn't count; what breaks production counts against you. The rank rewards real contribution and subtracts for the rework left behind, which is why it can go negative.
We don't rank by tokens, lines of code, or tickets closed. Those are easy to count and easy to game. Aidealy measures contribution to production, and its quality - the things that actually matter.
You can't improve what you can't measure. Aidealy gives you the measurement - clearly, fairly, and with the evidence right behind it.
Bring the one you most want answered - about your engineers, your AI agents, or the code they ship together. Once you're set up, Aidealy answers it on your own R&D, with the evidence right behind it.
Setup is light - connect your Git provider, roll the extension out to your Cursor users, and Aidealy backfills your Git history (a year by default; longer available as a paid add-on).
No dashboards to configure, no queries to write. Then you just start asking.
Talk to usTalk to us - a real person walks you through what Aidealy measures, then gets you set up on your own R&D. No sales sequence.