From 8 Hours to 15 Minutes: How We Automated Release Notes with AI
Releasing software is only half the job. We built an AI-powered pipeline that handles the other half: delivering complete, accurate release notes to customers the moment every product ships, automatically.
Release notes are operational guidance for our customers. Fleet operators and mobility managers use them to understand what changed in the software running their business, prepare their teams, and decide whether a workflow needs updating. Done well, they build confidence. Done poorly, or delivered late, they generate support calls and erode trust.
For years, writing them was a manual task that fell on one product lead at the end of every sprint. Four to eight hours of reading tickets, pull requests, and support threads, then translating that into a spreadsheet that customers could act on. It worked. It didn't scale.
The real cost of manual release documentation
The process was straightforward and expensive. A product lead pulled the ticket list for the release, read each one, cross-referenced the linked pull requests, checked for related customer support threads, and wrote a plain-English summary of what changed and why it mattered. For a release with 40 tickets, that's a full working day. Ridecell ships across four product lines on a rolling cadence, which means this was happening constantly, competing with every other priority in the sprint.
When documentation quality depends on one person's bandwidth at the end of a sprint, it will always be the first thing that slips. Customers are the ones who notice.
The inconsistency was a real problem. Product lines with disciplined product leads produced detailed, well-categorized notes. Others shipped sparse summaries or delayed delivery entirely when the sprint was heavy. Customers receiving updates across multiple Ridecell products were getting different documentation quality for the same caliber of engineering work. That's a customer experience problem masquerading as a process problem.
- 1Pull the ticket list for the release from the issue tracker
- 2Read each ticket, pull request, and commit message
- 3Search for related customer support tickets manually
- 4Write a plain-English customer summary for each item
- 5Format into a spreadsheet and distribute manually
- 1Sensor detects a new release branch or tag in GitHub
- 2Issue tracker queried for the authoritative release list
- 3Customer support tickets ingested and linked
- 4Claude drafts summaries, impact statements, categories
- 5Spreadsheet delivered to shared drive, Slack notification sent
Four products. Four release models. One pipeline.
Ridecell ships from four separate codebases, each with its own engineering team and release convention. Our Shared Mobility Backend cuts dated release branches. Intelligence Services uses calendar-versioned tags. The iOS app tags builds with semantic version plus build number. Asset Management aligns releases to sprint tags. Each convention made sense for the team that adopted it, but each one requires a completely different detection strategy, diff-scoping approach, and connection to the issue tracker.
| Product line | Release signal | How changes are scoped | Cadence |
|---|---|---|---|
| Shared Mobility Backend | Dated release branch | Merged pull requests between branch cuts | Every ~2 weeks |
| Intelligence Services | Calendar-versioned tag | Commits between consecutive release tags | Monthly |
| Mobile (iOS) | Semantic version + build number tag | Commits within the same version series | Per build |
| Asset Management | Sprint-aligned tag | Commits between consecutive sprint tags | Per sprint |
What unified them was the output we needed to produce: a consistently structured, complete set of release notes for every customer, every time. The pipeline had to handle all four release models while producing a single, consistent format, and it had to be extensible enough that adding a fifth product line wouldn't require re-engineering the core.
The engineering decisions that determined output quality
Start with the issue tracker, not the AI
The instinct when building an AI documentation tool is to focus on the prompt. We learned quickly that prompt quality is downstream of data quality. The most consequential pipeline stage is the one that runs first: querying the issue tracker for every ticket tagged with the release's fix version. When that query returns a full, well-structured result (issue description, linked pull requests, customer impact fields, and an engineer-written release note), Claude has everything it needs. Output is accurate, detailed, and fast.
When that query comes back empty and the pipeline falls back to extracting issue identifiers from commit messages, coverage drops significantly. No prompt engineering closes that gap. For some product lines, the fix-version naming convention was already in place. For others, we spent time working with engineering teams to establish and document the convention before the pipeline could reach its full potential. That alignment work, not the software itself, was the hardest part of onboarding each product line.
Commit history is noisier than you think
Our Shared Mobility Backend uses pull requests as its unit of change. Each one is a structured, titled, reviewed artifact. The other three product lines diff raw commits. In any active codebase, the commit history is full of automated entries that have no business in a customer release note: dependency bumps, deployment triggers, CI image builds, config auto-updates. For our iOS product, the build system auto-commits a version bump on every single build. For Intelligence Services, infrastructure automation produces several commits per deployment that look meaningful but aren't.
Without filtering, those entries consume LLM context, inflate the output, and, critically, lower the confidence scores on rows that would otherwise be high-quality. We built per-product noise filters that silently drop these commits before any issue extraction runs. The filters also handle a subtler problem: tag-based repos use version strings that appear sortable but aren't. Sprint 9 sorts after Sprint 10 lexicographically. Build numbers must be scoped within a semantic version series, not compared globally. Getting the previous-release baseline wrong means the diff is wrong, and the release notes document a different set of changes than what actually shipped.
Confidence scoring keeps humans in the loop on the rows that need them
Every row passes through Claude on AWS Bedrock for drafting. Each one receives a confidence score, a 0–1 signal that reflects the quality and quantity of evidence behind it. A row with an engineer-written release note in the issue tracker, a linked customer support thread, and a structured pull request description scores near 1.0. A row sourced only from a sparse commit message scores well below 0.5. Rows below the threshold are amber-highlighted in the output for human review before publishing. The rest go out automatically.
The score creates a useful incentive: engineers who write their own release notes in the issue tracker see that text used as Claude's primary source material, with measurably better output quality on those rows. Product leads spend their review time on the amber rows that genuinely need judgment, not on re-reading 40 rows that were already high-confidence. A per-release cost guard halts drafting if spend exceeds a defined threshold, preserving partial output rather than failing silently.
- Only a commit message matched
- No engineer release note in the tracker
- No linked customer support tickets
- Sparse or missing PR description
- Issue description present
- PR title and body available
- No linked support tickets
- No engineer-authored release note
- Engineer-written release note present
- Issue description plus linked support tickets
- Structured PR with clear title and body
- Multiple corroborating evidence sources
What it looks like in practice
Since the pipeline went live, every Ridecell release across all four product lines has shipped with complete release notes delivered before the release itself goes out. Zero delays, zero gaps. That's the headline metric, but it understates what changed operationally.
Product leads spend their time on editorial decisions, not data collection. The draft arrives automatically. Amber-flagged rows, typically 15–20% of a release, get a human read. The rest are reviewed quickly and published. A task that consumed most of a day now takes under an hour end-to-end.
Customers receive consistent documentation across every product line. Before automation, the quality and depth of release notes varied by author and by how much time they had. Now every release follows the same structure: product area, category, customer impact, technical detail, links to the source issues and support threads. A fleet operator managing both the platform and the mobile app gets the same documentation quality for both, every time.
Security-sensitive changes are handled correctly by default. The pipeline filters out pull requests touching sensitive file paths before any drafting occurs. Those changes do not appear in customer-facing output regardless of what issue keys they carry. No manual triage required.
Customer PII and sensitive data never reach the release notes. The pipeline drafts from engineering artifacts only: issue metadata, pull request titles and descriptions, engineer-written notes, and support ticket references. No personally identifiable information or sensitive customer data is sent to the model or written into the output.
The first time the Slack notification came through on its own, a full draft, all four product lines, before anyone on the team had opened their laptop, it was genuinely surprising. We'd spent weeks building it and still didn't fully believe it until that moment.
Product Lead, Ridecell Shared MobilityWhat we got right, and what we'd do differently
- We should have fixed the issue tracker conventions before writing any pipeline code. The most time we lost in this project was spent on LLM prompt tuning for product lines where the issue tracker data was sparse. None of that time moved the needle. When we instead worked with engineering teams to establish fix-version naming conventions and encouraged writing release notes directly in tickets, output quality improved more than any prompt change ever did. Data quality upstream determines everything downstream.
- Noise filtering deserved the same design attention as the AI integration. We underestimated how much automated commit noise would affect confidence scores and output quality on commit-based repos. Getting the filters right required reviewing actual commit histories for each product line, not just applying a generic list. The two weeks we spent on this were among the highest-leverage two weeks in the project.
- Claude is doing language work, not reasoning work, and that's the right scope. The structured data (issue metadata, PR titles, support ticket counts, engineer notes) tells Claude what happened. Claude's job is to make it readable for a customer who wasn't in the sprint. That's a translation task, not a judgment task. When we gave it good structured input, the output was reliable. When we asked it to infer customer impact from thin commit messages, the output was correctly flagged as low-confidence. The confidence score is doing real work.
- Shipping amber-flagging before full automation was the right call. The team needed several release cycles to build trust in the output before we removed the manual review gate. Amber rows gave product leads a concrete set of items to focus on rather than feeling like they had to validate everything. The feedback from those early reviews directly shaped improvements to the scoring model, and by the time we removed the gate for high-confidence rows, there was genuine confidence that the output was correct, not just hope.
The next problem to solve
The draft lands in a shared spreadsheet today. That's useful, but it still requires a manual step to publish to the documentation that customers actually read. The next integration we're building connects the pipeline directly to Ridecell's customer portal: high-confidence rows publish on release day without human intervention, and amber rows stay in review until a product lead clears them. The goal is to close the gap between "the release shipped" and "customers have the documentation" entirely.
We're also building a visibility layer for release notes quality across product lines: a view that shows, over time, which areas produce consistently low-confidence scores and where the gap between what was built and what was documented is largest. A product area whose confidence scores trend downward over several sprints is signaling something real, usually degrading issue-tracker hygiene or a team that's moving faster than their documentation practices. Surfacing that as a trend, rather than as individual amber rows, gives engineering leads something to act on before it becomes a customer experience problem.
The pipeline is live today across all four product lines. Every release (Shared Mobility Backend, Intelligence Services, Mobile, and Asset Management) ships with complete, confidence-scored release notes delivered to the team before the release goes out.