HN 提问:在代码审查之前,你会如何加固一个拥有百万行代码的遗留 SaaS 应用中的 AI 变更?
2 分•作者: thegreatkahuna•2 个月前
我并非软件工程师,但我一直在进行一项实验,以评估基于现有SaaS代码库的代理式开发(agentic development)能否产出一个有用的原型。
该代码库拥有超过100万行代码,已有15年历史,托管在Azure上,主要使用C#和React编写。
原型需要在9月份提供给客户进行测试。虽然我偶尔可以获得特定技术问题的帮助,但没有全职开发人员可以投入。一位工程师将在8月份评估实现情况,并决定我们对AI生成的代码有多大信心,以便我们决定如何将代码“转换”为生产级别。
我的问题是,在8月份的评估之前,我主要利用AI工具和手动测试,可以做些什么来使代码尽可能健壮和易于审查?我希望提高AI生成代码的质量,使其在转换为生产级别时,更接近于“复制粘贴”,而不是从头开始重建。
环境
原型正在一个单独的分支中开发,部署到一个单独的内部环境中,并连接到其自己的数据库和模式。
规划流程
我们在6月份访谈了客户,并将结果转化为MVP(最小可行产品)规范。
流程大致如下:
1. 编写PRD(产品需求文档)。
2. 使用LLM(大型语言模型)将PRD转换为架构文档,并由架构师进行审查。
3. 使用Claude Design创建产品设计,包括屏幕图像和包含交互细节及边缘情况的md文件。
4. 使用一个代理,以PRD和架构文档为指导,将工作分解为史诗(epics)。
史诗是我们最细粒度的规划产物,由我和架构师进行审查。
开发流程
开发流程旨在尽量减少干预:
1. 一个规划代理将史诗转换为md故事文件和Jira故事。
2. 一个编码代理实现故事,包括测试,并创建PR(拉取请求)。
3. 一个审查代理审查PR,请求修改,并将它们合并到原型分支。
编码代理会轮询PR以获取审查意见,并将问题反馈给规划代理。
代理还有一个“停止并询问”列表,用于它们不允许自主做出的决策。我(有时还有工程师)通过解决这些升级请求以及每天手动端到端测试累积的更改来参与其中。
大部分实现和初步审查都是使用基于Claude的代理完成的。对于风险较高的PR,我也使用了Codex作为审查者。第二次模型审查发现了Claude生成的代码中更多相关的实际问题,但令牌配额限制了其使用。
我还对代码进行了单独的重构和加固。
结果
规划以及环境和代理流程的设置大约花费了两周时间,然后代理在约两周内构建了整个MVP。从代码量上看,功能代码有13,000行,测试代码也占了相同数量。
我希望获得的建议
假设在审查之前我无法获得实质性的开发人员参与,我如何才能最大程度地提高代码接近生产级别的可能性?
以下是我一直在思考的一些问题:
1. 在工程师审查代码之前,哪些检查或开发循环可以带来最大的信心提升?
2. 如何使用独立的代理或模型来降低编码者和审查者做出相同错误假设的风险?
3. 测试是否应该由一个与编写实现的代理不同的代理来生成?
4. 哪些文档或证据可以使最终的工程审查更快、更可靠?
5. 如果只有几周时间来改进这个原型,然后将其交给工程师,我会优先考虑什么?
查看原文
I’m not a software engineer, but I’ve been running an experiment to see whether agentic development could produce a useful prototype on top of an existing SaaS codebase.<p>The codebase is 1M+ lines, 15 years old, hosted on Azure, and primarily written in C# and React.<p>The prototype needs to be available for customer testing in September. No developers were available to work on it full-time, although I could occasionally get help with specific technical issues. An engineer will evaluate the implementation in August and decide how much confidence we can have in the AI-generated code so that we can decide how to “convert” the code to production-grade.<p>My question is what I can do before August, primarily using AI tools and manual testing, to make the code as robust and reviewable as possible. I want to increase the likelihood that the AI-generated code would be so good that the path to production-grade would be closer to “copy-paste” than building everything again from scratch.<p>Environment
The prototype is being developed in a separate branch, deployed to a separate internal environment, and connected to its own database and schema.<p>Planning process
We interviewed customers in June and turned the resuls into an MVP spec.<p>The process was approximately:
1. Write a PRD.
2. Use an LLM to convert the PRD into an architecture document, which was reviewed by an architect.
3. Create product designs consisting of screen images and md files containing interaction details ans edge cases with Claude Design.
4. Use an agent to break the work into epics using the PRD and architecture document as guardrails.<p>The epics were the most granular planning artifacts that received review by me and the architect.<p>Development process
The development flow was intended to run with little intervention:
1. A planner agent converted epics into md story files and Jira stories
2. A coding agent implemented the stories including tests and opened PRs
3. A reviewer agent reviewed the PRs, requested changes, and merged them into the prototype branch<p>The coding agent polled PRs for review comments and could escalate issues back to the planner.<p>Agents also had a “stop and ask” list for decisions they were not allowed to make autonomously. I (and a few times an engineer) were involved by resolving those escalations and by manually testing the accumulated changes end to end each day.<p>Most implementation and initial review were done with Claude-based agents. For riskier PRs, I also used Codex as a reviewer. The second-model review found substantially more relevant issues in the Claude-generated code, but token quotas limited the use.<p>I had separate refactoring and harden runs for the code as well.<p>Results
The planning and setting up the environment and agentic flow took about two weeks and then the agents built the whole MVP in about two weeks. Size-wise it was 13k lines of functional code + the same amount for tests.<p>What I would like advice on
Assuming that I cannot get substantial developer involvement before the review, how can I increase the likelihood that the code is as close to production-grade as possible?<p>Here are some of the questions I have been thinking about:
1. What checks or development loops would give the largest improvement in confidence before an engineer reviews the code?
2. How would you use independent agents or models to reduce the risk that the coder and reviewer make the same incorrect assumptions?
3. Should tests be generated by a separate agent from the one that wrote the implementation?
4. What documentation or evidence would make the eventual engineering review faster and more reliable?
5. If you had only a few weeks to improve this prototype before handing it to an engineer, what would you prioritize?