我构建了一个真实虚假数据的来源。
1 分•作者: przeslijmi•大约 1 个月前
大家好。
我想向大家介绍一个我过去几个月一直在开发的工具——real-fake-data generator。
https://real-fake-data.com
它的基本功能是提供测试数据,这些数据(1)与真实数据相同但并非真实数据,(2)确保您的管道能够捕捉到真实世界的问题,而不仅仅是开发人员在开发时会想到的那些问题。
首先,现实是不断变化的。您可能有一个包含车辆注册号的提交表单,但您会遇到问题(不同国家,不同规则)。您要么花一个月的时间去挖掘各种场景,要么就放任不管,只设置一个“最大长度为7”,然后走开。您的工具运行良好——但总有一天,某个国家或地区决定车辆牌照的长度为8,然后……您的系统就会失败(用户无法输入新的车辆牌照),而您的测试管道却运行正常,并向您发送绿色的笑脸。
其次,您确实无法预测一切。您面向全球用户?很好,但他们使用不同的字母表。您精美的用户体验/用户界面可能会因为用户姓名变成4到5个字母而崩溃。或者姓名字段太短。它会接受即使在欧洲使用的非拉丁字符吗?
当您创建新软件时,您还需要一些初始数据——这样您就可以(在开发过程中)看到您的工具拥有一些模拟用户、模拟帖子、模拟产品、模拟交易等。同样,您要么花很多时间编写这些数据,即使有 AI 的帮助——要么您就只留下 user1、user2、user3,从而失去以真实用户视角审视您工具的能力。
我的工具——real-fake-data.com——解决了所有这些问题。
- 您想接受德国身份证件、美国车辆号、西班牙 18 岁以上人士、波兰真实存在的地址——我们有超过 300 个真实数据生成器(真实存在,校验和保证,遵循所有规则)。
- 您想创建一个包含 30 个用户、每人几笔订单、每笔订单都有付款详情和日志的种子数据库——具有有意义的真实时间戳——只需定义数据模式——您就搞定了——数据已生成。
- 您想在恶劣的环境中运行测试——我们也能做到——对于每个生成器,您可以从正常模式切换到:EDGE(正确的数据,但处于正确性的边缘——例如,出生日期是昨天,最长的姓氏,波兰最短的车辆号),EXTREME(正确的数据,但故意制造问题,例如未修剪的空格、换行符、隐藏的 UTF 字符等),以及 INVALID(不正确的数据,可用于检查您的表单是否会拒绝它或是否能正常处理)。
- 您想使用相同的数据重新运行测试——这可以通过 SEED 号码实现,您将始终获得相同的随机数据——因此,您可以选择何时需要随机性,何时需要重复的安全性。
- 您希望您的 Claude 以低成本为您提供数据?很高兴听到——MCP 服务器已准备好供您使用。
- 您想轻松编写 Playwright 测试,例如 `const person = await fakeData.plPerson({ sex: 'f' });` 并且您就搞定了。
- 您想使用 VSCode 插件,直接在浏览器中获取数据而无需离开?很酷——Ctrl+Shift+P “生成假 UUID”——您无需离开 VS Code 屏幕即可获得所需数据。
- 您想确保任何数据都不会使您的软件面临风险——很好——确保种子是随机的,打开边缘模式,您就可以确信,一旦真实世界出现新的数据格式,您的管道就会针对它们进行测试。
它的成本是多少?简单使用是免费的。没有隐藏费用,无需信用卡,没有任何形式的月度订阅。每月免费 2000 个 token。
我非常希望收到更多见解,并听到您对该工具的看法。
它还附带 MCP 插件、Playwright 插件和 VS Code 扩展,让您无需离开 IDE 即可获取任何数据。
此致,
Karol Nowakowski
查看原文
Hello guys.<p>I want to interest You guys in a tool I've been working on for a last few months - real-fake-data generator.
https://real-fake-data.com<p>The basic need it covers is to deliver test data that is (1) identical to the real data but not real, (2) make sure Your pipelines will catch real life problems - not only those which developer will remember to cover when developing.<p>First of all - reality changes. You can have a submission form with vehicle registration number and You have a problem (different countries, different rules). You will either spend month on digging various scenarios or You will just let it go and put a "max length 7" and go away. Your tool is up and working - BUT someday some country or state decides to have lenght-8 for vehicle plates and .... your system fails (users cant input new vehicle plates) while your test pipelines work and are sending You fake-green smile.<p>Second of all - You really cant predict everythin. You are open for users from around the world? Nice but they use different alphabets. Your beautiful UX/UI will crash with user initials becoming 4 or 5 letters. Or the field for surname will be just too short. Or will it accept non-latin characters used even in Europe?<p>When You create new software You also need a seed starting data - so You can see (while developing) Your tool with some mock users, mock posts, mock products, mock transactions etc. Again - You either spend a lot of time writing it - even with help of AI - either You just leave user1, user2, user3 and You loose ability to look at Your tool in a way real life user will be looking.<p>My tool - real-fake-data.com - fixes all of it.
- you want to accept german ID document, USA vehicle number, spanish person over 18 y.o., real existing address from Poland - we have over 300 generators of real data (really existing, checksum guaranteed, following all the rules)
- you want to create a seed database of 30 users, each with few orders, each with payment details, and logs - with REAL timestamps that make sense - just define data schema - and You are covered - data is generated
- you want to run Your tests in a hostile environment - we got this - You can switch for every generator from normal mode into: EDGE (correct data but on edge of correctness - eg. born date yesterday, longest possible surname, shortest vehicle number from Poland), EXTREME (correct data but deliberately made problematic with spaces untrimmed, new lines, hidden UTF characters, etc.) and INVALID (incorrect data that is useful to check if Your form with reject it or behave correctly)
- you want to rerun the test with THE SAME data - it's there with a SEED number You will always receive the same random data - so You choose whenever You want random things and whenever You want repeated security
- you want Your Claude to deliver You data on low cost? Great to hear - MCP server is ready for You to use
- you want to write playwright tests easily nice `const person = await fakeData.plPerson({ sex: 'f' });` and You are covered
- you want to use VSCode Addon to have test data directly on the browser without leaving it? Cool - Ctrl+Shift+P "generate fake UUID" - and You have it ready without leaving the VS code screen
- you want to be sure that none data will ever put Your software at risk - great - make sure seeds are random, turn on edge mode and You will be sure as soon as some new formats of data will be there in real world Your pipeline will be tested against them.<p>How much does it cost? For simple using it costs Nothing. No hidden fees, no credit card required, no monthly subscription of any kind. 2000 free tokens/month.<p>I would really love to receive more insight and hear Your opinion on the tool.<p>It also have MCP addon, Playwright addon on VS Code extension that allows You to grab any data without ever leaving the IDE.<p>with regards,
Karol Nowakowski