Show HN:BlazeRules – 用于流式数据的 YAML 规则引擎,每秒可处理高达 500 万条记录
1 分•作者: jspuri•大约 2 个月前
<a href="https://github.com/purijs/blazerules" rel="nofollow">https://github.com/purijs/blazerules</a><p>我最初想用 C++ 编写一个亚毫秒级的日志解析器,但后来演变成了一个可嵌入的决策引擎,可以对传入数据运行 YAML 定义的规则。
这些规则通过先重新投影为列式格式(如果尚未是列式格式)来以向量化格式在传入数据上执行。根据载荷大小和规则的复杂性,性能从每秒 200,000 条记录到每秒超过一百万条记录不等,吞吐量平均约为每秒 200 MiB 到每秒 3 GiB。</p><p>规则也可以是 SQL 表达式,或者 ONNX 模型(数值型)、窗口操作以及支持的许多其他操作。</p><p>它与 DuckDB 类似,但适用于流式数据和即时决策。</p>
查看原文
<a href="https://github.com/purijs/blazerules" rel="nofollow">https://github.com/purijs/blazerules</a><p>I initially wanted to make a sub-millisecond log parser in C++ but that blew into a embeddable decision engine, that can run YAML defined rules on incoming data.
The rules are executed in a vectorized format on incoming data by reprojecting into a columnar format first, if it's not already. Depending on the payload size and rules complexity, the performance goes from 200K records/s to more than million records/sec, in terms of througput this would be around 200 MiB/s to 3 GiB/s on average.<p>Rules can be sql expressions too, or onnx models (numeric), window ops and quite a few more operations are supported.<p>It's comparable to DuckDB but for streaming data and on the fly decisions.