Ratatool

一个用于数据采样、数据生成和数据区分的工具。「A tool for data sampling, data generation, and data diffing」

Github星跟踪图

Ratatool

CircleCI
codecov.io
GitHub license
Maven Central
Scala Steward badge

A tool for random data sampling and generation

Features

  • ScalaCheck Generators - ScalaCheck generators (Gen[T]) for property-based testing for scala case classes, Avro, Protocol Buffers, BigQuery TableRow
  • IO - utilities for reading and writing records in Avro, Parquet (via Avro GenericRecord), BigQuery and TableRow JSON files. Local file system, HDFS and Google Cloud Storage are supported.
  • Samplers - random data samplers for Avro, BigQuery and Parquet. True random sampling is supported for Avro only while head mode (sampling from the start) is supported for all sources.
  • Diffy - field-level record diff tool for Avro, Protobuf and BigQuery TableRow.
  • BigDiffy - Scio library for pairwise field-level statistical diff of data sets. See slides for more.
  • Command line tool - command line tool for local sampler, or executing BigDiffy and BigSampler.
  • Shapeless - An extension for Case Class Diffing via Shapeless.

For more information or documentation, project level READMEs are provided.

Usage

If you use sbt add the following dependency to your build file:

libraryDependencies += "com.spotify" %% "ratatool-scalacheck" % "0.3.10" % "test"

If needed, the following other libraries are published:

  • ratatool-diffy
  • ratatool-sampling

Or install via our Homebrew tap if you're on a Mac:

brew tap spotify/public
brew install ratatool
ratatool

Or download the release jar and run it.

wget https://github.com/spotify/ratatool/releases/download/v0.3.10/ratatool-cli-0.3.10.tar.gz
bin/ratatool directSampler

The command line tool can be used to sample from local file system or Google Cloud Storage directly if Google Cloud SDK is installed and authenticated.

bin/ratatool bigSampler avro --head -n 1000 --in gs://path/to/dataset --out out.avro
bin/ratatool bigSampler parquet --head -n 1000 --in gs://path/to/dataset --out out.parquet

# write output to both JSON file and BigQuery table
bin/ratatool bigSampler bigquery --head -n 1000 --in project_id:dataset_id.table_id \
    --out out.json--tableOut project_id:dataset_id.table_id

It can also be used to sample from HDFS with if core-site.xml and hdfs-site.xml are available.

bin/ratatool bigSampler avro \
    --head -n 10 --in hdfs://namenode/path/to/dataset --out file:///path/to/out.avro

Or execute BigDiffy directly

bin/ratatool bigDiffy \
    --input-mode=avro \
    --key=record.key \
    --lhs=gs://path/to/left \
    --rhs=gs://path/to/right \
    --output=gs://path/to/output \
    --runner=DataflowRunner ....

Development

Testing local changes to the CLI before releasing

To test local changes before release:

$ sbt
> project ratatoolCli
> packArchive

and then find the built CLI at ratatool-cli/target/ratatool-cli-{version}.tar.gz

License

Copyright 2016-2018 Spotify AB.

Licensed under the Apache License, Version 2.0: http://www.apache.org/licenses/LICENSE-2.0

主要指标

概览
名称与所有者spotify/ratatool
主编程语言Scala
编程语言Scala (语言数: 2)
平台
许可证Apache License 2.0
所有者活动
创建于2016-08-01 17:33:25
推送于2025-04-04 15:12:51
最后一次提交2025-04-04 17:12:51
发布数61
最新版本名称v0.4.10 (发布于 2024-06-07 11:02:57)
第一版名称v0.1.0 (发布于 2016-08-01 13:52:52)
用户参与
星数342
关注者数27
派生数54
提交数754
已启用问题?
问题数95
打开的问题数18
拉请求数330
打开的拉请求数11
关闭的拉请求数368
项目设置
已启用Wiki?
已存档?
是复刻?
已锁定?
是镜像?
是私有?