Skip to content

Repository files navigation

schemate

A small, checkmate-first schema DSL for R data.

CRAN_Status_Badge CRAN RStudio mirror downloads

schemate provides a small, checkmate-first schema DSL for R data. It can infer schemas from example objects, edit schema documents, save them as JSON, read them back, and validate new inputs against the schema.

During validation, each check node is translated to the corresponding checkmate::check_*() call. For example, "check": { "kind": "int", "lower": 1 } validates a value with checkmate::check_int(x, lower = 1) at the matched input path.

The package is meant for package authors and pipeline authors who want a compact R-native schema format without adopting the full JSON Schema vocabulary. It is a good fit for structural contracts around R objects, nested lists, JSON-like payloads, or package-facing inputs. If you need standards-compliant JSON document validation or tabular data-quality reporting, use a tool built for that domain.

A typical workflow is:

  1. infer a conservative schema with schema_infer();
  2. edit it with schema_*() authoring verbs;
  3. save it with schema_write();
  4. read it back with schema_read();
  5. validate inputs with schema_validate().

Installation

install.packages("schemate")

Development Version

To get a bug fix or to use a feature from the development version, you can install the development version of schemate from GitHub.

# install.packages("pak")
pak::pak("hongyuanjia/schemate")

Quick Start

The public API uses a single schema_ prefix and works well in pipelines. Start from an example object and infer a conservative schema. Schema objects print as the compact JSON-style DSL that can be reviewed or stored.

library(schemate)

person <- list(id = 1L, name = "Ada")
person_schema <- schema_infer(person, keys = "named")
person_schema
## {
##   "check": {
##     "kind": "list"
##   },
##   "keys": {
##     "type": "named"
##   },
##   "fields": {
##     "id": {
##       "check": {
##         "kind": "int"
##       }
##     },
##     "name": {
##       "check": {
##         "kind": "string"
##       }
##     }
##   }
## }

For larger payloads, compact the inferred schema into something easier to edit and review.

payload <- list(
    items = list(
        list(id = 1L, name = "alpha", label = "Alpha", slug = "alpha"),
        list(id = 2L, name = "beta", label = "Beta", slug = "beta")
    )
)

schema <- payload |>
    schema_infer(keys = "named", arrays = "rest") |>
    schema_compact() |>
    schema_set_desc("$items", "Repository-like result items")

schema |>
    schema_validate(payload, mode = "test")
## [1] TRUE

schema_validate() defaults to assert mode: invalid input raises an error and valid input is returned invisibly. Other modes are available when you need a message or a boolean result.

bad_payload <- payload
bad_payload$items[[1L]]$id <- "bad"

schema |>
    schema_validate(bad_payload, mode = "check", name = "payload")
## [1] "payload$items[[1]]$id: Must be of type 'single integerish value', not 'character'"
schema |>
    schema_validate(bad_payload, mode = "test", name = "payload")
## [1] FALSE

When validating many payloads against the same schema, flatten once and reuse the flattened schema.

flat <- schema_flatten(schema)
schema_validate(flat, payload, mode = "test")
## [1] TRUE

For a data frame example, see the Get started article.

JSON Workflow

Schemas are stored as a compact JSON DSL. The DSL is not JSON Schema; it is a thin representation of checkmate checks, field schemas, local definitions, and combinators. See the Schema DSL article for the complete format reference. schema_read() and schema_write() require the suggested package jsonlite.

path <- tempfile(fileext = ".json")
schema_write(schema, path)

restored <- schema_read(path)

restored |>
    schema_validate(payload, mode = "test")
## [1] TRUE

You can also write the compact JSON DSL directly when a schema is easier to review as data:

json_schema <- schema_read('{
  "check": { "kind": "list" },
  "keys": { "type": "named" },
  "fields": {
    "id": { "check": { "kind": "int", "lower": 1 } }
  },
  "rest": { "check": { "kind": "string" } }
}')

good <- list(id = 1L, label = "alpha")
bad <- list(id = 1L, label = 2L)

schema_validate(json_schema, good, mode = "test", name = "payload")
## [1] TRUE
schema_validate(json_schema, bad, mode = "check", name = "payload")
## [1] "payload$label: Must be of type 'string', not 'integer'"

Example schema files are installed under inst/extdata:

system.file("extdata", "person-schema.json", package = "schemate")

Validation Modes

schema_validate() supports four modes:

Mode Return value on success Return value on failure
assert invisibly returns the input throws an error
check TRUE diagnostic string
test TRUE FALSE
expect testthat-style expectation object expectation failure object

Use assert inside application code, check when displaying diagnostics, test for control flow, and expect in tests.

Standalone Use

schemate also publishes a generated standalone bundle for packages that want the schema features without depending on schemate at runtime.

usethis::use_standalone("hongyuanjia/schemate", "schema", ref = "standalone")

The imported file contains the core schema lifecycle: inference, edit helpers, JSON IO, flattening, and validation. The target package still needs to declare the standalone imports: checkmate and S7 for core schema work, and jsonlite in Suggests or Imports if it calls schema_read() or schema_write(). See the Standalone Use article for details.

Relation to Other Tools

schemate is closest in spirit to checkmate: schemas ultimately validate R objects by calling checkmate checks. It adds a schema lifecycle around those checks: infer, edit, serialize, read, and validate.

pointblank is a better fit for tabular data quality workflows, reporting, and column-oriented validation plans. schemate is deliberately narrower and more structural: it describes R values, R object names, nested lists, JSON-like payloads, and package-facing input contracts. It is not a replacement for JSON Schema or jsonvalidate, which are better choices when you need standards-compliant JSON document validation.

The R validation ecosystem is broad:

  • validate captures data validation rules that can be documented, stored, and applied to data sets.
  • assertr is designed for assertive data checks inside analysis pipelines.
  • data.validator focuses on dataset validation with reporting.
  • vetr provides template-based structural checks for R objects.
  • testthat is the right home for unit-test expectations; schema_validate(..., mode = "expect") is intended to fit into that style.

License

The project is released under the terms of MIT License.

About

A small, checmkate-first schema DSL for R data.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages