Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions docs/reference/sql/create.md
Original file line number Diff line number Diff line change
Expand Up @@ -158,13 +158,33 @@ Users can add table options by using `WITH`. The valid options contain the follo
| `append_mode` | Whether the table is append-only | String value. Default is 'false', which removes duplicate rows by primary keys and timestamps according to the `merge_mode`. Setting it to 'true' to enable append mode and create an append-only table which keeps duplicate rows. |
| `merge_mode` | The strategy to merge duplicate rows | String value. Only available when `append_mode` is 'false'. Default is `last_row`, which keeps the last row for the same primary key and timestamp. Setting it to `last_non_null` to keep the last non-null field for the same primary key and timestamp. |
| `sst_format` | The format of SST files | String value, supports `primary_key`, `flat`. Default is `flat`. `flat` is recommended for tables which have a large number of unique primary keys. |
| `experimental_sst_float_field_encoding` | Experimental SST encoding for floating-point field columns | String value: `default` (the default) or `byte_stream_split`. `byte_stream_split` applies Parquet byte-stream-split encoding to every `FLOAT` and `DOUBLE` field column of the table and disables dictionary encoding for those columns. Tag columns, the time index column, and columns of other types are not affected. |
| `comment` | Table level comment | String value. |
| `skip_wal` | Whether to disable Write-Ahead-Log for this table | String type. When set to `'true'`, the data written to the table will not be persisted to the write-ahead log, which can avoid storage wear and improve write throughput. However, when the process restarts, any unflushed data will be lost. Please use this feature only when the data source itself can ensure reliability. |
| `write_buffer_size` | Per-region write buffer stall threshold for this table | String type, such as `'512MB'` or `'1GB'`. For a positive value, GreptimeDB schedules a flush when mutable memtable usage reaches half the value, stalls writes at the value, and rejects writes at twice the value. The table option overrides `region_engine.mito.default_region_write_buffer_size`. An explicit `'0'` disables the per-region limit even when the engine default is nonzero. Unset the option to remove the table override and fall back to the engine default. |
| `auto_flush_interval` | How long a region of this table may go without a flush before one is triggered | String type, a time duration such as `'5m'` or `'1h'`. Must be greater than zero. The table option overrides the engine-wide `region_engine.mito.auto_flush_interval`. Set it to `NULL` with `ALTER TABLE` to drop the override and fall back to the engine setting. |
| `max_row_group_row_count` | Maximum number of rows in a Parquet row group | String type representing an integer from `1` through `10485760` (`10 * 1024 * 1024`). The default is `102400` (`100 * 1024`) when this option is not set. |
| `index.type` | Index type | **Only for metric engine** String value, supports `none`, `skipping`. |

#### Create a table with byte-stream-split float encoding

`byte_stream_split` stores the bytes of each floating-point value grouped by byte position before the Parquet writer compresses them, which can improve the compression ratio of `FLOAT` and `DOUBLE` field columns. Try it when floating-point fields make up a large share of the SST data, such as metrics and telemetry workloads: smaller SST files reduce storage usage, and queries that read those columns may need less I/O.

The option applies to every floating-point field column of the table, not only to a single value column. In a physical table for the metric engine, for example, it covers `greptime_value` and all other `FLOAT`/`DOUBLE` field columns. Tag columns (`PRIMARY KEY`), the `TIME INDEX` column, and columns of other types keep their existing encoding. This is a Mito table option, so any table that stores float fields can use it; it is not specific to the metric engine.

```sql
CREATE TABLE float_metrics (
ts TIMESTAMP TIME INDEX,
host STRING PRIMARY KEY,
val DOUBLE
) ENGINE=mito
WITH ('experimental_sst_float_field_encoding' = 'byte_stream_split');
```

Omitting the option or setting it to `default` keeps the current Parquet writer behavior, and the option is experimental. It changes how values are stored, not the values themselves or query results, and GreptimeDB reads SST files written with either encoding. Byte-stream-split does not use dictionary encoding, so the affected float field columns are written without dictionary encoding.

Whether the encoding reduces storage and I/O depends on the data distribution and the compression codec, and latency can change in either direction. Data that already compresses well, for example where dictionary encoding is effective, may not benefit. Benchmark your own workload before enabling it in production. See [byte-stream-split encoding for Prometheus remote write](/user-guide/ingest-data/for-observability/prometheus.md#byte-stream-split-encoding-for-float-fields) for a metric engine example. Set the option when you create the table; `ALTER TABLE` does not support changing it.

#### Create a table with a custom row group size

```sql
Expand Down
23 changes: 23 additions & 0 deletions docs/user-guide/ingest-data/for-observability/prometheus.md
Original file line number Diff line number Diff line change
Expand Up @@ -351,6 +351,29 @@ Setting `http.timeout = "0s"` (the default) disables the HTTP timeout entirely.

By default, the metric engine will automatically create a physical table named `greptime_physical_table` if it does not already exist. For performance optimization, you may choose to create a physical table with customized configurations.

### Byte-stream-split encoding for float fields

Comment thread
discord9 marked this conversation as resolved.
Create this table before sending data to it. If the default physical table already exists, choose a new physical table name and route remote writes to it; this example does not change an existing table.

If floating-point fields dominate the data you store, you can create a custom physical table with the [`experimental_sst_float_field_encoding`](/reference/sql/create.md#table-options) table option set to `byte_stream_split`:

```sql
CREATE TABLE greptime_physical_table (
Comment thread
discord9 marked this conversation as resolved.
greptime_timestamp TIMESTAMP(3) NOT NULL,
greptime_value DOUBLE NULL,
TIME INDEX (greptime_timestamp)
)
ENGINE = metric
WITH (
"physical_metric_table" = "",
"experimental_sst_float_field_encoding" = "byte_stream_split"
);
```

The encoding groups the bytes of each floating-point value by byte position before compression, which can reduce the size of the SST files that store metrics and the I/O needed to read float columns. It applies to every `FLOAT`/`DOUBLE` field column of the physical table, not only `greptime_value`, and never to tag or timestamp columns. It is a Mito table option, so it is not limited to the metric engine.

The option is opt-in: `default` keeps the existing writer behavior, stored values and query results are unchanged, and SST files written with either encoding can be read. Whether it reduces storage and I/O depends on your data distribution and compression setup, and latency can change in either direction; data that already compresses well may not benefit, so benchmark your workload before enabling it in production. This experimental option is set only when the physical table is created: it cannot be switched back to `default` on that table because `ALTER TABLE` does not support it. To use the default encoding for newly created metrics, create another physical table without this option and point their Remote Write requests to it. Existing logical metric tables remain associated with their original physical table: changing the URL does not move their data or switch their encoding. SSTs written with either encoding remain readable. The `x-greptime-hints` remote-write header does not set this physical table option. If you use a different physical table name, pass it in the `physical_table` parameter of the remote write URL; the example above uses the default name, so no URL change is needed.

### Enable skipping index

By default, the metric engine won't create indexes for columns. You can enable it by setting the `index.type` to `skipping`.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -160,13 +160,33 @@ GreptimeDB 提供了丰富的索引实现来加速查询,请在[索引](/user-
| `append_mode` | 该表是否时 append-only 的 | 字符串值。默认值为 'false',根据 'merge_mode' 按主键和时间戳删除重复行。设置为 'true' 可以开启 append 模式和创建 append-only 表,保留所有重复的行 |
| `merge_mode` | 合并重复行的策略 | 字符串值。只有当 `append_mode` 为 'false' 时可用。默认值为 `last_row`,保留相同主键和时间戳的最后一行。设置为 `last_non_null` 则保留相同主键和时间戳的最后一个非空字段。 |
| `sst_format` | SST 文件的格式 | 字符串值,支持 `primary_key`,`flat`。默认为 `flat`。`flat` 格式建议用于具有高基数主键的表。 |
| `experimental_sst_float_field_encoding` | 实验性浮点字段 SST 编码 | 字符串值,支持 `default` 和 `byte_stream_split`,默认值为 `default`。`byte_stream_split` 对表中所有 `FLOAT` 和 `DOUBLE` 字段列应用 Parquet 的 byte-stream-split 编码,并关闭这些列的字典编码;标签列、时间索引列以及其他类型的列不受影响。 |
| `comment` | 表级注释 | 字符串值。 |
| `index.type` | Index 类型 | **仅用于 metric engine** 字符串值,支持 `none`, `skipping`. |
| `skip_wal` | 是否关闭表的预写日志 | 字符串类型。当设置为 `'true'` 时表的写入数据将不会持久化到预写日志,可以避免存储磨损同时提升写入吞吐。但是当进程重启时,尚未 flush 的数据会丢失。请仅在数据源本身可以确保可靠性的情况下使用此功能。 |
| `write_buffer_size` | 该表的单 region 写缓冲区阻塞阈值 | 字符串类型,例如 `'512MB'` 或 `'1GB'`。设置为正值后,mutable memtable 内存用量达到该值的一半时,GreptimeDB 会调度 flush;达到该值时会阻塞写入,达到该值的 2 倍时会拒绝写入。该表选项会覆盖 `region_engine.mito.default_region_write_buffer_size`。即使引擎默认值非零,显式设置为 `'0'` 也会禁用单 region 限制。取消设置会移除表级覆盖,并回退到引擎默认值。 |
| `auto_flush_interval` | 该表的 region 最长多久没有 flush 就触发一次 flush | 字符串类型,是一个时间范围字符串,例如 `'5m'` 或 `'1h'`,必须大于 0。该表选项会覆盖引擎级的 `region_engine.mito.auto_flush_interval`。用 `ALTER TABLE` 将其设为 `NULL` 可以移除表级覆盖、回退到引擎级配置。 |
| `max_row_group_row_count` | Parquet row group 的最大行数 | 字符串类型,表示 `1` 到 `10485760`(`10 * 1024 * 1024`)之间的整数。未设置该选项时,默认值为 `102400`(`100 * 1024`)。 |

#### 创建使用 byte-stream-split 浮点编码的表

`byte_stream_split` 会在 Parquet 写入器压缩之前,按字节位置重新排列每个浮点值的字节,从而提升 `FLOAT` 和 `DOUBLE` 字段列的压缩率。当浮点字段在 SST 数据中占比较大时可以尝试启用,例如 metrics 和 telemetry 场景:更小的 SST 文件可以降低存储占用,读取这些列的查询也可能减少所需的 I/O。

该选项作用于表中的所有浮点字段列,而不只是某一个取值列。例如在 metric engine 的物理表中,它同时覆盖 `greptime_value` 以及其他所有 `FLOAT`/`DOUBLE` 字段列;标签列(`PRIMARY KEY`)、`TIME INDEX` 列以及其他类型的列保持原有编码。这是 Mito 的表选项,任何存储浮点字段的表都可以使用,并非 metric engine 专用。

```sql
CREATE TABLE float_metrics (
ts TIMESTAMP TIME INDEX,
host STRING PRIMARY KEY,
val DOUBLE
) ENGINE=mito
WITH ('experimental_sst_float_field_encoding' = 'byte_stream_split');
```

省略该选项或设置为 `default` 时保持当前 Parquet 写入行为,该选项属于实验性功能。它只改变数据的存储编码,不改变存储的值和查询结果,两种编码写出的 SST 文件都可以读取。byte-stream-split 不使用字典编码,因此受影响的浮点字段列不会采用字典编码。

该编码能否减少存储占用和 I/O 取决于数据分布和压缩算法,读写延迟的表现也因工作负载而异;原本压缩效果已经很好的数据(例如字典编码更有效的场景)可能并不会受益。请先在自己的工作负载上进行基准测试,再决定是否在生产环境启用。metric engine 物理表的配置示例请参考 [Prometheus Remote Write 的 byte-stream-split 编码](/user-guide/ingest-data/for-observability/prometheus.md#byte-stream-split-encoding-for-float-fields)。该选项需要在建表时设置,`ALTER TABLE` 不支持修改。

#### 创建自定义 row group 大小的表

```sql
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -345,6 +345,31 @@ GreptimeDB 会将其自动调整为该值并输出警告日志。
默认情况下,metric engine 会自动创建一个名为 `greptime_physical_table` 的物理表。
为了优化性能,你可以选择创建一个具有自定义配置的物理表。

<AnchorAlias id="byte-stream-split-encoding-for-float-fields" />

### 浮点字段使用 byte-stream-split 编码

请在向该表写入数据之前创建物理表。如果默认物理表已经存在,请使用新的物理表名称,并将 Remote Write 指向该表;以下示例不会修改已有表。

如果存储的数据以浮点字段为主,可以在自定义物理表时通过 [`experimental_sst_float_field_encoding`](/reference/sql/create.md#table-options) 表选项启用 `byte_stream_split` 编码:

```sql
CREATE TABLE greptime_physical_table (
greptime_timestamp TIMESTAMP(3) NOT NULL,
greptime_value DOUBLE NULL,
TIME INDEX (greptime_timestamp)
)
ENGINE = metric
WITH (
"physical_metric_table" = "",
"experimental_sst_float_field_encoding" = "byte_stream_split"
);
```

如果物理表使用其他名称,需要在 Remote Write URL 的 `physical_table` 参数中指定该名称。该编码会在压缩前按字节位置重新排列每个浮点值的字节,可以减小存储指标数据的 SST 文件大小,并减少读取浮点列所需的 I/O。它作用于物理表中的所有 `FLOAT`/`DOUBLE` 字段列,而不只是 `greptime_value`,也不会作用于标签列和时间戳列。这是 Mito 的表选项,并不限于 metric engine。

该选项默认关闭:`default` 保持当前写入行为,存储的值和查询结果都不变,两种编码写出的 SST 文件都可以读取。能否减少存储占用和 I/O 取决于数据分布和压缩配置,读写延迟的表现也因工作负载而异;原本压缩效果已经很好的数据可能不会受益,请先在自己的工作负载上进行基准测试,再在生产环境启用。这是实验性选项,只能在创建物理表时设置;`ALTER TABLE` 不支持将已有表改回 `default`。如果希望新创建的指标表使用默认编码,可以新建未设置该选项的物理表,并通过 Remote Write URL 将新指标写入该表。已有的逻辑指标表仍关联原物理表;修改 URL 不会迁移其已有数据,也不会切换其写入编码。两种编码写出的已有 SST 文件仍可读取。Remote Write 的 `x-greptime-hints` 请求头不能设置这个物理表选项。示例中使用的是默认物理表名称,因此无需修改 Remote Write URL。

### 启用跳数索引

默认情况下,metric engine 不会为列创建索引。你可以通过设置 `index.type` 为 `skipping` 来设置索引类型。
Expand Down
Loading