这是本节的多页打印视图。 .
HugeGraph-Server 配置
- 1: Server 启动指南
- 2: Server 完整配置手册
- 3: HugeGraph 内置用户权限与扩展权限配置及使用
- 4: 配置 HugeGraphServer 使用 https 协议
- 5: 配置 RocksDB 后端
- 6: 配置 HStore 分布式后端
- 7: 配置 HBase 后端
本节介绍 HugeGraph-Server 的配置文件、可用选项、认证和 HTTPS 设置。
后端配置
1 - Server 启动指南
1 概述
配置文件的目录为 hugegraph-release/conf,所有关于服务和图本身的配置都在此目录下。
主要的配置文件包括:gremlin-server.yaml、rest-server.properties 和 hugegraph.properties
HugeGraphServer 内部集成了 GremlinServer 和 RestServer,而 gremlin-server.yaml 和 rest-server.properties 就是用来配置这两个 Server 的。
- GremlinServer:GremlinServer 接收 Gremlin 请求并调用图引擎。
- RestServer:提供 RESTful API,根据不同的 HTTP 请求,调用对应的 Core API,如果用户请求体是 gremlin 语句,则会转发给 GremlinServer,实现对图数据的操作。
下面对这三个配置文件逐一介绍。
2 gremlin-server.yaml
gremlin-server.yaml 的主要结构如下。示例省略了部分导入项;完整内容以发布包中的文件为准。
通常只需关注 channelizer、host 和 port。图不在 Gremlin Server 的 graphs 段加载;是否读取本地图配置由 rest-server.properties 中的 graph.load_from_local_config 控制。
- channelizer:默认的
WsAndHttpChannelizer同时支持 WebSocket 和 HTTP。Gremlin-Console 使用 WebSocket,HugeGraph-Client、Loader 和 Hubble 使用 HTTP;
默认 GremlinServer 是服务在 127.0.0.1:8182,如果需要修改,配置 host、port 即可
- host:部署 GremlinServer 机器的机器名或 IP,GremlinServer 不直接暴露给用户,由 RestServer 转发 Gremlin 请求;
- port:部署 GremlinServer 机器的端口;
同时需要在 rest-server.properties 中增加对应的配置项 gremlinserver.url=http://host:port
3 rest-server.properties
下面是可用的 rest-server.properties 示例。当前上游发布模板没有写出 graph.load_from_local_config,而源码默认值为 false;使用 conf/graphs 中的本地图配置时必须显式设为 true。
- restserver.url:RestServer 提供服务的 url,根据实际环境修改。如果其他 IP 地址无法访问,可以尝试修改为特定的地址;或修改为
http://0.0.0.0来监听来自任何 IP 地址的请求,这种方案较为便捷,但需要留意服务可被访问的网络范围; - graphs:图配置文件所在目录,默认值是
./conf/graphs。init-store会扫描该目录;Server 仅在graph.load_from_local_config=true时加载其中的 properties 文件; - graph.load_from_local_config:是否在 Server 启动时读取本地图配置,源码默认值为
false;
当前上游模板中的 Arthas 键仍写作
arthas.telnet_port、arthas.http_port和arthas.disabled_commands,但ServerOptions读取的是下方示例中的 camelCase 名称。自定义配置应使用arthas.telnetPort、arthas.httpPort和arthas.disabledCommands。
配置项 gremlinserver.url 是 GremlinServer 为 RestServer 提供服务的 url,该配置项默认为 http://127.0.0.1:8182,如需修改,需要和 gremlin-server.yaml 中的 host 和 port 相匹配;该值可以像模板那样省略协议前缀,缺失时会自动补上
http://。
4 hugegraph.properties
hugegraph.properties 是一类文件,因为如果系统存在多个图,则会有多个相似的文件。该文件用来配置与图存储和查询相关的参数,文件的默认内容如下:
重点关注未注释的几项:
- gremlin.graph:GremlinServer 的启动入口,用户不要修改此项;开启鉴权时才改为
org.apache.hugegraph.auth.HugeFactoryAuthProxy; - vertex.cache_type / edge.cache_type:缓存实现,可选值为
l1和l2,默认l2; - backend:使用的后端存储。1.7.0 支持 memory、rocksdb、hstore 和 hbase;
- serializer:schema、vertex 和 edge 写入后端时使用的序列化器。RocksDB 使用 binary;
- store:图在后端使用的存储名称;
- task.schedule_period、task.retry、task.wait_timeout:异步任务的调度周期(秒)、重试次数和等待超时(秒)。调度器由后端决定,
hstore使用分布式调度器,其余后端使用本地调度器;旧的task.scheduler_type键已被忽略; - search.text_analyzer / search.text_analyzer_mode:全文索引使用的分词器及其模式。可选分词器为
ansj、hanlp、smartcn、jieba、jcseg、mmseg4j和ikanalyzer,每种分词器有各自的模式取值; - rocksdb.data_path:backend 为 rocksdb 时此项才有意义,rocksdb 的数据目录,默认为
rocksdb-data/data - rocksdb.wal_path:backend 为 rocksdb 时此项才有意义,rocksdb 的日志目录,默认为
rocksdb-data/wal
5 多图配置
一个 Server 可以加载多个图,每个图使用单独的 properties 文件。下面创建 RocksDB 图 hugegraph_rocksdb 和内存图 hugegraph_memory。
[可选]:修改 rest-server.properties
通过修改 rest-server.properties 中的 graphs 配置项来设置图的配置文件目录。默认配置为 graphs=./conf/graphs,如果想要修改为其它目录则调整 graphs 配置项,比如调整为 graphs=/etc/hugegraph/graphs,示例如下:
在 conf/graphs 路径下基于 hugegraph.properties 创建 hugegraph_memory.properties 和 hugegraph_rocksdb.properties。
hugegraph_memory.properties 修改如下:
hugegraph_rocksdb.properties 修改如下:
停止 Server,初始化执行 init-store.sh(为新的图创建数据库),重新启动 Server
查看创建的图:
查看某个图的信息:
2 - Server 完整配置手册
Gremlin Server 配置项
对应配置文件gremlin-server.yaml
| config option | default value | description |
|---|---|---|
| host | 127.0.0.1 | The host or ip of Gremlin Server. |
| port | 8182 | The listening port of Gremlin Server. |
| graphs | {} | 图由 Server 动态加载,不要在此处配置。 |
| evaluationTimeout | 30000 | Gremlin 脚本执行超时,单位为毫秒。 |
| channelizer | org.apache.tinkerpop.gremlin.server.channel.WsAndHttpChannelizer | 同时处理 WebSocket 和 HTTP 请求。 |
| maxContentLength | 65536 | Server 接受的单个请求的最大字节数。 |
| maxChunkSize | 8192 | HTTP 请求分块的最大字节数。 |
| maxHeaderSize | 8192 | HTTP 请求头的最大字节数。 |
| resultIterationBatchSize | 64 | 流式返回结果集时,每批返回的结果条数。 |
| ssl.enabled | false | Gremlin Server 是否启用 TLS。 |
| authentication | 未配置 | 启用认证时配置认证器、处理器和 rest-server.properties 路径。 |
Rest Server & API 配置项
对应配置文件rest-server.properties
| config option | default value | description |
|---|---|---|
| graphs | ./conf/graphs | 图配置 properties 文件所在目录。 |
| graph.load_from_local_config | false | 是否在 Server 启动时读取 graphs 目录;使用本地图配置时需设为 true。 |
| graphs.enable_dynamic_create_drop | true | Whether to enable create or drop graph dynamically. |
| init_store.enabled | true | Whether init-store initializes the local backend stores and the built-in admin account. Set false in distributed deployments (PD/HStore) where the storage side already owns the metadata. |
| server.id | 空字符串 | The optional legacy id of hugegraph-server. |
| server.role | master | The role of nodes in the cluster, available types are [master, worker, computer] |
| server.role_election | false | Whether to enable role election, if enabled, the server will elect a master node in the cluster. |
| server.node_id | node-id1 | The node id of the server. |
| server.node_role | worker | The node role of the server. |
| server.graphspace | DEFAULT | The graph space of the server. |
| server.service_id | DEFAULT | The service id of the server. |
| server.path_graphspace | DEFAULT | The default path graph space of the server. |
| server.start_ignore_single_graph_error | true | Whether to start ignore single graph error. |
| server.event_hub_threads | 1 | The event hub threads of server. |
| restserver.url | http://127.0.0.1:8080 | The url for listening of graph server. |
| ssl.keystore_file | conf/hugegraph-server.keystore | The path of server keystore file used when https protocol is enabled. |
| ssl.keystore_password | hugegraph | The password of the server keystore file when the https protocol is enabled. |
| white_ip.status | disable | The status of whether enable white ip. |
| restserver.max_worker_threads | 2 * CPUs | The maximum worker threads of rest server. |
| restserver.task_threads | max(4, CPUs / 2) | The task threads of rest server. |
| restserver.min_free_memory | 64 | The minimum free memory(MB) of rest server, requests will be rejected when the available memory of system is lower than this value. |
| restserver.request_timeout | 30 | The time in seconds within which a request must complete, -1 means no timeout. |
| restserver.connection_idle_timeout | 30 | The time in seconds to keep an inactive connection alive, -1 means no timeout. |
| restserver.connection_max_requests | 256 | The max number of HTTP requests allowed to be processed on one keep-alive connection, -1 means unlimited. |
| gremlinserver.url | http://127.0.0.1:8182 | The url of gremlin server. |
| gremlinserver.max_route | 2 * CPUs | The max route number for gremlin server. |
| gremlinserver.timeout | 30 | The timeout in seconds of waiting for gremlin server. |
| batch.max_edges_per_batch | 2500 | The maximum number of edges submitted per batch. |
| batch.max_vertices_per_batch | 2500 | The maximum number of vertices submitted per batch. |
| batch.max_write_ratio | 70 | The maximum thread ratio for batch writing, only take effect if the batch.max_write_threads is 0. |
| batch.max_write_threads | 0 | The maximum threads for batch writing, if the value is 0, the actual value will be set to batch.max_write_ratio * restserver.max_worker_threads. |
| raft.group_peers | 127.0.0.1:8090 | The rpc address of raft group initial peers. |
| auth.authenticator | The class path of authenticator implementation. e.g., org.apache.hugegraph.auth.StandardAuthenticator, or a custom implementation. | |
| auth.graph_store | hugegraph | The name of graph used to store authentication information, like users, only for org.apache.hugegraph.auth.StandardAuthenticator. |
| auth.admin_pa | pa | 内置 admin 账户的初始密码,仅首次启动时生效;部署前必须修改。 |
| auth.audit_log_rate | 1000.0 | The max rate of audit log output per user, default value is 1000 records per second. |
| auth.cache_capacity | 10240 | The max cache capacity of each auth cache item. |
| auth.cache_expire | 600 | The expiration time in seconds of auth cache in auth client and auth server. |
| auth.remote_url | If the address is empty, it provide auth service, otherwise it is auth client and also provide auth service through rpc forwarding. The remote url can be set to multiple addresses, which are concat by ‘,’. | |
| auth.token_expire | 86400 | The expiration time in seconds after token created |
| auth.token_secret | 启动时随机生成 | HS256 的密钥;需要跨重启保持既有 token 有效时应显式配置。 |
| exception.allow_trace | true | Whether to allow exception trace stack. |
| memory_monitor.threshold | 0.85 | Threshold for JVM memory usage monitoring, 1 means disabling the memory monitoring task. |
| memory_monitor.period | 2000 | The period in ms of JVM memory usage monitoring, in each period we will detect the jvm memory usage and take corresponding actions. |
| log.slow_query_threshold | 1000 | The threshold time(ms) of logging slow query, 0 means logging slow query is disabled. |
| log.slow_query_body_limit | 512 | 慢查询日志记录的请求体最大字节数,0 表示不记录。记录的前缀原样写入,可能包含敏感的 Gremlin 或 Cypher 字面量。 |
对应配置文件rest-server.properties,仅在 server.role_election=true 时生效。
| config option | default value | description |
|---|---|---|
| server.role.node_external_url | http://127.0.0.1:8080 | The url of external accessibility. |
| server.role.base_timeout | 500 | The role state machine candidate state base timeout time, in ms. |
| server.role.random_timeout | 1000 | The random timeout in ms that be used when candidate node request to become master state to reduce competitive voting. |
| server.role.heartbeat_interval | 2 | The role state machine heartbeat interval second time. |
| server.role.fail_count | 5 | When the node failed count of update or query heartbeat is reaches this threshold, the node will become abdication state to guardsafe property. |
| server.role.master_dead_times | 10 | When the worker node detects that the number of times the master node fails to update heartbeat reaches this threshold, the worker node will become to a candidate node. |
PD/Meta 配置项 (分布式模式)
对应配置文件rest-server.properties
| config option | default value | description |
|---|---|---|
| usePD | false | Whether use pd. |
| pd.peers | 127.0.0.1:8686 | The pd server peers, separated with commas. |
| cluster | hg-test | The cluster name. |
| metrics.data_to_pd | true | Whether to report metrics data to pd. |
| meta.endpoints | http://127.0.0.1:2379 | meta 端点的 URL。当前代码中没有任何地方读取该配置项,设置后不会生效;meta 连接由 pd.peers 建立。 |
| meta.use_ca | false | Whether to use ca to meta server. |
| meta.ca | The ca file of meta server. | |
| meta.client_ca | The client ca file of meta server. | |
| meta.client_key | The client key file of meta server. |
HStore 后端还会从图配置文件 {graph-name}.properties 中读取以下两项,默认值 0 表示由 PD 决定:
| config option | default value | description |
|---|---|---|
| hstore.partition_count | 0 | Number of partitions, which PD controls partitions based on. |
| hstore.shard_count | 0 | Number of copies, which PD controls partition copies based on. |
基本配置项
基本配置项及后端配置项对应配置文件:{graph-name}.properties,如hugegraph.properties
| config option | default value | description |
|---|---|---|
| gremlin.graph | org.apache.hugegraph.HugeFactory | Gremlin entrance to create graph. |
| backend | memory | The data store type. For version 1.7.0+ the allowed values are [memory, rocksdb, hstore, hbase]; the shipped conf/graphs/hugegraph.properties sets rocksdb and conf/graphs/hstore.properties.template sets hstore. Note: cassandra, scylladb, mysql, postgresql were removed in 1.7.0 (use <= 1.5.x for legacy backends). |
| serializer | text | The serializer for backend store, built-in values are [text, binary, binaryscatter]; a backend may register its own, like hbase. The shipped graph templates set binary. |
| serializer.buffer_max_capacity | 134217728 | The process-wide max capacity of one serialization buffer in bytes. |
| store | hugegraph | The backend database namespace. |
| store.connection_detect_interval | 600 | The interval in seconds for detecting connections, if the idle time of a connection exceeds this value, detect it and reconnect if needed before using, value 0 means detecting every time. |
| store.graph | g | The graph table name, which store vertex, edge and property. |
| graphspace | DEFAULT | The graph space name. |
| alias.graph.id | The graph alias id. | |
| graph.read_mode | OLTP_ONLY | The graph read mode, which could be ALL | OLTP_ONLY | OLAP_ONLY. |
| pd.peers | 127.0.0.1:8686 | The addresses of pd nodes, separated with commas. Only used by the hstore backend. |
| schema.illegal_name_regex | .\s+$|~. | The regex specified the illegal format for schema name. |
| schema.cache_capacity | 10000 | The max cache size(items) of schema cache. |
| schema.init_template | The template schema used to init graph. | |
| schema.index_rebuild_using_pushdown | true | Whether to use pushdown when to create/rebuild index. |
| vertex.cache_type | l2 | The type of vertex cache, allowed values are [l1, l2]. |
| vertex.cache_capacity | 10000000 | The max cache size(items) of vertex cache. |
| vertex.cache_expire | 600 | The expiration time in seconds of vertex cache. |
| vertex.check_customized_id_exist | false | Whether to check the vertices exist for those using customized id strategy. |
| vertex.default_label | vertex | The default vertex label. |
| vertex.tx_capacity | 10000 | The max size(items) of vertices(uncommitted) in transaction. |
| vertex.check_adjacent_vertex_exist | false | Whether to check the adjacent vertices of edges exist. |
| vertex.lazy_load_adjacent_vertex | true | Whether to lazy load adjacent vertices of edges. |
| vertex.part_edge_commit_size | 5000 | Whether to enable the mode to commit part of edges of vertex, enabled if commit size > 0, 0 means disabled. |
| vertex.encode_primary_key_number | true | Whether to encode number value of primary key in vertex id. |
| vertex.remove_left_index_at_overwrite | false | Whether remove left index at overwrite. |
| edge.cache_type | l2 | The type of edge cache, allowed values are [l1, l2]. |
| edge.cache_capacity | 1000000 | The max cache size(items) of edge cache. |
| edge.cache_expire | 600 | The expiration time in seconds of edge cache. |
| edge.tx_capacity | 10000 | The max size(items) of edges(uncommitted) in transaction. |
| query.page_size | 500 | The size of each page when querying by paging. |
| query.batch_size | 1000 | The size of each batch when querying by batch. |
| query.ignore_invalid_data | true | Whether to ignore invalid data of vertex or edge. |
| query.index_intersect_threshold | 1000 | The maximum number of intermediate results to intersect indexes when querying by multiple single index properties. |
| query.max_indexes_available | 1 | The upper limit of the number of indexes that can be used to query. |
| query.dedup_option | limit | The way to dedup data, allowed values are [limit, global]. |
| query.trust_index | false | Whether to trust index. |
| query.ramtable_edges_capacity | 20000000 | The maximum number of edges in ramtable, include OUT and IN edges. |
| query.ramtable_enable | false | Whether to enable ramtable for query of adjacent edges. |
| query.ramtable_vertices_capacity | 10000000 | The maximum number of vertices in ramtable, generally the largest vertex id is used as capacity. |
| query.optimize_aggregate_by_index | false | Whether to optimize aggregate query(like count) by index. |
| oltp.concurrent_depth | 10 | The min depth to enable concurrent oltp algorithm. |
| oltp.concurrent_threads | max(10, CPUs / 2) | Thread number to concurrently execute oltp algorithm. |
| oltp.collection_type | EC | The implementation type of collections used in oltp algorithm, allowed values are [JCF, EC, FU]. |
| oltp.query_batch_size | 10000 | The size of each batch when executing oltp algorithm. |
| oltp.query_batch_avg_degree_ratio | 0.95 | The ratio of exponential approximation for average degree of iterator when executing oltp algorithm. |
| oltp.query_batch_expect_degree | 100000000 | The expect sum of degree in each batch when executing oltp algorithm. |
| rate_limit.read | 0 | The max rate(times/s) to execute query of vertices/edges. |
| rate_limit.write | 0 | The max rate(items/s) to add/update/delete vertices/edges. |
| task.schedule_period | 10 | Period time in seconds when scheduler to schedule task. |
| task.wait_timeout | 10 | Timeout in seconds for waiting for the task to complete, such as when truncating or clearing the backend. |
| task.retry | 0 | Task retry times, allowed range is [0, 3]. |
| task.input_size_limit | 16777216 | The job input size limit in bytes. |
| task.result_size_limit | 16777216 | The job result size limit in bytes. |
| task.sync_deletion | false | Whether to delete schema or expired data synchronously. |
| task.ttl_delete_batch | 1 | The batch size used to delete expired data. |
| computer.config | ./conf/computer.yaml | The config file path of computer job. |
| k8s.operator_template | ./conf/operator-template.yaml | The path of operator container template. |
| k8s.quota_template | ./conf/resource-quota-template.yaml | The path of resource quota template. |
| search.text_analyzer | ikanalyzer | Choose a text analyzer for searching the vertex/edge properties, available type are [ansj, hanlp, smartcn, jieba, jcseg, mmseg4j, ikanalyzer]. The shipped graph templates set jieba. If use ‘ikanalyzer’, need download jar from ‘https://github.com/apache/hugegraph-doc/raw/ik_binary/dist/server/ikanalyzer-2012_u6.jar' to lib directory |
| search.text_analyzer_mode | smart | Specify the mode for the text analyzer, the available mode of analyzer are {ansj: [BaseAnalysis, IndexAnalysis, ToAnalysis, NlpAnalysis], hanlp: [standard, nlp, index, nShort, shortest, speed], smartcn: [], jieba: [SEARCH, INDEX], jcseg: [Simple, Complex], mmseg4j: [Simple, Complex, MaxWord], ikanalyzer: [smart, max_word]}. |
| snowflake.datacenter_id | 0 | The datacenter id of snowflake id generator. |
| snowflake.force_string | false | Whether to force the snowflake long id to be a string. |
| snowflake.worker_id | 0 | The worker id of snowflake id generator. |
| memory.mode | off-heap | The memory mode used for query in HugeGraph. |
| memory.max_capacity | 1073741824 | The maximum memory capacity in bytes that can be managed for all queries in HugeGraph. |
| memory.one_query_max_capacity | 104857600 | The maximum memory capacity in bytes that can be managed for a query in HugeGraph. |
| memory.alignment | 8 | The alignment used for round memory size. |
发行包中的图配置模板已将这些配置项标注为废弃。它们仅在 raft.mode=true 时生效,
且 raft.group_peers 从 rest-server.properties 读取,而不是图配置文件。
| config option | default value | description |
|---|---|---|
| raft.mode | false | Whether the backend storage works in raft mode. |
| raft.safe_read | false | Whether to use linearly consistent read. |
| raft.path | ./raftlog | The log path of current raft node. |
| raft.use_replicator_pipeline | true | Whether to use replicator line, when turned on it multiple logs can be sent in parallel, and the next log doesn’t have to wait for the ack message of the current log to be sent. |
| raft.election_timeout | 10000 | Timeout in milliseconds to launch a round of election. |
| raft.snapshot_interval | 3600 | The interval in seconds to trigger snapshot save. |
| raft.snapshot_threads | 4 | The thread number used to do snapshot. |
| raft.snapshot_parallel_compress | false | Whether to enable parallel compress. |
| raft.snapshot_compress_threads | 4 | The thread number used to do snapshot compress. |
| raft.snapshot_decompress_threads | 4 | The thread number used to do snapshot decompress. |
| raft.backend_threads | CPUs | The thread number used to apply task to backend. |
| raft.read_index_threads | 8 | The thread number used to execute reading index. |
| raft.read_strategy | ReadOnlyLeaseBased | The linearizability of read strategy, allowed values are [ReadOnlyLeaseBased, ReadOnlySafe]. |
| raft.apply_batch | 1 | The apply batch size to trigger disruptor event handler. |
| raft.queue_size | 16384 | The disruptor buffers size for jraft RaftNode, StateMachine and LogManager. |
| raft.queue_publish_timeout | 60 | The timeout in second when publish event into disruptor. |
| raft.rpc_threads | max(CPUs * 2, 80) | The rpc threads for jraft RPC layer. |
| raft.rpc_connect_timeout | 5000 | The rpc connect timeout in milliseconds for jraft rpc. |
| raft.rpc_timeout | 60 | The general rpc timeout in seconds for jraft rpc. |
| raft.install_snapshot_rpc_timeout | 36000 | The install snapshot rpc timeout in seconds for jraft rpc. |
| raft.rpc_buf_low_water_mark | 10485760 | The ChannelOutboundBuffer’s low water mark of netty, when buffer size less than this size, the method ChannelOutboundBuffer.isWritable() will return true, it means that low downstream pressure or good network. |
| raft.rpc_buf_high_water_mark | 20971520 | The ChannelOutboundBuffer’s high water mark of netty, only when buffer size exceed this size, the method ChannelOutboundBuffer.isWritable() will return false, it means that the downstream pressure is too great to process the request or network is very congestion, upstream needs to limit rate at this time. |
RocksDB 后端配置项
| config option | default value | description |
|---|---|---|
| backend | Must be set to rocksdb. | |
| serializer | Must be set to binary. | |
| rocksdb.data_path | rocksdb-data/data | The path for storing data of RocksDB. |
| rocksdb.wal_path | rocksdb-data/wal | The path for storing WAL of RocksDB. |
| rocksdb.sst_path | The path for ingesting SST file into RocksDB. | |
| rocksdb.data_disks | [] | The optimized disks for storing data of RocksDB. The format of each element: STORE/TABLE: /path/disk.Allowed keys are [g/vertex, g/edge_out, g/edge_in, g/vertex_label_index, g/edge_label_index, g/range_int_index, g/range_float_index, g/range_long_index, g/range_double_index, g/secondary_index, g/search_index, g/shard_index, g/unique_index, g/olap] |
| rocksdb.log_level | INFO | The info log level of RocksDB. |
| rocksdb.num_levels | 7 | Set the number of levels for this database. |
| rocksdb.compaction_style | LEVEL | Set compaction style for RocksDB: LEVEL/UNIVERSAL/FIFO. |
| rocksdb.optimize_mode | true | Optimize for heavy workloads and big datasets. |
| rocksdb.bulkload_mode | false | Switch to the mode to bulk load data into RocksDB. |
| rocksdb.compression_per_level | [none, none, snappy, snappy, snappy, snappy, snappy] | The compression algorithms for different levels of RocksDB, allowed values are none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd. |
| rocksdb.bottommost_compression | none | The compression algorithm for the bottommost level of RocksDB, allowed values are none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd. |
| rocksdb.compression | snappy | The compression algorithm for compressing blocks of RocksDB, allowed values are none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd. |
| rocksdb.max_background_jobs | 8 | Maximum number of concurrent background jobs, including flushes and compactions. |
| rocksdb.max_subcompactions | 4 | The value represents the maximum number of threads per compaction job. |
| rocksdb.delayed_write_rate | 16777216 | The rate limit in bytes/s of user write requests when need to slow down if the compaction gets behind. |
| rocksdb.max_open_files | -1 | The maximum number of open files that can be cached by RocksDB, -1 means no limit. |
| rocksdb.max_manifest_file_size | 104857600 | The max size of manifest file in bytes. |
| rocksdb.skip_stats_update_on_db_open | false | Whether to skip statistics update when opening the database, setting this flag true allows us to not update statistics. |
| rocksdb.skip_check_sst_size_on_db_open | false | Whether to skip checking sizes of all sst files when opening the database. |
| rocksdb.max_file_opening_threads | 16 | The max number of threads used to open files. |
| rocksdb.max_total_wal_size | 0 | Total size of WAL files in bytes. Once WALs exceed this size, we will start forcing the flush of column families related, 0 means no limit. |
| rocksdb.bytes_per_sync | 0 | Allows OS to incrementally sync SST files to disk while they are being written, asynchronously in the background. Issue one request for every bytes_per_sync written. 0 turns it off. |
| rocksdb.wal_bytes_per_sync | 0 | Allows OS to incrementally sync WAL files to disk while they are being written, asynchronously in the background. Issue one request for every bytes_per_sync written. 0 turns it off. |
| rocksdb.strict_bytes_per_sync | false | When true, guarantees SST/WAL files have at most bytes_per_sync/wal_bytes_per_sync bytes submitted for writeback at any given time. This can be used to handle cases where processing speed exceeds I/O speed. |
| rocksdb.db_write_buffer_size | 0 | Total size of write buffers in bytes across all column families, 0 means no limit. |
| rocksdb.log_readahead_size | 0 | The number of bytes to prefetch when reading the log. 0 means the prefetching is disabled. |
| rocksdb.compaction_readahead_size | 0 | The number of bytes to perform bigger reads when doing compaction. If running RocksDB on spinning disks, you should set this to at least 2MB. 0 means the prefetching is disabled. |
| rocksdb.row_cache_capacity | 0 | The capacity in bytes of global cache for table-level rows. 0 means the row_cache is disabled. |
| rocksdb.delete_obsolete_files_period | 21600 | The periodicity in seconds when obsolete files get deleted, 0 means always do full purge. |
| rocksdb.write_buffer_size | 134217728 | Amount of data in bytes to build up in memory. |
| rocksdb.max_write_buffer_number | 6 | The maximum number of write buffers that are built up in memory. |
| rocksdb.min_write_buffer_number_to_merge | 2 | The minimum number of write buffers that will be merged together. |
| rocksdb.max_write_buffer_number_to_maintain | 0 | The total maximum number of write buffers to maintain in memory for conflict checking when transactions are used. |
| rocksdb.memtable_bloom_size_ratio | 0.0 | If prefix-extractor is set and memtable_bloom_size_ratio is not 0, or if memtable_whole_key_filtering is set true, create bloom filter for memtable with the size of write_buffer_size * memtable_bloom_size_ratio. If it is larger than 0.25, it is santinized to 0.25. |
| rocksdb.memtable_whole_key_filtering | false | Enable whole key bloom filter in memtable, it can potentially reduce CPU usage for point-look-ups. Note this will only take effect if memtable_bloom_size_ratio > 0. |
| rocksdb.memtable_huge_page_size | 0 | The page size for huge page TLB for bloom in memtable. If <= 0, not allocate from huge page TLB but from malloc. |
| rocksdb.inplace_update_support | false | Allows thread-safe inplace updates if a put key exists in current memtable and sizeof new value is smaller. |
| rocksdb.level_compaction_dynamic_level_bytes | false | Whether to enable level_compaction_dynamic_level_bytes, if it’s enabled we give max_bytes_for_level_multiplier a priority against max_bytes_for_level_base, the bytes of base level is dynamic for a more predictable LSM tree, it is useful to limit worse case space amplification. Turning this feature on/off for an existing DB can cause unexpected LSM tree structure so it’s not recommended. |
| rocksdb.max_bytes_for_level_base | 536870912 | The upper-bound of the total size of level-1 files in bytes. |
| rocksdb.max_bytes_for_level_multiplier | 10.0 | The ratio between the total size of level (L+1) files and the total size of level L files for all L. |
| rocksdb.target_file_size_base | 67108864 | The target file size for compaction in bytes. |
| rocksdb.target_file_size_multiplier | 1 | The size ratio between a level L file and a level (L+1) file. |
| rocksdb.level0_file_num_compaction_trigger | 2 | Number of files to trigger level-0 compaction. |
| rocksdb.level0_slowdown_writes_trigger | 20 | Soft limit on number of level-0 files for slowing down writes. |
| rocksdb.level0_stop_writes_trigger | 36 | Hard limit on number of level-0 files for stopping writes. |
| rocksdb.soft_pending_compaction_bytes_limit | 68719476736 | The soft limit to impose on pending compaction in bytes. |
| rocksdb.hard_pending_compaction_bytes_limit | 274877906944 | The hard limit to impose on pending compaction in bytes. |
| rocksdb.allow_mmap_writes | false | Allow the OS to mmap file for writing. |
| rocksdb.allow_mmap_reads | false | Allow the OS to mmap file for reading sst tables. |
| rocksdb.use_direct_reads | false | Enable the OS to use direct I/O for reading sst tables. |
| rocksdb.use_direct_io_for_flush_and_compaction | false | Enable the OS to use direct read/writes in flush and compaction. |
| rocksdb.use_fsync | false | If true, then every store to stable storage will issue a fsync. |
| rocksdb.atomic_flush | false | If true, flushing multiple column families and committing their results atomically to MANIFEST. Note that it’s not necessary to set atomic_flush=true if WAL is always enabled. |
| rocksdb.format_version | 5 | The format version of BlockBasedTable, allowed values are 0~5. |
| rocksdb.index_type | kBinarySearch | The index type used to lookup between data blocks with the sst table, allowed values are [kBinarySearch,kHashSearch,kTwoLevelIndexSearch,kBinarySearchWithFirstKey]. |
| rocksdb.data_block_index_type | kDataBlockBinarySearch | The search type used to point lookup in data block with the sst table, allowed values are [kDataBlockBinarySearch,kDataBlockBinaryAndHash]. |
| rocksdb.data_block_hash_table_util_ratio | 0.75 | The hash table utilization ratio value of entries/buckets. It is valid only when data_block_index_type=kDataBlockBinaryAndHash. |
| rocksdb.block_size | 4096 | Approximate size of user data packed per block, Note that it corresponds to uncompressed data. |
| rocksdb.block_size_deviation | 10 | The percentage of free space used to close a block. |
| rocksdb.block_restart_interval | 16 | The block restart interval for delta encoding in blocks. |
| rocksdb.block_cache_capacity | 8388608 | The amount of block cache in bytes that will be used by RocksDB, 0 means no block cache. |
| rocksdb.cache_index_and_filter_blocks | true | Set this option true if we’d put index/filter blocks to the block cache. |
| rocksdb.pin_l0_filter_and_index_blocks_in_cache | true | Set this option true if we’d pin L0 index/filter blocks to the block cache. |
| rocksdb.bloom_filter_bits_per_key | -1 | The bits per key in bloom filter, a good value is 10, which yields a filter with ~ 1% false positive rate. Set bloom_filter_bits_per_key > 0 to enable bloom filter, -1 means no bloom filter (0~0.5 round down to no filter). |
| rocksdb.bloom_filter_block_based_mode | false | If bloom filter is enabled, set this option true to use block based filter rather than full filter. |
| rocksdb.bloom_filter_whole_key_filtering | true | If bloom filter is enabled, set this option true to place whole keys in the bloom filter, else place the prefix of keys when prefix-extractor is set. |
| rocksdb.optimize_filters_for_hits | true | If bloom filter is enabled, this flag allows us to not store filters for the last level. set this option true to optimize the filters mainly for cases where keys are found rather than also optimize for keys missed. |
| rocksdb.partition_filters_and_indexes | false | If bloom filter is enabled, set this option true to use partitioned full filters and indexes for each sst file. This option is incompatible with block-based filters. |
| rocksdb.pin_top_level_index_and_filter | true | If partition_filters_and_indexes is set true, set this option true if we’d pin top-level index of partitioned filter and index blocks to the block cache. |
| rocksdb.prefix_extractor_n_bytes | 0 | The prefix-extractor uses the first N bytes of a key as its prefix, it will use the full key when a key is shorter than the N. 0 means unset prefix-extractor. |
对应配置文件rest-server.properties
| config option | default value | description |
|---|---|---|
| server.use_k8s | false | Whether to use k8s to support multiple tenancy. |
| server.deploy_in_k8s | false | Whether to deploy server in k8s. |
| server.urls_to_pd | http://0.0.0.0:8080 | Used as the server address reserved for PD and provided to clients, only used when starting the server in k8s. |
| server.k8s_url | https://127.0.0.1:8888 | The url of k8s. |
| server.k8s_use_ca | false | Whether to use ca to k8s api server. |
| server.k8s_ca | The ca file of k8s api server. | |
| server.k8s_client_ca | The client ca file of k8s api server. | |
| server.k8s_client_key | The client key file of k8s api server. | |
| k8s.api | false | The k8s api start status when the computer service is enabled. |
| k8s.namespace | hugegraph-computer-system | The namespace used for k8s work when the computer service is enabled. |
| k8s.kubeconfig | The k8s kube config file when the computer service is enabled. | |
| k8s.hugegraph_url | The hugegraph url for k8s work when the computer service is enabled. | |
| k8s.enable_internal_algorithm | true | Whether to open k8s internal algorithm. |
| service.access_pd_name | hg | Service name for server to access pd service. |
| service.access_pd_token | Service token for server to access pd service. | |
| server.k8s_oltp_image | 127.0.0.1/kgs_bd/hugegraphserver:3.0.0 | The oltp server image of k8s. |
| server.k8s_olap_image | hugegraph/hugegraph-server:v1 | The olap server image of k8s. |
| server.k8s_storage_image | hugegraph/hugegraph-server:v1 | The storage server image of k8s. |
| server.default_oltp_k8s_namespace | hugegraph-server | The default oltp namespace for HugeGraph default graph space. |
| server.default_olap_k8s_namespace | hugegraph-computer-system | The default olap namespace for HugeGraph default graph space. |
| k8s.internal_algorithm | [page-rank, degree-centrality, wcc, triangle-count, rings, rings-with-filter, betweenness-centrality, closeness-centrality, lpa, links, kcore, louvain, clustering-coefficient, ppr, subgraph-match] | The names of the built-in k8s algorithms. |
| k8s.algorithms | See ServerOptions.K8S_ALGORITHMS | The name:paramsClass mapping of the built-in k8s algorithms. |
对应配置文件rest-server.properties
| config option | default value | description |
|---|---|---|
| arthas.telnetPort | 8562 | Arthas telnet port. |
| arthas.httpPort | 8561 | Arthas HTTP port. |
| arthas.ip | 0.0.0.0 | Arthas bind IP. |
| arthas.disabledCommands | jad | Disabled Arthas commands, separated by commas. |
对应配置文件rest-server.properties
| config option | default value | description |
|---|---|---|
| rpc.server_host | The hosts/ips bound by rpc server to provide services, empty value means not enabled. | |
| rpc.server_port | 8090 | The port bound by rpc server to provide services. |
| rpc.server_adaptive_port | false | Whether the bound port is adaptive, if it’s enabled, when the port is in use, automatically +1 to detect the next available port. Note that this process is not atomic, so there may still be port conflicts. |
| rpc.server_timeout | 30 | The timeout(in seconds) of rpc server execution. |
| rpc.remote_url | The remote urls of rpc peers, it can be set to multiple addresses, which are concat by ‘,’, empty value means not enabled. | |
| rpc.client_connect_timeout | 20 | The timeout(in seconds) of rpc client connect to rpc server. |
| rpc.client_reconnect_period | 10 | The period(in seconds) of rpc client reconnect to rpc server. |
| rpc.client_read_timeout | 40 | The timeout(in seconds) of rpc client read from rpc server. |
| rpc.client_retries | 3 | Failed retry number of rpc client calls to rpc server. |
| rpc.client_load_balancer | consistentHash | The rpc client uses a load-balancing algorithm to access multiple rpc servers in one cluster. Default value is ‘consistentHash’, means forwarding by request parameters. |
| rpc.protocol | bolt | Rpc communication protocol, client and server need to be specified the same value. |
| rpc.serialization | hessian2 | Rpc serialization type, client and server must set the same value. Note: If you choose ‘protobuf’, you need to add the relative IDL file. (Could refer PD/Store *.proto) |
| rpc.config_order | 999 | Sofa-RPC configuration file loading order, the larger the more later loading. |
| rpc.logger_impl | com.alipay.sofa.rpc.log.SLF4JLoggerImpl | Sofa-RPC log implementation class. |
| config option | default value | description |
|---|---|---|
| backend | Must be set to hbase. | |
| serializer | Must be set to hbase. | |
| hbase.hosts | localhost | The hostnames or ip addresses of HBase zookeeper, separated with commas. |
| hbase.port | 2181 | The port address of HBase zookeeper. |
| hbase.threads_max | 64 | The max threads num of hbase connections. |
| hbase.znode_parent | /hbase | The znode parent path of HBase zookeeper. |
| hbase.zk_retry | 3 | The recovery retry times of HBase zookeeper. |
| hbase.truncate_timeout | 30 | The timeout in seconds of waiting for store truncate. |
| hbase.aggregation_timeout | 43200 | The timeout in seconds of waiting for aggregation. |
| hbase.kerberos_enable | false | Is Kerberos authentication enabled for HBase. |
| hbase.kerberos_keytab | The HBase’s key tab file for kerberos authentication. | |
| hbase.kerberos_principal | The HBase’s principal for kerberos authentication. | |
| hbase.krb5_conf | /etc/krb5.conf | Kerberos configuration file, including KDC IP, default realm, etc. |
| hbase.hbase_site | /etc/hbase/conf/hbase-site.xml | The HBase’s configuration file |
| hbase.enable_partition | true | Is pre-split partitions enabled for HBase. |
| hbase.vertex_partitions | 10 | The number of partitions of the HBase vertex table. |
| hbase.edge_partitions | 30 | The number of partitions of the HBase edge table. |
≤ 1.5 版本配置 (Legacy)
以下后端存储在 1.7.0+ 版本中不再支持,仅在 1.5.x 及更早版本中可用:
| config option | default value | description |
|---|---|---|
| backend | Must be set to cassandra. | |
| serializer | Must be set to cassandra. | |
| cassandra.host | localhost | The seeds hostname or ip address of cassandra cluster. |
| cassandra.port | 9042 | The seeds port address of cassandra cluster. |
| cassandra.connect_timeout | 5 | The cassandra driver connect server timeout(seconds). |
| cassandra.read_timeout | 20 | The cassandra driver read from server timeout(seconds). |
| cassandra.keyspace.strategy | SimpleStrategy | The replication strategy of keyspace, valid value is SimpleStrategy or NetworkTopologyStrategy. |
| cassandra.keyspace.replication | [3] | The keyspace replication factor of SimpleStrategy, like ‘[3]’.Or replicas in each datacenter of NetworkTopologyStrategy, like ‘[dc1:2,dc2:1]’. |
| cassandra.username | The username to use to login to cassandra cluster. | |
| cassandra.password | The password corresponding to cassandra.username. | |
| cassandra.compression_type | none | The compression algorithm of cassandra transport: none/snappy/lz4. |
| cassandra.jmx_port=7199 | 7199 | The port of JMX API service for cassandra. |
| cassandra.aggregation_timeout | 43200 | The timeout in seconds of waiting for aggregation. |
| config option | default value | description |
|---|---|---|
| backend | Must be set to scylladb. | |
| serializer | Must be set to scylladb. |
其它与 Cassandra 后端一致。
| config option | default value | description |
|---|---|---|
| backend | Must be set to mysql. | |
| serializer | Must be set to mysql. | |
| jdbc.driver | com.mysql.jdbc.Driver | The JDBC driver class to connect database. |
| jdbc.url | jdbc:mysql://127.0.0.1:3306 | The url of database in JDBC format. |
| jdbc.username | root | The username to login database. |
| jdbc.password | ****** | The password corresponding to jdbc.username. |
| jdbc.ssl_mode | false | The SSL mode of connections with database. |
| jdbc.reconnect_interval | 3 | The interval(seconds) between reconnections when the database connection fails. |
| jdbc.reconnect_max_times | 3 | The reconnect times when the database connection fails. |
| jdbc.storage_engine | InnoDB | The storage engine of backend store database, like InnoDB/MyISAM/RocksDB for MySQL. |
| jdbc.postgresql.connect_database | template1 | The database used to connect when init store, drop store or check store exist. |
| config option | default value | description |
|---|---|---|
| backend | Must be set to postgresql. | |
| serializer | Must be set to postgresql. |
其它与 MySQL 后端一致。
PostgreSQL 后端的 driver 和 url 应该设置为:
jdbc.driver=org.postgresql.Driverjdbc.url=jdbc:postgresql://localhost:5432/
3 - HugeGraph 内置用户权限与扩展权限配置及使用
概述
HugeGraph 内置 StandardAuthenticator,支持多用户认证和基于“用户、用户组、操作、资源”的权限控制。
StandardAuthenticator 模式的几个核心设计:
- 初始化时创建超级管理员 (
admin) 用户,后续通过超级管理员创建其它用户,新创建的用户被分配足够权限后,可以创建或管理更多的用户 - 支持动态创建用户、用户组、资源,支持动态分配或取消权限
- 用户可以属于一个或多个用户组,每个用户组可以拥有对任意个资源的操作权限,操作类型包括:读、写、删除、执行等种类
- “资源” 描述了图数据库中的数据,比如符合某一类条件的顶点,每一个资源包括
type、label、properties三个要素,共有 18 种类型、任意 label、任意 properties 可组合形成的资源,一个资源的内部条件是且关系,多个资源之间的条件是或关系
举例说明:
配置用户认证
HugeGraph 目前默认未启用用户认证功能,需通过修改配置文件来启用该功能。
⚠️ SEC 提醒:图查询语言 (Gremlin/Cypher) 的安全性
鉴于图查询语言的灵活性可能带来的潜在系统安全隐患,不要把 Gremlin、Cypher 等查询接口直接暴露到公网。生产环境应同时启用鉴权、IP 白名单和审计日志,并通过 Docker 或 Kubernetes 隔离 Server 进程。
StandardAuthenticator 支持多用户认证和细粒度权限控制。也可以实现 HugeAuthenticator 接口来接入已有的用户系统。
用户认证使用 HTTP Basic Authentication。Basic 后面的值是 用户名:密码 的 Base64 编码。使用 curl 时可直接通过 -u 传入凭据:
警告:在 1.5.0 之前版本的 HugeGraph-Server 在鉴权模式下存在 JWT 相关的安全隐患,请务必使用新版本或自行修改 JWT token 的 secretKey。
修改方式为在配置文件rest-server.properties中重写auth.token_secret信息:(1.5.0 后会默认生成随机值则无需配置)
也可以通过下面的命令实现:
由于默认值在每次启动时随机生成,当 token 需要在重启后继续有效、或者需要被多个服务节点接受时,必须显式配置该项。token 的有效期由
auth.token_expire 决定,默认为 86400 秒。
StandardAuthenticator 模式
StandardAuthenticator模式是通过在数据库后端存储用户信息来支持用户认证和权限控制,该实现基于数据库存储的用户的名称与密码进行认证(密码已被加密),基于用户的角色来细粒度控制用户权限。下面是具体的配置流程(重启服务生效):
在配置文件gremlin-server.yaml中配置authenticator及其rest-server文件路径:
在 rest-server.properties 中配置认证器和权限数据存储图:
其中,graph_store配置项是指使用哪一个图来存储用户信息,如果存在多个图的话,选取任意一个均可。
在配置文件hugegraph{n}.properties中配置gremlin.graph信息:
权限 API 的调用方式见 Authentication API 文档。
自定义用户认证系统
如果需要支持更加灵活的用户系统,可自定义 authenticator 进行扩展,自定义 authenticator 实现接口org.apache.hugegraph.auth.HugeAuthenticator即可,然后修改配置文件中authenticator配置项指向该实现。
基于鉴权模式启动
首次执行 init-store.sh 时,如果尚未创建 admin 用户,命令会要求输入管理员密码。对于已经初始化的持久化后端,init-store.sh 会补充认证所需的系统信息,无需删除原有图数据。
使用 Docker 时开启鉴权模式
对于镜像 hugegraph/hugegraph 大于等于 1.2.0 的版本,我们可以在启动 docker 镜像的同时开启鉴权模式
具体做法如下:
1. 采用 docker run
在 docker run 中添加环境变量 PASSWORD=xxx(密码可以自由设置)即可开启鉴权模式::
2. 采用 docker-compose
使用 docker-compose 在环境变量中设置 PASSWORD=xxx即可
3. 进入容器后重新开启鉴权模式
首先进入容器:
之后参照 基于鉴权模式启动 即可
4 - 配置 HugeGraphServer 使用 https 协议
概述
HugeGraphServer 默认使用的是 http 协议,如果用户对请求的安全性有要求,可以配置成 https。
服务端配置
修改 conf/rest-server.properties 配置文件,将 restserver.url 的 schema 部分改为 https。
由于 keystore 文件没有声明许可证,发行包中并不包含它。当 restserver.url 以 https 开头而 conf/hugegraph-server.keystore
不存在时,bin/start-hugegraph.sh 会在启动前从 hugegraph-doc 仓库的 binary-1.5 分支下载该文件,其密码为 hugegraph。
这两项都是 ssl.keystore_file 和 ssl.keystore_password 的默认值,用户可以生成自己的 keystore 文件及密码,然后修改这两个配置项。
客户端配置
在 HugeGraph-Client 中使用 https
在构造 HugeClient 时传入 https 相关的配置,代码示例:
注意:HugeGraph-Client 在 1.9.0 版本以前是直接以 new 的方式创建,并且不支持 https 协议,在 1.9.0 版本以后改成以 builder 的方式创建,并支持配置 https 协议。
在 HugeGraph-Loader 中使用 https
启动导入任务时,在命令行中添加如下选项:
hugegraph-loader 的 conf 目录下已经放了一个默认的客户端证书文件 hugegraph.truststore,其密码是 hugegraph。
在 HugeGraph-Tools 中使用 https
执行命令时,在命令行中添加如下选项:
hugegraph-tools 的 conf 目录下已经放了一个默认的客户端证书文件 hugegraph.truststore,其密码是 hugegraph。
如何生成证书文件
本部分给出生成证书的示例,如果默认的证书已经够用,或者已经知晓如何生成,可跳过。
服务端
- ⽣成服务端私钥,并且导⼊到服务端 keystore ⽂件中,server.keystore 是给服务端⽤的,其中保存着⾃⼰的私钥
过程中根据需求填写描述信息,默认证书的描述信息如下:
- 根据服务端私钥,导出服务端证书
server.crt 就是服务端的证书
客户端
client.truststore 是给客户端⽤的,其中保存着受信任的证书
5 - 配置 RocksDB 后端
概述
RocksDB 是一个嵌入式的 LSM-tree 键值存储。使用 rocksdb 后端时,HugeGraph-Server 把全部图数据保存在
服务进程内部的 RocksDB 实例中,不需要额外部署存储服务。发布包中的 conf/graphs/hugegraph.properties
默认使用的就是这个后端。
从 1.7.0 版本开始,服务端只接受 memory、rocksdb、hbase 和 hstore 作为后端。rocksdb 后端把
数据写在单台服务器的本地磁盘上,不支持共享存储,因此多个服务无法基于同一个数据目录提供同一个图。
分布式部署请使用 hstore 后端,配合 PD 与 Store。
RocksDB 的 JNI 依赖在 hugegraph-rocksdb/pom.xml 中固定为 8.10.2 版本,因此磁盘格式与各配置项的
语义都以 RocksDB 8.10 为准。
该后端上报的驱动版本是 1.11,在初始化图时会写入 system store 的 meta 表中。
选择后端
在图配置文件(conf/graphs/<graph>.properties)中设置后端与序列化器:
backend=rocksdb选择 RocksDB 存储实现。serializer=binary是发布包模板为该后端使用的序列化器。内置的序列化器为binary、binaryscatter和text。store是该图在后端中的库名,同时也是存储实现拿到的图名的一部分。
首次启动前执行一次 bin/init-store.sh 创建各个 store,然后再启动服务。bin/init-store.sh 与
bin/hugegraph-server.sh 都会加载 RocksDB 库,数据目录在执行这些脚本的机器上创建。
发布包会为打包时 backend.properties 中列出的每个后端注册配置项空间和存储实现,该文件的取值来自
hugegraph.backends 构建属性。默认构建会注册 rocksdb, hbase, hstore;使用 -Drocksdb-only 构建会
激活 rocksdb-only profile,产出的发布包只注册 rocksdb。未注册的后端在启动时会报
Not exists BackendStoreProvider。
注册过程还会额外注册一个名字 rocksdbsst,它对应的实现写出 SST 文件而不是打开一个可用的数据库。
该名字不在允许的后端列表中,因此 backend=rocksdbsst 会被拒绝并报 backend is illegal: rocksdbsst。
如果要把 SST 文件导入普通的 rocksdb 图,请使用下面介绍的 rocksdb.sst_path。
数据目录结构
有两个目录需要关注:rocksdb.data_path(默认 rocksdb-data/data)和 rocksdb.wal_path
(默认 rocksdb-data/wal)。相对路径基于服务的工作目录解析,也就是安装目录。
每个图会打开三个 store:m 存放 schema,g 存放图数据,s 是 system store。store 名会拼接到上面
两个路径之后,因此一个默认的单图安装目录如下:
后端的每张表在所属 store 中对应一个 RocksDB 列族,名字形如 <database>+<table>,其中 database 由图名
推导得到。已有数据目录中的列族总是会被重新打开,因此旧版本创建的表仍然可读。
还需要注意:
- 两个图不能共用同一个数据路径。通过 API 基于已有配置克隆创建图时,存储实现会在
rocksdb.data_path和rocksdb.wal_path后面追加_<newGraph>。删除这样的图会同时删除这两个目录。 - 快照创建在数据目录旁边:数据路径的最后两段会加上前缀重写,因此在默认路径下 graph store 的快照位于
rocksdb-data/<prefix>_data/g。恢复快照时会先关闭实例,删除数据目录,再把快照移动到原位置。 - 设置了
rocksdb.data_disks时,其中列出的表会在指定路径下作为独立的 RocksDB 实例打开,而不再放在rocksdb.data_path下。服务最多并发打开 8 个实例,打开最多等待 600 秒,会话关闭最多等待 30 秒。
路径与日志配置项
| config option | default value | description |
|---|---|---|
| rocksdb.data_path | rocksdb-data/data | RocksDB 数据存储路径,不允许为空。 |
| rocksdb.data_disks | [] | 为部分表指定独立磁盘,每个元素格式为 STORE/TABLE: /path/disk。允许的键为 [g/vertex, g/edge_out, g/edge_in, g/vertex_label_index, g/edge_label_index, g/range_int_index, g/range_float_index, g/range_long_index, g/range_double_index, g/secondary_index, g/search_index, g/shard_index, g/unique_index, g/olap]。磁盘路径不能与 rocksdb.data_path 相同。 |
| rocksdb.wal_path | rocksdb-data/wal | RocksDB WAL 存储路径,不允许为空。 |
| rocksdb.sst_path | (空) | 待导入 RocksDB 的 SST 文件所在路径,为空表示不导入。 |
| rocksdb.log_level | INFO | RocksDB 的日志级别,可选值:DEBUG、INFO、WARN、ERROR、FATAL、HEADER。 |
压缩与合并配置项
| config option | default value | description |
|---|---|---|
| rocksdb.num_levels | 7 | 数据库的层数,取值范围 1 到 2^31-1。 |
| rocksdb.compaction_style | LEVEL | RocksDB 的 compaction 策略:LEVEL/UNIVERSAL/FIFO。 |
| rocksdb.optimize_mode | true | 针对高负载和大数据量做优化,具体行为见下文的配置项生效方式一节。 |
| rocksdb.bulkload_mode | false | 切换到批量导入数据的模式。 |
| rocksdb.compression_per_level | [none, none, snappy, snappy, snappy, snappy, snappy] | 各层使用的压缩算法,可选值为 none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd。列表必须为空,或者元素个数正好等于 rocksdb.num_levels。 |
| rocksdb.bottommost_compression | none | 最底层使用的压缩算法,可选值为 none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd。 |
| rocksdb.compression | snappy | 压缩数据块使用的压缩算法,可选值为 none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd。 |
数据库级配置项
| config option | default value | description |
|---|---|---|
| rocksdb.max_background_jobs | 8 | 后台任务(包括 flush 和 compaction)的最大并发数,取值范围 1 到 2^31-1。 |
| rocksdb.max_subcompactions | 4 | 单个 compaction 任务使用的最大线程数,取值范围 1 到 2^31-1。 |
| rocksdb.delayed_write_rate | 16777216(16 MB/s) | 当 compaction 落后需要降速时,用户写请求的限速值,单位字节每秒。 |
| rocksdb.max_open_files | -1 | RocksDB 可缓存的最大打开文件数,-1 表示不限制。 |
| rocksdb.max_manifest_file_size | 104857600(100 MB) | manifest 文件的最大字节数。 |
| rocksdb.skip_stats_update_on_db_open | false | 打开数据库时是否跳过统计信息更新,设为 true 表示不更新统计信息。 |
| rocksdb.skip_check_sst_size_on_db_open | false | 打开数据库时是否跳过检查所有 sst 文件的大小。 |
| rocksdb.max_file_opening_threads | 16 | 打开文件使用的最大线程数,取值范围 1 到 2^31-1。 |
| rocksdb.max_total_wal_size | 0 | WAL 文件的总大小上限,单位字节。超过后会强制 flush 相关的列族,0 表示不限制。 |
| rocksdb.bytes_per_sync | 0 | 允许操作系统在后台异步地增量同步正在写入的 SST 文件,每写入这么多字节发起一次请求,0 表示关闭。 |
| rocksdb.wal_bytes_per_sync | 0 | 同上,作用于 WAL 文件,0 表示关闭。 |
| rocksdb.strict_bytes_per_sync | false | 为 true 时保证任意时刻提交回写的 SST/WAL 数据不超过 bytes_per_sync/wal_bytes_per_sync 字节,可用于处理写入速度超过 I/O 速度的场景。 |
| rocksdb.db_write_buffer_size | 0 | 所有列族 write buffer 的总大小上限,单位字节,0 表示不限制。 |
| rocksdb.log_readahead_size | 0 | 读取日志时预取的字节数,0 表示关闭预取。 |
| rocksdb.compaction_readahead_size | 0 | compaction 时批量读取的字节数。如果 RocksDB 跑在机械盘上,建议至少设为 2MB,0 表示关闭预取。 |
| rocksdb.row_cache_capacity | 0 | 表级行缓存的全局容量,单位字节,0 表示关闭 row_cache。 |
| rocksdb.delete_obsolete_files_period | 21600(6 小时) | 删除废弃文件的周期,单位秒,0 表示每次都做完整清理。该值传给 RocksDB 前会换算成微秒。 |
Memtable 配置项
| config option | default value | description |
|---|---|---|
| rocksdb.write_buffer_size | 134217728(128 MB) | 在内存中累积的数据量,单位字节,最小 1 MB。该值针对单个列族生效。 |
| rocksdb.max_write_buffer_number | 6 | 内存中累积的 write buffer 的最大个数,取值范围 1 到 2^31-1。 |
| rocksdb.min_write_buffer_number_to_merge | 2 | 会被合并到一起的 write buffer 的最小个数,取值范围 1 到 2^31-1。 |
| rocksdb.max_write_buffer_number_to_maintain | 0 | 使用事务时,为冲突检查在内存中保留的 write buffer 总数上限。 |
| rocksdb.memtable_bloom_size_ratio | 0.0 | 设置了 prefix-extractor 且该值不为 0,或者 memtable_whole_key_filtering 为 true 时,为 memtable 创建大小为 write_buffer_size * memtable_bloom_size_ratio 的布隆过滤器。大于 0.25 的取值会被收敛为 0.25,取值范围 0.0 到 1.0。 |
| rocksdb.memtable_whole_key_filtering | false | 在 memtable 中启用全键布隆过滤器,可以降低点查的 CPU 开销。只有 memtable_bloom_size_ratio > 0 时才生效。 |
| rocksdb.memtable_huge_page_size | 0 | memtable 中布隆过滤器使用的 huge page TLB 的页大小,小于等于 0 时不从 huge page TLB 分配,改用 malloc。 |
| rocksdb.inplace_update_support | false | 当写入的键已存在于当前 memtable 且新值更小时,允许线程安全的原地更新。 |
层级大小与写入限流配置项
| config option | default value | description |
|---|---|---|
| rocksdb.level_compaction_dynamic_level_bytes | false | 是否启用 level_compaction_dynamic_level_bytes。启用后 max_bytes_for_level_multiplier 的优先级高于 max_bytes_for_level_base,基准层的大小是动态的,LSM tree 更可预测,有助于限制最差情况下的空间放大。对已有数据库开关该特性可能造成异常的 LSM tree 结构,因此不推荐改动。 |
| rocksdb.max_bytes_for_level_base | 536870912(512 MB) | level-1 文件总大小的上限,单位字节,最小 1 MB。 |
| rocksdb.max_bytes_for_level_multiplier | 10.0 | 对所有层 L,第 L+1 层文件总大小与第 L 层文件总大小的比值,最小 1.0。 |
| rocksdb.target_file_size_base | 67108864(64 MB) | compaction 的目标文件大小,单位字节,最小 1 MB。 |
| rocksdb.target_file_size_multiplier | 1 | 第 L 层文件与第 L+1 层文件的大小比值。 |
| rocksdb.level0_file_num_compaction_trigger | 2 | 触发 level-0 compaction 的文件数。 |
| rocksdb.level0_slowdown_writes_trigger | 20 | 触发写入降速的 level-0 文件数软上限。 |
| rocksdb.level0_stop_writes_trigger | 36 | 触发停写的 level-0 文件数硬上限。 |
| rocksdb.soft_pending_compaction_bytes_limit | 68719476736(64 GB) | 待 compaction 数据量的软上限,单位字节,最小 1 GB。 |
| rocksdb.hard_pending_compaction_bytes_limit | 274877906944(256 GB) | 待 compaction 数据量的硬上限,单位字节,最小 1 GB。 |
文件 I/O 配置项
| config option | default value | description |
|---|---|---|
| rocksdb.allow_mmap_writes | false | 允许操作系统以 mmap 方式写文件。 |
| rocksdb.allow_mmap_reads | false | 允许操作系统以 mmap 方式读取 sst 文件。 |
| rocksdb.use_direct_reads | false | 读取 sst 文件时使用直接 I/O。 |
| rocksdb.use_direct_io_for_flush_and_compaction | false | flush 和 compaction 时使用直接读写。 |
| rocksdb.use_fsync | false | 为 true 时每次持久化都会执行 fsync。 |
| rocksdb.atomic_flush | false | 为 true 时多个列族的 flush 结果会原子地提交到 MANIFEST。WAL 一直开启的情况下不必设置该项。 |
SST 表格式与块缓存配置项
| config option | default value | description |
|---|---|---|
| rocksdb.format_version | 5 | BlockBasedTable 的格式版本,可选值为 0~5。 |
| rocksdb.index_type | kBinarySearch | sst 文件中数据块之间查找使用的索引类型,可选值为 [kBinarySearch, kHashSearch, kTwoLevelIndexSearch, kBinarySearchWithFirstKey]。 |
| rocksdb.data_block_index_type | kDataBlockBinarySearch | sst 文件数据块内点查使用的查找类型,可选值为 [kDataBlockBinarySearch, kDataBlockBinaryAndHash]。 |
| rocksdb.data_block_hash_table_util_ratio | 0.75 | 哈希表 entries/buckets 的使用率,仅在 data_block_index_type=kDataBlockBinaryAndHash 时有效,取值范围 0.0 到 1.0。 |
| rocksdb.block_size | 4096(4 KB) | 每个块中打包的用户数据的近似大小,注意对应的是未压缩的数据。 |
| rocksdb.block_size_deviation | 10 | 用于结束一个块的空闲空间百分比,取值范围 0 到 100。 |
| rocksdb.block_restart_interval | 16 | 块内增量编码的 restart 间隔。 |
| rocksdb.block_cache_capacity | 8388608(8 MB) | RocksDB 使用的块缓存大小,单位字节,0 表示不使用块缓存。每个列族都会创建一个该大小的独立缓存。 |
布隆过滤器配置项
只有 rocksdb.bloom_filter_bits_per_key 大于等于 0 时,本组的配置项才会被读取。默认值 -1 表示不启用
布隆过滤器,此时本表中其他配置项都不生效,包括索引和过滤块的缓存相关项。
| config option | default value | description |
|---|---|---|
| rocksdb.bloom_filter_bits_per_key | -1 | 布隆过滤器中每个键占用的位数,10 是一个不错的取值,对应约 1% 的误判率。大于 0 表示启用布隆过滤器,-1 表示不启用(0~0.5 向下取整为不启用)。 |
| rocksdb.bloom_filter_block_based_mode | false | 启用布隆过滤器时,设为 true 表示使用 block based filter 而不是 full filter。 |
| rocksdb.bloom_filter_whole_key_filtering | true | 启用布隆过滤器时,设为 true 表示把完整的键放入布隆过滤器,否则在设置了 prefix-extractor 时放入键的前缀。 |
| rocksdb.cache_index_and_filter_blocks | true | 设为 true 表示把索引块和过滤块放入块缓存。 |
| rocksdb.pin_l0_filter_and_index_blocks_in_cache | true | 设为 true 表示把 L0 的索引块和过滤块固定在块缓存中。 |
| rocksdb.optimize_filters_for_hits | true | 启用布隆过滤器时,该开关允许不为最后一层存储过滤器,设为 true 表示主要针对命中的场景优化过滤器,而不是同时优化未命中的场景。该项在过滤器关闭时也会生效。 |
| rocksdb.partition_filters_and_indexes | false | 启用布隆过滤器时,设为 true 表示每个 sst 文件使用分区的 full filter 和索引。该项与 block based filter 不兼容。开启后索引类型会被强制为 kTwoLevelIndexSearch,元数据块大小取 rocksdb.block_size。 |
| rocksdb.pin_top_level_index_and_filter | true | 当 partition_filters_and_indexes 为 true 时,设为 true 表示把分区过滤块和索引块的顶层索引固定在块缓存中。 |
| rocksdb.prefix_extractor_n_bytes | 0 | prefix-extractor 取键的前 N 个字节作为前缀,键长度小于 N 时使用完整键,0 表示不设置 prefix-extractor。 |
配置项的生效方式
服务为每个 store 和每个列族构建一次 RocksDB 的选项对象,因此上面任何配置项的修改都在下次启动服务时 生效。
rocksdb.optimize_mode=true会在应用上表中的取值之前先套用一组预设:数据库层面把并行度提高到可用 处理器数的一半(至少 1),允许 memtable 并发写入,并开启写线程自适应让出;列族层面调用 RocksDB 的 level-style 和 universal-style compaction 预设。显式配置项在之后应用,因此配置文件中写明的取值会覆盖 预设。rocksdb.bulkload_mode=true会关闭自动 compaction,把三个 level-0 触发阈值提高到 int 最大值,把两个 待 compaction 上限提高到 long 最大值。导入结束后要关闭它并重启,否则 compaction 不会运行。rocksdb.block_cache_capacity=0表示彻底关闭块缓存,而不是不限制大小。rocksdb.prefix_extractor_n_bytes大于 0 时会安装一个该长度的 capped prefix extractor。- 所有列族都使用
uint64addmerge 操作符,计数器表依赖它。 - 数据库不存在时会自动创建,
avoid_unnecessary_blocking_io和write_dbid_to_manifest始终开启。
内存说明
RocksDB 的缓存和 write buffer 都是本地内存分配,不属于 bin/hugegraph-server.sh 中设置的 JVM 堆。
GET /metrics/backend 接口会返回存储的使用量:内存数值是所有已打开列族的块缓存用量、固定在块缓存中的
用量、预估的 table reader 内存(索引块和过滤块)以及全部 memtable 大小之和,取自 RocksDB 的属性。
有两个配置项的实际占用会随列族数量成倍增长:
rocksdb.block_cache_capacity为每个列族创建一个缓存实例,因此一台服务的块缓存总量大致等于该值乘以 所有图的m、g、s三个 store 中已打开表的数量,再加上rocksdb.data_disks额外打开的实例。rocksdb.write_buffer_size乘以rocksdb.max_write_buffer_number限定的是单个列族的 memtable 内存。rocksdb.db_write_buffer_size限制一个 store 内所有列族的总量,默认值 0 表示没有这个限制。
rocksdb.row_cache_capacity 不同:它是每个 store 一个缓存,0 表示关闭。
导入 SST 文件
设置 rocksdb.sst_path 即开启导入。打开 store 时以及每次创建表时,服务会遍历
<sst_path>/<column family>/ 目录,收集其中所有非空的 *.sst 文件,导入到对应的列族。导入采用移动
文件的方式而不是复制,因此源目录会被导入过程消耗掉。
raft 模式
RocksDB 后端仍然可以运行在 raft 状态机之后:raft.mode=true 时,本地后端的存储实现会被 raft 实现包装。
包装层会拒绝共享存储的后端,因此 rocksdb 可用而 hbase 不可用。raft 模式下 RocksDB 会话写入时关闭
WAL 且不做 sync,因为状态机可以通过快照加 raft 日志恢复,而该后端支持快照。
使用时需要注意:
bin/init-store.sh在初始化后端时会强制把raft.mode置为 false,因此初始化过程不会走 raft。- 发布包中的
conf/graphs/hugegraph.properties已把 raft 相关配置标记为废弃。1.7.0 及之后版本的分布式 部署改用hstore后端,配合 PD 与 Store。 - raft 成员管理接口位于
graphspaces/{graphspace}/graphs/{graph}/raft/之下,包括list_peers、get_leader、set_leader、transfer_leader、add_peer和remove_peer。bin/raft-tools.sh封装了 同样的操作,但它拼接的 URL 中仍然没有 graphspace 段,在 1.7.0 的服务上需要调整路径才能使用。 - 其余
raft.*配置项见 Server 配置选项。
后端能力
该后端的特性开关决定了哪些操作可以下推给存储:
- 支持按键前缀扫描、按键范围扫描、分页查询、范围条件和 order-by。
- RocksDB 内部没有索引,因此按名字查询 schema、按标签查询以及按标签删除边由服务端完成,而不是由存储 完成。
- 通过 RocksDB 的 write batch 支持事务。
- 支持快照,raft 模式和备份依赖该能力。
- 不支持共享存储,一个数据目录属于一台服务。
- 支持 olap 属性,对应的表会作为额外的列族创建。
- 存储本身不会让数据过期,因此服务端在读取时过滤掉 TTL 已到期的元素。
- 存储层不支持
in、contains、contains_key条件,不支持聚合属性,也不支持原地更新顶点或边的属性。
riscv64 平台说明
在 Linux riscv64 上,RocksDB 的 JNI 库需要 libatomic.so.1。bin/util.sh 会查找该库并在
bin/hugegraph-server.sh、bin/init-store.sh 和 bin/dump-store.sh 启动 JVM 之前把它加入
LD_PRELOAD。如果找不到,这些脚本会以
RISC-V RocksDB requires libatomic.so.1; install libatomic1 退出,安装 libatomic1 包即可解决。
6 - 配置 HStore 分布式后端
1 概述
hstore 是 HugeGraph 的分布式存储后端。图使用该后端时,HugeGraph-Server 本地磁盘上不保存任何图数据,
数据由另外两个进程负责:
- HugeGraph-PD(Placement Driver)保存集群元数据:已注册的 Store 列表、每个图的分区布局、分区到 Store 的映射关系、图的 Schema 以及 Schema 的 id 计数器。
- HugeGraph-Store 保存实际的键值数据,并通过 Raft 在多个 Store 节点之间复制。
Server 进程内嵌了 PD 客户端和 Store 客户端。每次读写时,它先向 PD 查询该 key 属于哪个分区、当前哪个 Store 节点是这个分区的 leader,然后把请求直接发给这个 Store 节点。
服务端的适配层是 hugegraph-hstore 模块,它以后端名 hstore 注册,驱动版本为 1.13。
选择 hstore 影响的不只是数据写到哪里,Server 还会根据后端类型切换下列行为:
| 方面 | 使用 hstore | 使用本地后端 |
|---|---|---|
| Schema 存储 | 通过 PD 元数据驱动读写 Schema | Schema 保存在 m store 中 |
| Schema id | 通过 PD 客户端由 PD 分配 | 由 schema store 分配 |
| System store | 没有独立的 system store,系统数据写入 graph store | 独立的 s store |
| 任务调度器 | distributed | local |
| 权限管理器 | StandardAuthManagerV2 | StandardAuthManager |
| 后端版本校验 | 读取 graph store | 读取 system store |
init-store.sh | 跳过该图,元数据由 PD 和 Store 负责 | 创建本地 store |
2 前置条件
hstore 不能独立工作。在 Server 打开 hstore 图之前,PD 集群和至少一个 Store 节点必须已经运行,
并且启动顺序如下:
- PD,先启动以便组成 Raft 组。
- Store,通过 gRPC 向 PD 注册。gRPC 地址出现在 PD 自身
pd.initial-store-list中的 Store 会直接进入Up状态;不在该列表中、并且 PD 从未见过它处于Up或Offline的 Store 会注册为Pending, 需要先激活才能提供数据服务。 - Server,随后从 PD 读回 Store 列表。
服务端需要关注的默认端口:
| 进程 | gRPC 端口 | REST 端口 |
|---|---|---|
| PD | 8686 | 8620 |
| Store | 8500 | 8520 |
服务端的 pd.peers 指向 PD 的 gRPC 端口,而不是 REST 端口。
另外两个进程的安装与配置方式,参见 安装/构建 HugeGraph-PD 和 安装/构建 HugeGraph-Store。
3 选择 hstore 后端
3.1 图配置文件
在图的属性文件(例如 conf/graphs/hugegraph.properties)中设置后端:
关于这四个配置项:
backend=hstore选择该适配层。自 1.7.0 起允许的取值为memory、rocksdb、hbase和hstore。 发行包中做校验的位置对该值不区分大小写。serializer=binary是必需的。注册hstore后端时只注册了配置空间和存储 provider,并没有注册自己的 序列化器,适配层就是按二进制序列化器编写的。serializer的内置默认值是text,因此必须显式写出该项。store=hugegraph是 PD 看到的图名中的命名空间部分。Server 以<graphspace>/<store>打开 provider, 每个底层 store 再追加自己的后缀,因此 PD 中每个 store 对应一个图条目:图数据是DEFAULT/hugegraph/g,schema store 位是DEFAULT/hugegraph/m。graphspace默认为DEFAULT,g和m是固定的。pd.peers是以逗号分隔的 PD gRPC 地址列表。适配层从图配置中读取该项,而不是从rest-server.properties中读取,图级别的元数据连接也使用同一个值。
如果图配置文件中没有 pd.peers,那么在加载图时,只要 usePD 为 true 或者后端是 hstore,
Server 会把 rest-server.properties 中的值复制到图配置里。不过在图配置文件中显式写出该项更清晰。
3.2 rest-server.properties
usePD=true 让 Server 在启动时从 PD 加载元数据。在这条路径上,它会把元数据管理器连接到 PD,创建内置的
admin 账号和默认图空间,加载图空间与服务,创建内部的系统图(后端固定为 hstore),并加载 PD 中保存的
图配置。
它和图级别的 backend=hstore 是两个独立的开关:一个图可以使用 hstore 而 usePD 保持默认的 false,
此时 Server 不会走基于 PD 的元数据路径。发行包自带的测试启动脚本在后端为 hstore 时会设置该项。
3.3 发行包中的模板文件
发行包在 conf/graphs/hstore.properties.template 中提供了一份该后端的现成图配置文件。它与
hugegraph.properties 的差别是:把 backend 设为 hstore、不注释 pd.peers=127.0.0.1:8686、
并且不包含内存管理配置段。
hstore 的 Docker 镜像会自动套用这份模板:它删除 conf/graphs/hugegraph.properties,再把模板重命名过去,
因此容器启动时就已经选好了 hstore 后端。
本地构建的发行包默认编译了 hstore provider。rocksdb-only 这个 Maven profile 会把编译进去的后端列表
收窄为只有 rocksdb,用这种方式构建出来的发行包会以 Unsupported backend type 拒绝 backend=hstore。
4 hstore 配置项
hstore 配置空间中只有下面两个配置项,它们写在图的属性文件里。
| 配置项 | 默认值 | 说明 |
|---|---|---|
| hstore.partition_count | 0 | 分区数量,PD 依据该值控制分区(Number of partitions)。 |
| hstore.shard_count | 0 | 副本数量,PD 依据该值控制分区副本(Number of copies)。 |
4.1 hstore.partition_count
每个 graph store 第一次被打开时,Server 会把这个数字连同图名一起发给 PD。取负值会在此处被拒绝,
报错信息为 The value of hstore.partition_count cannot be less than 0.
PD 对该值的处理方式:
0,也就是默认值,表示交给 PD 决定。对图数据 store,PD 使用自身集群级别的分区总数,该总数由pd.initial-store-list中的条目数量、partition.store-max-shard-count和partition.default-shard-count推算得出;对/m和/sstore 固定使用1。- 取值在
1到该总数之间时,按原值使用。 - 取值大于该总数时,会被下调到该总数。
该数字在 store 首次向 PD 注册时生效,之后再修改属性文件不会让已有的图重新分区。
4.2 hstore.shard_count
hstore.shard_count 声明在 hstore 配置空间中,属性文件里也接受该项,但当前版本服务端没有任何代码读取它:
适配层读取的只有 hstore.partition_count 一项。实际生效的副本数由 PD 的配置决定,即 PD application.yml
中的 partition.default-shard-count。
5 只在 hstore 模式下生效的其他配置项
下列配置项位于公共的 rest-server.properties 和图属性文件中,但只有在使用 PD 和 hstore 后端时才生效,
或者才会改变行为。source 列给出该配置项在 HugeGraph master 分支上的声明位置(文件与行号)。
| 配置项 | 文件 | 默认值 | 在 hstore 模式下的作用 | source |
|---|---|---|---|---|
| pd.peers | rest-server.properties | 127.0.0.1:8686 | 用于元数据、服务发现和系统图的 PD 地址 | ServerOptions.java:195-201 |
| pd.peers | {graph}.properties | 127.0.0.1:8686 | 后端适配层自身使用的 PD 地址 | CoreOptions.java:649-654 |
| usePD | rest-server.properties | false | Server 启动时是否从 PD 加载元数据 | ServerOptions.java:390-396 |
| cluster | rest-server.properties | hg-test | 集群名,作为所有 PD 元数据 key 的前缀 | ServerOptions.java:187-193 |
| init_store.enabled | rest-server.properties | true | PD/Store 部署下应设为 false,元数据已由存储侧负责 | ServerOptions.java:371-380 |
| graph.load_from_local_config | rest-server.properties | false | 启动时是否在 PD 中的图配置之外,额外扫描 conf/graphs | ServerOptions.java:355-361 |
| auth.graph_store | rest-server.properties | hugegraph | 保存权限数据的图,关闭 init-store 时会校验它使用 hstore 后端 | ServerOptions.java:591-598 |
| graphspace | {graph}.properties | DEFAULT | PD 看到的图名的第一段 | CoreOptions.java:679-685 |
init-store.sh 从不初始化 hstore 图。在开启的路径上,它扫描 conf/graphs 并跳过后端为 hstore 的每一个
图。如果用 init_store.enabled=false 整体关闭这一步,它会改为校验 admin 账号仍然能在 PD 启动路径上被创建:
usePD 必须为 true、权限图必须存在于本地配置中且后端为 hstore、auth.admin_pa 必须显式设置为非空值。
否则启动会直接失败,而不是使用公开的默认密码创建账号。
6 Server 如何通过 PD 发现 Store
适配层在进程中第一次打开 hstore 图时,一次性构建这些客户端:
- 用
pd.peers构建 PD 客户端配置,带上 PD 的鉴权凭据,并开启客户端侧的分区缓存。 - 创建进程级的 PD 客户端。
- 用该 PD 客户端创建进程级的 Store 客户端。
创建 Store 客户端时,会把一个基于 PD 的分区器同时注册为 Store 客户端节点管理器的 node provider、 partitioner 和 notifier。路由逻辑全部在这个分区器中:
- 单点和前缀请求:向 PD 查询拥有该 key 的分区,取该分区的 leader 副本,把请求发到对应的 store id。
- 按 code 的范围扫描:按 code 逐个遍历分区直到覆盖整个范围,每个分区产生一个目标 Store。
- 全图扫描:向 PD 查询该图的活跃 Store,并向全部 Store 扇出请求。
- Store 地址解析:通过 PD 把 store id 解析成主机和端口。
- 缓存失效:当某个 Store 返回分区 leader 已迁移时,notifier 会更新 PD 客户端缓存中的分区 leader 并使过期的分区条目失效,之后的请求就会跟随新的 leader。
由于 Store 列表来自 PD 而不是配置文件,增删 Store 节点只需要针对同一个 PD 集群启动或停止它, 服务端不需要改任何配置。
7 后端能力
hstore 并不支持本地后端的所有查询形式。对用户可见的差异如下:
| 特性 | 是否支持 |
|---|---|
| 按 key 前缀扫描 | 支持 |
| 按 key 范围扫描 | 支持 |
| 带范围条件的查询 | 支持 |
| 带 order by 的查询 | 支持 |
| 分页查询 | 支持 |
| OLAP 属性 | 支持 |
| Task 和 Server 顶点 | 支持 |
| Scan token | 不支持 |
| 按名称查询 Schema | 不支持 |
| 按 label 查询 | 不支持 |
带 in 条件的查询 | 不支持 |
带 contains 的查询 | 不支持 |
带 contains key 的查询 | 不支持 |
| 按输入 id 顺序排序 | 不支持 |
| 按 label 删除边 | 不支持 |
| 更新顶点属性 | 不支持 |
| 更新边属性 | 不支持 |
| 事务 | 不支持 |
| Number 类型 | 不支持 |
| 聚合属性 | 不支持 |
| TTL | 不支持 |
不支持按输入 id 顺序排序,是因为多节点批量扫描会按 Store 对输入 key 分组,从而丢失全局顺序; 不支持更新顶点和边属性,是因为属性被存放在单个 cell 中。
8 验证
Server 启动后,后端指标接口会返回 PD 当前认为处于活跃状态的 Store 数量:
响应中的 nodes 就是 PD 返回的活跃 Store 数量。nodes 为 0 说明 Server 连上了 PD,但 PD 中没有状态为
Up 的 Store,通常是 Store 节点还没注册,或者因为不在 PD 的 pd.initial-store-list 中而注册成了
Pending。
7 - 配置 HBase 后端
概述
HBase 后端将图数据存储在 Apache HBase 表中。HugeGraph 仅作为 HBase 客户端:它通过 HBase 的 ZooKeeper 集群连接,为每个图创建一个 HBase namespace,并在其中创建该图的 schema 表、数据表和索引表。计数查询由 HBase 的 AggregateImplementation 协处理器完成,HugeGraph 在创建每张表时都会挂载该协处理器。
注意:HBase 后端已废弃,计划在 HugeGraph 2.0 中移除。新部署请使用
hstore(分布式)或rocksdb(内嵌,默认值),已有的 HBase 部署请规划迁移。
自 1.7.0 起,发行包内置的后端只有 hstore、rocksdb、hbase 和 memory。HBase provider 上报的后端驱动版本为 1.12。
支持的 HBase 版本
客户端 jar 固定为 HBase 2.6.5(hbase-endpoint 加 hbase-shaded-client)。服务端要求 HBase 2.x:当检测到的 HBase 版本低于 2.0 时,scan 逻辑会把 inclusive stop row 改写为 exclusive 并追加一个 0 字节,因为该版本之前 inclusive stop row 不生效。CI 任务和本地 Docker 镜像都使用 HBase 2.6.5,后端也是针对这个版本做测试的。
选择该后端
修改需要使用 HBase 的图的 conf/graphs/hugegraph.properties:
注意:
serializer必须设置为hbase,而不是binary。HBase 序列化器是BinarySerializer的子类,它不在 rowkey 中写入 id 前缀,并写入预分区的顶点表和边表所需要的分区前缀。使用serializer=binary时这两点都不生效。
然后初始化后端并启动服务:
默认发行包构建时包含 rocksdb, hbase, hstore 三个后端,无需额外引入 jar。使用 rocksdb-only Maven profile 构建的发行包不包含 HBase 后端,此时 backend=hbase 会以 Not exists BackendStoreProvider: hbase 打开失败。
下面所有配置项都位于图配置文件(conf/graphs/hugegraph.properties)中,而不是 rest-server.properties。只有当发行包包含 hbase 后端时,这些配置项才会被注册。
连接配置项
| 配置项 | 默认值 | 说明 |
|---|---|---|
| hbase.hosts | localhost | HBase ZooKeeper 的主机名或 IP 地址,多个以逗号分隔,不允许为空。对应 hbase.zookeeper.quorum。 |
| hbase.port | 2181 | HBase ZooKeeper 的端口,取值范围 1 到 65535。对应 hbase.zookeeper.property.clientPort。 |
| hbase.znode_parent | /hbase | HBase ZooKeeper 的 znode 父路径,不允许为空。对应 zookeeper.znode.parent。 |
| hbase.zk_retry | 3 | HBase ZooKeeper 的恢复重试次数,取值范围 0 到 1000。对应 zookeeper.recovery.retry。 |
| hbase.threads_max | 64 | HBase 连接的最大线程数,取值范围 1 到 1000。对应 hbase.hconnection.threads.max,HBase 自身默认值为 256,这里取更小的值以避免内存溢出。 |
超时配置项
| 配置项 | 默认值 | 说明 |
|---|---|---|
| hbase.truncate_timeout | 30 | 等待后端 truncate 的超时时间,单位秒,必须为正数。该超时按 store 计算,而一个图有三个 store,因此一次 truncate 最多耗时该值的三倍。 |
| hbase.aggregation_timeout | 43200(12 小时) | 等待聚合的超时时间,单位秒,必须为正数。它会设置计数查询所用聚合客户端的 hbase.rpc.timeout。 |
Kerberos 与 HBase 配置文件配置项
| 配置项 | 默认值 | 说明 |
|---|---|---|
| hbase.kerberos_enable | false | 是否为 HBase 启用 Kerberos 认证。 |
| hbase.krb5_conf | /etc/krb5.conf | Kerberos 配置文件,包含 KDC IP、默认 realm 等。会被设置为 java.security.krb5.conf 系统属性。 |
| hbase.hbase_site | /etc/hbase/conf/hbase-site.xml | HBase 的配置文件。无论是否启用 Kerberos,每次建立连接时都会把它作为配置资源加载。 |
| hbase.kerberos_principal | (空) | Kerberos 认证使用的 HBase principal。 |
| hbase.kerberos_keytab | (空) | Kerberos 认证使用的 HBase keytab 文件。 |
当 hbase.kerberos_enable=true 时,HugeGraph 会在连接上把 hadoop.security.authentication 和 hbase.security.authentication 设置为 kerberos,然后在打开连接之前用配置的 principal 从 keytab 登录。因此 Kerberos 环境下 hbase.krb5_conf、hbase.hbase_site、hbase.kerberos_principal 和 hbase.kerberos_keytab 四项都必须有效:
即使关闭 Kerberos,hbase.hbase_site 也会被读取,路径不存在时相当于加载了一个空资源。当需要上述配置项之外的 HBase 设置时,把它指向集群自身的 hbase-site.xml。
预分区配置项
| 配置项 | 默认值 | 说明 |
|---|---|---|
| hbase.enable_partition | true | 是否为 HBase 启用预分区。它同时决定后端是否声明支持前缀扫描和范围扫描。 |
| hbase.vertex_partitions | 10 | HBase 顶点表的分区数,不允许为负数。 |
| hbase.edge_partitions | 30 | HBase 边表的分区数,不允许为负数。 |
启用预分区后,顶点表按 hbase.vertex_partitions 个 region 创建,两张边表各按 hbase.edge_partitions 个 region 创建,序列化器会在 rowkey 前面加上 id 哈希得到的分区前缀。
注意:请在初始化后端之前,按实际数据量和 region server 数量调整分区数。它对导入速度影响很大,并且只在建表时生效。
关闭 hbase.enable_partition 会恢复不带前缀的原始 rowkey。作为交换,后端此时会声明支持前缀扫描和范围扫描,这两类扫描在预分区 rowkey 下无法工作。
Namespace 与表结构
每个图对应一个 HBase namespace,名称为 <graphspace>/<store> 转小写,并把 / 替换为 _,因为 HBase namespace 名称只允许字母数字和 _ 字符。在默认配置 graphspace=DEFAULT、store=hugegraph 下,namespace 为 default_hugegraph。
在该 namespace 内,一个图包含三个 store:schema store m、graph store g 和 system store s:
| Store | 表 |
|---|---|
schema (m) | VL、EL、PK、IL、C、m_si |
graph (g) | g_v、g_oe、g_ie、g_si、g_vi、g_ei、g_ii、g_fi、g_li、g_di、g_ai、g_hi、g_ui |
system (s) | s_v、s_oe、s_ie、s_si、s_vi、s_ei、s_ii、s_fi、s_li、s_di、s_ai、s_hi、s_ui、M |
g_v 是顶点表,g_oe 和 g_ie 分别是出边表和入边表,其余 g_* 表依次是二级索引、顶点标签索引、边标签索引、范围索引(int、float、long、double)、全文索引、shard 索引和唯一索引表。所有表都只有一个名为 f 的列族,并且都在建表时挂载了 org.apache.hadoop.hbase.coprocessor.AggregateImplementation 协处理器。只有 g_v、g_oe 和 g_ie 会预分区,system store 中同名的那几张表按单个 region 创建。
system store 中的 M 表保存 init-store.sh 写入的后端版本。truncate 图时会排除该表,因为丢失它会导致下次启动的版本校验失败。清空图会删除这些表;连同存储空间一起清空则会删除整个 namespace。
GET /metrics/backend 会返回 HBase 集群状态:cluster_id、master_name、average_load、hbase_version、region_count、leaving_servers、nodes、region_servers,以及一个 servers map,其中包含每个 region server 的堆内存、磁盘、请求数和 region 明细。PUT /graphspaces/{graphspace}/graphs/{name}/compact 会请求 HBase 对该图的所有表做 compaction。
使用 Docker 做本地测试
Server 仓库中的 docker/hbase 会构建一个 HBase 2.6.5 单机镜像(hugegraph/hbase:2.6.5,容器名 hg-hbase-test),用于本地开发和测试。以下命令都在仓库根目录执行。
为运行在宿主机上的 HugeGraph 启动 HBase:
为运行在同一个 Docker 网络中的容器化 HugeGraph 启动 HBase:
对外公布的主机名很重要:容器启动时会把 HBASE_MASTER_HOSTNAME 和 HBASE_REGIONSERVER_HOSTNAME 写入自己的 hbase-site.xml,未设置时回退到 HBASE_HOSTNAME(默认 hbase)。如果客户端无法解析这个主机名,即使 ZooKeeper 可用,也会报 UnknownHostException: hbase:16000。
映射到宿主机的端口:
| 端口 | 服务 |
|---|---|
| 2181 | ZooKeeper,与 hbase.port 默认值一致 |
| 16000 | HBase Master RPC |
| 16010 | HBase Master Web UI,http://localhost:16010 |
| 16020 | HBase RegionServer RPC |
| 16030 | HBase RegionServer Web UI,http://localhost:16030 |
针对它运行后端测试:
停止并删除数据卷:
该镜像会分别启动 ZooKeeper、master 和 region server 三个守护进程,并等到 master 上报有存活的 server 之后才开始 tail 日志,因此首次启动会比较慢。请给 Docker 分配至少 4 GB 内存。compose 的健康检查也因此设置了 90 秒的 start period。
限制
HBase 后端不支持以下特性:
- 事务。rollback 只会丢弃尚未提交的批次,而 commit 是逐表写入的,因此跨表不是原子的。
- 原地更新单个顶点或边属性,以及合并顶点属性。属性存放在一个 cell 中,因此会重写整个属性列。
- 按名称查询 schema,以及仅按标签查询顶点或边。这两者都需要 HBase 二级索引。
- 按标签删除边。
- 带
in条件、contains条件或contains_key条件的查询。 - 聚合属性和 OLAP 属性。
- 原生数值类型(后端特性
supportsNumberType为关闭状态)。 - scan token。
hbase.enable_partition为true时的前缀扫描和范围扫描。- 除
count以外的聚合函数,其它聚合函数会被拒绝。 - 快照。创建或恢复后端快照会抛出
UnsupportedOperationException。
已支持的特性包括顶点和边的 TTL、分页查询、order by 查询、范围条件,以及按输入 id 排序。