跳转到主要内容

1 - Server 启动指南

1 概述

配置文件的目录为 hugegraph-release/conf,所有关于服务和图本身的配置都在此目录下。

主要的配置文件包括:gremlin-server.yaml、rest-server.properties 和 hugegraph.properties

HugeGraphServer 内部集成了 GremlinServer 和 RestServer,而 gremlin-server.yaml 和 rest-server.properties 就是用来配置这两个 Server 的。

  • GremlinServer:GremlinServer 接收 Gremlin 请求并调用图引擎。
  • RestServer:提供 RESTful API,根据不同的 HTTP 请求,调用对应的 Core API,如果用户请求体是 gremlin 语句,则会转发给 GremlinServer,实现对图数据的操作。

下面对这三个配置文件逐一介绍。

2 gremlin-server.yaml

gremlin-server.yaml 的主要结构如下。示例省略了部分导入项;完整内容以发布包中的文件为准。

conf/gremlin-server.yaml
# host and port of gremlin server, need to be consistent with host and port in rest-server.properties
#host: 127.0.0.1
#port: 8182

# Gremlin 查询中的超时时间(以毫秒为单位)
evaluationTimeout: 30000

channelizer: org.apache.tinkerpop.gremlin.server.channel.WsAndHttpChannelizer
# 不要在此处设置图形,此功能将在支持动态添加图形后再进行处理
graphs: {
}
scriptEngines: {
  gremlin-groovy: {
    staticImports: [
      org.opencypher.gremlin.process.traversal.CustomPredicates.*',
      org.opencypher.gremlin.traversal.CustomFunctions.*
    ],
    plugins: {
      org.apache.hugegraph.plugin.HugeGraphGremlinPlugin: {},
      org.apache.tinkerpop.gremlin.server.jsr223.GremlinServerGremlinPlugin: {},
      org.apache.tinkerpop.gremlin.jsr223.ImportGremlinPlugin: {
        classImports: [
          java.lang.Math,
          org.apache.hugegraph.backend.id.IdGenerator,
          org.apache.hugegraph.type.define.Directions,
          org.apache.hugegraph.type.define.NodeRole,
          org.apache.hugegraph.masterelection.GlobalMasterInfo,
          org.apache.hugegraph.util.DateUtil,
          org.apache.hugegraph.traversal.algorithm.CollectionPathsTraverser,
          org.apache.hugegraph.traversal.algorithm.CountTraverser,
          org.apache.hugegraph.traversal.algorithm.CustomizedCrosspointsTraverser,
          org.apache.hugegraph.traversal.algorithm.CustomizePathsTraverser,
          org.apache.hugegraph.traversal.algorithm.FusiformSimilarityTraverser,
          org.apache.hugegraph.traversal.algorithm.HugeTraverser,
          org.apache.hugegraph.traversal.algorithm.JaccardSimilarTraverser,
          org.apache.hugegraph.traversal.algorithm.KneighborTraverser,
          org.apache.hugegraph.traversal.algorithm.KoutTraverser,
          org.apache.hugegraph.traversal.algorithm.MultiNodeShortestPathTraverser,
          org.apache.hugegraph.traversal.algorithm.NeighborRankTraverser,
          org.apache.hugegraph.traversal.algorithm.PathsTraverser,
          org.apache.hugegraph.traversal.algorithm.PersonalRankTraverser,
          org.apache.hugegraph.traversal.algorithm.SameNeighborTraverser,
          org.apache.hugegraph.traversal.algorithm.ShortestPathTraverser,
          org.apache.hugegraph.traversal.algorithm.SingleSourceShortestPathTraverser,
          org.apache.hugegraph.traversal.algorithm.SubGraphTraverser,
          org.apache.hugegraph.traversal.algorithm.TemplatePathsTraverser,
          org.apache.hugegraph.traversal.algorithm.steps.EdgeStep,
          org.apache.hugegraph.traversal.algorithm.steps.RepeatEdgeStep,
          org.apache.hugegraph.traversal.algorithm.steps.WeightedEdgeStep,
          org.apache.hugegraph.traversal.optimize.ConditionP,
          org.apache.hugegraph.traversal.optimize.Text,
          org.apache.hugegraph.traversal.optimize.TraversalUtil,
          org.opencypher.gremlin.traversal.CustomFunctions,
          org.opencypher.gremlin.traversal.CustomPredicate
        ],
        methodImports: [
          java.lang.Math#*,
          org.opencypher.gremlin.traversal.CustomPredicate#*,
          org.opencypher.gremlin.traversal.CustomFunctions#*
        ]
      },
      org.apache.tinkerpop.gremlin.jsr223.ScriptFileGremlinPlugin: {
        files: [scripts/empty-sample.groovy]
      }
    }
  }
}
serializers:
  - { className: org.apache.tinkerpop.gremlin.driver.ser.GraphBinaryMessageSerializerV1,
      config: {
        serializeResultToString: false,
        ioRegistries: [org.apache.hugegraph.io.HugeGraphIoRegistry]
      }
  }
  - { className: org.apache.tinkerpop.gremlin.driver.ser.GraphSONMessageSerializerV1d0,
      config: {
        serializeResultToString: false,
        ioRegistries: [org.apache.hugegraph.io.HugeGraphIoRegistry]
      }
  }
  - { className: org.apache.tinkerpop.gremlin.driver.ser.GraphSONMessageSerializerV2d0,
      config: {
        serializeResultToString: false,
        ioRegistries: [org.apache.hugegraph.io.HugeGraphIoRegistry]
      }
  }
  - { className: org.apache.tinkerpop.gremlin.driver.ser.GraphSONMessageSerializerV3d0,
      config: {
        serializeResultToString: false,
        ioRegistries: [org.apache.hugegraph.io.HugeGraphIoRegistry]
      }
  }
metrics: {
  consoleReporter: {enabled: false, interval: 180000},
  csvReporter: {enabled: false, interval: 180000, fileName: ./metrics/gremlin-server-metrics.csv},
  jmxReporter: {enabled: false},
  slf4jReporter: {enabled: false, interval: 180000},
  gangliaReporter: {enabled: false, interval: 180000, addressingMode: MULTICAST},
  graphiteReporter: {enabled: false, interval: 180000}
}
maxInitialLineLength: 4096
maxHeaderSize: 8192
maxChunkSize: 8192
maxContentLength: 65536
maxAccumulationBufferComponents: 1024
resultIterationBatchSize: 64
writeBufferLowWaterMark: 32768
writeBufferHighWaterMark: 65536
ssl: {
  enabled: false
}

通常只需关注 channelizer、host 和 port。图不在 Gremlin Server 的 graphs 段加载;是否读取本地图配置由 rest-server.properties 中的 graph.load_from_local_config 控制。

  • channelizer:默认的 WsAndHttpChannelizer 同时支持 WebSocket 和 HTTP。Gremlin-Console 使用 WebSocket,HugeGraph-Client、Loader 和 Hubble 使用 HTTP;

默认 GremlinServer 是服务在 127.0.0.1:8182,如果需要修改,配置 host、port 即可

  • host:部署 GremlinServer 机器的机器名或 IP,GremlinServer 不直接暴露给用户,由 RestServer 转发 Gremlin 请求;
  • port:部署 GremlinServer 机器的端口;

同时需要在 rest-server.properties 中增加对应的配置项 gremlinserver.url=http://host:port

3 rest-server.properties

下面是可用的 rest-server.properties 示例。当前上游发布模板没有写出 graph.load_from_local_config,而源码默认值为 false;使用 conf/graphs 中的本地图配置时必须显式设为 true。

# bind url
# could use '0.0.0.0' or specified (real)IP to expose external network access
restserver.url=http://127.0.0.1:8080
#restserver.enable_graphspaces_filter=false
# gremlin server url, need to be consistent with host and port in gremlin-server.yaml
#gremlinserver.url=127.0.0.1:8182

graphs=./conf/graphs
graph.load_from_local_config=true

# The maximum thread ratio for batch writing, only take effect if the batch.max_write_threads is 0
batch.max_write_ratio=80
batch.max_write_threads=0

# configuration of arthas
arthas.telnetPort=8562
arthas.httpPort=8561
arthas.ip=127.0.0.1
arthas.disabledCommands=jad

# authentication configs
#auth.authenticator=org.apache.hugegraph.auth.StandardAuthenticator
# for admin password, By default, it is pa and takes effect upon the first startup
#auth.admin_pa=pa
#auth.graph_store=hugegraph

# use pd
# usePD=true

# slow query log
log.slow_query_threshold=1000
# bytes of request body recorded as-is (may contain sensitive literals), 0 to disable
log.slow_query_body_limit=512

# jvm(in-heap) memory usage monitor, set 1 to disable it
memory_monitor.threshold=0.85
memory_monitor.period=2000
  • restserver.url:RestServer 提供服务的 url,根据实际环境修改。如果其他 IP 地址无法访问,可以尝试修改为特定的地址;或修改为 http://0.0.0.0 来监听来自任何 IP 地址的请求,这种方案较为便捷,但需要留意服务可被访问的网络范围;
  • graphs:图配置文件所在目录,默认值是 ./conf/graphs。init-store 会扫描该目录;Server 仅在 graph.load_from_local_config=true 时加载其中的 properties 文件;
  • graph.load_from_local_config:是否在 Server 启动时读取本地图配置,源码默认值为 false;

当前上游模板中的 Arthas 键仍写作 arthas.telnet_port、arthas.http_port 和 arthas.disabled_commands,但 ServerOptions 读取的是下方示例中的 camelCase 名称。自定义配置应使用 arthas.telnetPort、arthas.httpPort 和 arthas.disabledCommands。

配置项 gremlinserver.url 是 GremlinServer 为 RestServer 提供服务的 url,该配置项默认为 http://127.0.0.1:8182,如需修改,需要和 gremlin-server.yaml 中的 host 和 port 相匹配;该值可以像模板那样省略协议前缀,缺失时会自动补上 http://。

4 hugegraph.properties

hugegraph.properties 是一类文件,因为如果系统存在多个图,则会有多个相似的文件。该文件用来配置与图存储和查询相关的参数,文件的默认内容如下:

# gremlin entrance to create graph
# auth config: org.apache.hugegraph.auth.HugeFactoryAuthProxy
gremlin.graph=org.apache.hugegraph.HugeFactory

# cache config
#schema.cache_capacity=100000
# vertex-cache default is 1000w, 10min expired
vertex.cache_type=l2
#vertex.cache_capacity=10000000
#vertex.cache_expire=600
# edge-cache default is 100w, 10min expired
edge.cache_type=l2
#edge.cache_capacity=1000000
#edge.cache_expire=600


# schema illegal name template
#schema.illegal_name_regex=\s+|~.*

#vertex.default_label=vertex

# NOTE: since 1.7.0, only hstore, rocksdb, hbase, memory are supported for backend.
# if you want to use Cassandra/MySql/PG... as backend, please use version < 1.7.0
backend=rocksdb
serializer=binary
# The process-wide max capacity of one serialization buffer in bytes
#serializer.buffer_max_capacity=134217728

store=hugegraph

# pd config
#pd.peers=127.0.0.1:8686

# task config
task.schedule_period=10
task.retry=0
task.wait_timeout=10

# search config
search.text_analyzer=jieba
search.text_analyzer_mode=INDEX

# rocksdb backend config
#rocksdb.data_path=/path/to/disk
#rocksdb.wal_path=/path/to/disk

# hbase backend config
#hbase.hosts=localhost
#hbase.port=2181
#hbase.znode_parent=/hbase
#hbase.threads_max=64
# IMPORTANT: recommend to modify the HBase partition number
#            by the actual/env data amount & RS amount before init store
#            It will influence the load speed a lot
#hbase.enable_partition=true
#hbase.vertex_partitions=10
#hbase.edge_partitions=30

# WARNING: These raft configurations are deprecated, please use the latest version instead.
# raft.mode=false

# memory management config
#memory.mode=off-heap
#memory.max_capacity=1073741824
#memory.one_query_max_capacity=104857600
#memory.alignment=8

重点关注未注释的几项:

  • gremlin.graph:GremlinServer 的启动入口,用户不要修改此项;开启鉴权时才改为 org.apache.hugegraph.auth.HugeFactoryAuthProxy;
  • vertex.cache_type / edge.cache_type:缓存实现,可选值为 l1 和 l2,默认 l2;
  • backend:使用的后端存储。1.7.0 支持 memory、rocksdb、hstore 和 hbase;
  • serializer:schema、vertex 和 edge 写入后端时使用的序列化器。RocksDB 使用 binary;
  • store:图在后端使用的存储名称;
  • task.schedule_period、task.retry、task.wait_timeout:异步任务的调度周期(秒)、重试次数和等待超时(秒)。调度器由后端决定,hstore 使用分布式调度器,其余后端使用本地调度器;旧的 task.scheduler_type 键已被忽略;
  • search.text_analyzer / search.text_analyzer_mode:全文索引使用的分词器及其模式。可选分词器为 ansj、hanlp、smartcn、jieba、jcseg、mmseg4j 和 ikanalyzer,每种分词器有各自的模式取值;
  • rocksdb.data_path:backend 为 rocksdb 时此项才有意义,rocksdb 的数据目录,默认为 rocksdb-data/data
  • rocksdb.wal_path:backend 为 rocksdb 时此项才有意义,rocksdb 的日志目录,默认为 rocksdb-data/wal

5 多图配置

一个 Server 可以加载多个图,每个图使用单独的 properties 文件。下面创建 RocksDB 图 hugegraph_rocksdb 和内存图 hugegraph_memory。

[可选]:修改 rest-server.properties

通过修改 rest-server.properties 中的 graphs 配置项来设置图的配置文件目录。默认配置为 graphs=./conf/graphs,如果想要修改为其它目录则调整 graphs 配置项,比如调整为 graphs=/etc/hugegraph/graphs,示例如下:

graphs=./conf/graphs
graph.load_from_local_config=true

在 conf/graphs 路径下基于 hugegraph.properties 创建 hugegraph_memory.properties 和 hugegraph_rocksdb.properties。

hugegraph_memory.properties 修改如下:

backend=memory
serializer=text
store=hugegraph_memory

hugegraph_rocksdb.properties 修改如下:

backend=rocksdb
serializer=binary

store=hugegraph_rocksdb

停止 Server,初始化执行 init-store.sh(为新的图创建数据库),重新启动 Server

$ ./bin/stop-hugegraph.sh
$ ./bin/init-store.sh

Initializing HugeGraph Store...
2023-06-11 14:16:14 [main] [INFO] o.a.h.u.ConfigUtil - Scanning option 'graphs' directory './conf/graphs'
2023-06-11 14:16:14 [main] [INFO] o.a.h.c.InitStore - Init graph with config file: ./conf/graphs/hugegraph_rocksdb.properties
...
2023-06-11 14:16:15 [main] [INFO] o.a.h.StandardHugeGraph - Graph 'hugegraph_rocksdb' has been initialized
2023-06-11 14:16:15 [main] [INFO] o.a.h.c.InitStore - Init graph with config file: ./conf/graphs/hugegraph_memory.properties
...
2023-06-11 14:16:16 [main] [INFO] o.a.h.StandardHugeGraph - Graph 'hugegraph_memory' has been initialized
2023-06-11 14:16:16 [main] [INFO] o.a.h.StandardHugeGraph - Close graph standardhugegraph[hugegraph_rocksdb]
...
2023-06-11 14:16:16 [main] [INFO] o.a.h.HugeFactory - HugeFactory shutdown
2023-06-11 14:16:16 [hugegraph-shutdown] [INFO] o.a.h.HugeFactory - HugeGraph is shutting down
Initialization finished.
$ ./bin/start-hugegraph.sh

Starting HugeGraphServer in daemon mode...
Connecting to HugeGraphServer (http://127.0.0.1:8080/graphs)...OK
Started [pid 21614]

查看创建的图:

curl http://127.0.0.1:8080/graphspaces/DEFAULT/graphs

{"graphs":["hugegraph_rocksdb","hugegraph_memory"]}

查看某个图的信息:

curl http://127.0.0.1:8080/graphspaces/DEFAULT/graphs/hugegraph_memory

{"name":"hugegraph_memory","backend":"memory"}
curl http://127.0.0.1:8080/graphspaces/DEFAULT/graphs/hugegraph_rocksdb

{"name":"hugegraph_rocksdb","backend":"rocksdb"}

2 - Server 完整配置手册

Gremlin Server 配置项

对应配置文件gremlin-server.yaml

config optiondefault valuedescription
host127.0.0.1The host or ip of Gremlin Server.
port8182The listening port of Gremlin Server.
graphs{}图由 Server 动态加载,不要在此处配置。
evaluationTimeout30000Gremlin 脚本执行超时,单位为毫秒。
channelizerorg.apache.tinkerpop.gremlin.server.channel.WsAndHttpChannelizer同时处理 WebSocket 和 HTTP 请求。
maxContentLength65536Server 接受的单个请求的最大字节数。
maxChunkSize8192HTTP 请求分块的最大字节数。
maxHeaderSize8192HTTP 请求头的最大字节数。
resultIterationBatchSize64流式返回结果集时,每批返回的结果条数。
ssl.enabledfalseGremlin Server 是否启用 TLS。
authentication未配置启用认证时配置认证器、处理器和 rest-server.properties 路径。

Rest Server & API 配置项

对应配置文件rest-server.properties

config optiondefault valuedescription
graphs./conf/graphs图配置 properties 文件所在目录。
graph.load_from_local_configfalse是否在 Server 启动时读取 graphs 目录;使用本地图配置时需设为 true。
graphs.enable_dynamic_create_droptrueWhether to enable create or drop graph dynamically.
init_store.enabledtrueWhether init-store initializes the local backend stores and the built-in admin account. Set false in distributed deployments (PD/HStore) where the storage side already owns the metadata.
server.id空字符串The optional legacy id of hugegraph-server.
server.rolemasterThe role of nodes in the cluster, available types are [master, worker, computer]
server.role_electionfalseWhether to enable role election, if enabled, the server will elect a master node in the cluster.
server.node_idnode-id1The node id of the server.
server.node_roleworkerThe node role of the server.
server.graphspaceDEFAULTThe graph space of the server.
server.service_idDEFAULTThe service id of the server.
server.path_graphspaceDEFAULTThe default path graph space of the server.
server.start_ignore_single_graph_errortrueWhether to start ignore single graph error.
server.event_hub_threads1The event hub threads of server.
restserver.urlhttp://127.0.0.1:8080The url for listening of graph server.
ssl.keystore_fileconf/hugegraph-server.keystoreThe path of server keystore file used when https protocol is enabled.
ssl.keystore_passwordhugegraphThe password of the server keystore file when the https protocol is enabled.
white_ip.statusdisableThe status of whether enable white ip.
restserver.max_worker_threads2 * CPUsThe maximum worker threads of rest server.
restserver.task_threadsmax(4, CPUs / 2)The task threads of rest server.
restserver.min_free_memory64The minimum free memory(MB) of rest server, requests will be rejected when the available memory of system is lower than this value.
restserver.request_timeout30The time in seconds within which a request must complete, -1 means no timeout.
restserver.connection_idle_timeout30The time in seconds to keep an inactive connection alive, -1 means no timeout.
restserver.connection_max_requests256The max number of HTTP requests allowed to be processed on one keep-alive connection, -1 means unlimited.
gremlinserver.urlhttp://127.0.0.1:8182The url of gremlin server.
gremlinserver.max_route2 * CPUsThe max route number for gremlin server.
gremlinserver.timeout30The timeout in seconds of waiting for gremlin server.
batch.max_edges_per_batch2500The maximum number of edges submitted per batch.
batch.max_vertices_per_batch2500The maximum number of vertices submitted per batch.
batch.max_write_ratio70The maximum thread ratio for batch writing, only take effect if the batch.max_write_threads is 0.
batch.max_write_threads0The maximum threads for batch writing, if the value is 0, the actual value will be set to batch.max_write_ratio * restserver.max_worker_threads.
raft.group_peers127.0.0.1:8090The rpc address of raft group initial peers.
auth.authenticatorThe class path of authenticator implementation. e.g., org.apache.hugegraph.auth.StandardAuthenticator, or a custom implementation.
auth.graph_storehugegraphThe name of graph used to store authentication information, like users, only for org.apache.hugegraph.auth.StandardAuthenticator.
auth.admin_papa内置 admin 账户的初始密码,仅首次启动时生效;部署前必须修改。
auth.audit_log_rate1000.0The max rate of audit log output per user, default value is 1000 records per second.
auth.cache_capacity10240The max cache capacity of each auth cache item.
auth.cache_expire600The expiration time in seconds of auth cache in auth client and auth server.
auth.remote_urlIf the address is empty, it provide auth service, otherwise it is auth client and also provide auth service through rpc forwarding. The remote url can be set to multiple addresses, which are concat by ‘,’.
auth.token_expire86400The expiration time in seconds after token created
auth.token_secret启动时随机生成HS256 的密钥;需要跨重启保持既有 token 有效时应显式配置。
exception.allow_tracetrueWhether to allow exception trace stack.
memory_monitor.threshold0.85Threshold for JVM memory usage monitoring, 1 means disabling the memory monitoring task.
memory_monitor.period2000The period in ms of JVM memory usage monitoring, in each period we will detect the jvm memory usage and take corresponding actions.
log.slow_query_threshold1000The threshold time(ms) of logging slow query, 0 means logging slow query is disabled.
log.slow_query_body_limit512慢查询日志记录的请求体最大字节数,0 表示不记录。记录的前缀原样写入,可能包含敏感的 Gremlin 或 Cypher 字面量。
角色选举配置项 (可选)

对应配置文件rest-server.properties,仅在 server.role_election=true 时生效。

config optiondefault valuedescription
server.role.node_external_urlhttp://127.0.0.1:8080The url of external accessibility.
server.role.base_timeout500The role state machine candidate state base timeout time, in ms.
server.role.random_timeout1000The random timeout in ms that be used when candidate node request to become master state to reduce competitive voting.
server.role.heartbeat_interval2The role state machine heartbeat interval second time.
server.role.fail_count5When the node failed count of update or query heartbeat is reaches this threshold, the node will become abdication state to guardsafe property.
server.role.master_dead_times10When the worker node detects that the number of times the master node fails to update heartbeat reaches this threshold, the worker node will become to a candidate node.

PD/Meta 配置项 (分布式模式)

对应配置文件rest-server.properties

config optiondefault valuedescription
usePDfalseWhether use pd.
pd.peers127.0.0.1:8686The pd server peers, separated with commas.
clusterhg-testThe cluster name.
metrics.data_to_pdtrueWhether to report metrics data to pd.
meta.endpointshttp://127.0.0.1:2379meta 端点的 URL。当前代码中没有任何地方读取该配置项,设置后不会生效;meta 连接由 pd.peers 建立。
meta.use_cafalseWhether to use ca to meta server.
meta.caThe ca file of meta server.
meta.client_caThe client ca file of meta server.
meta.client_keyThe client key file of meta server.

HStore 后端还会从图配置文件 {graph-name}.properties 中读取以下两项,默认值 0 表示由 PD 决定:

config optiondefault valuedescription
hstore.partition_count0Number of partitions, which PD controls partitions based on.
hstore.shard_count0Number of copies, which PD controls partition copies based on.

基本配置项

基本配置项及后端配置项对应配置文件:{graph-name}.properties,如hugegraph.properties

config optiondefault valuedescription
gremlin.graphorg.apache.hugegraph.HugeFactoryGremlin entrance to create graph.
backendmemoryThe data store type. For version 1.7.0+ the allowed values are [memory, rocksdb, hstore, hbase]; the shipped conf/graphs/hugegraph.properties sets rocksdb and conf/graphs/hstore.properties.template sets hstore. Note: cassandra, scylladb, mysql, postgresql were removed in 1.7.0 (use <= 1.5.x for legacy backends).
serializertextThe serializer for backend store, built-in values are [text, binary, binaryscatter]; a backend may register its own, like hbase. The shipped graph templates set binary.
serializer.buffer_max_capacity134217728The process-wide max capacity of one serialization buffer in bytes.
storehugegraphThe backend database namespace.
store.connection_detect_interval600The interval in seconds for detecting connections, if the idle time of a connection exceeds this value, detect it and reconnect if needed before using, value 0 means detecting every time.
store.graphgThe graph table name, which store vertex, edge and property.
graphspaceDEFAULTThe graph space name.
alias.graph.idThe graph alias id.
graph.read_modeOLTP_ONLYThe graph read mode, which could be ALL | OLTP_ONLY | OLAP_ONLY.
pd.peers127.0.0.1:8686The addresses of pd nodes, separated with commas. Only used by the hstore backend.
schema.illegal_name_regex.\s+$|~.The regex specified the illegal format for schema name.
schema.cache_capacity10000The max cache size(items) of schema cache.
schema.init_templateThe template schema used to init graph.
schema.index_rebuild_using_pushdowntrueWhether to use pushdown when to create/rebuild index.
vertex.cache_typel2The type of vertex cache, allowed values are [l1, l2].
vertex.cache_capacity10000000The max cache size(items) of vertex cache.
vertex.cache_expire600The expiration time in seconds of vertex cache.
vertex.check_customized_id_existfalseWhether to check the vertices exist for those using customized id strategy.
vertex.default_labelvertexThe default vertex label.
vertex.tx_capacity10000The max size(items) of vertices(uncommitted) in transaction.
vertex.check_adjacent_vertex_existfalseWhether to check the adjacent vertices of edges exist.
vertex.lazy_load_adjacent_vertextrueWhether to lazy load adjacent vertices of edges.
vertex.part_edge_commit_size5000Whether to enable the mode to commit part of edges of vertex, enabled if commit size > 0, 0 means disabled.
vertex.encode_primary_key_numbertrueWhether to encode number value of primary key in vertex id.
vertex.remove_left_index_at_overwritefalseWhether remove left index at overwrite.
edge.cache_typel2The type of edge cache, allowed values are [l1, l2].
edge.cache_capacity1000000The max cache size(items) of edge cache.
edge.cache_expire600The expiration time in seconds of edge cache.
edge.tx_capacity10000The max size(items) of edges(uncommitted) in transaction.
query.page_size500The size of each page when querying by paging.
query.batch_size1000The size of each batch when querying by batch.
query.ignore_invalid_datatrueWhether to ignore invalid data of vertex or edge.
query.index_intersect_threshold1000The maximum number of intermediate results to intersect indexes when querying by multiple single index properties.
query.max_indexes_available1The upper limit of the number of indexes that can be used to query.
query.dedup_optionlimitThe way to dedup data, allowed values are [limit, global].
query.trust_indexfalseWhether to trust index.
query.ramtable_edges_capacity20000000The maximum number of edges in ramtable, include OUT and IN edges.
query.ramtable_enablefalseWhether to enable ramtable for query of adjacent edges.
query.ramtable_vertices_capacity10000000The maximum number of vertices in ramtable, generally the largest vertex id is used as capacity.
query.optimize_aggregate_by_indexfalseWhether to optimize aggregate query(like count) by index.
oltp.concurrent_depth10The min depth to enable concurrent oltp algorithm.
oltp.concurrent_threadsmax(10, CPUs / 2)Thread number to concurrently execute oltp algorithm.
oltp.collection_typeECThe implementation type of collections used in oltp algorithm, allowed values are [JCF, EC, FU].
oltp.query_batch_size10000The size of each batch when executing oltp algorithm.
oltp.query_batch_avg_degree_ratio0.95The ratio of exponential approximation for average degree of iterator when executing oltp algorithm.
oltp.query_batch_expect_degree100000000The expect sum of degree in each batch when executing oltp algorithm.
rate_limit.read0The max rate(times/s) to execute query of vertices/edges.
rate_limit.write0The max rate(items/s) to add/update/delete vertices/edges.
task.schedule_period10Period time in seconds when scheduler to schedule task.
task.wait_timeout10Timeout in seconds for waiting for the task to complete, such as when truncating or clearing the backend.
task.retry0Task retry times, allowed range is [0, 3].
task.input_size_limit16777216The job input size limit in bytes.
task.result_size_limit16777216The job result size limit in bytes.
task.sync_deletionfalseWhether to delete schema or expired data synchronously.
task.ttl_delete_batch1The batch size used to delete expired data.
computer.config./conf/computer.yamlThe config file path of computer job.
k8s.operator_template./conf/operator-template.yamlThe path of operator container template.
k8s.quota_template./conf/resource-quota-template.yamlThe path of resource quota template.
search.text_analyzerikanalyzerChoose a text analyzer for searching the vertex/edge properties, available type are [ansj, hanlp, smartcn, jieba, jcseg, mmseg4j, ikanalyzer]. The shipped graph templates set jieba. If use ‘ikanalyzer’, need download jar from ‘https://github.com/apache/hugegraph-doc/raw/ik_binary/dist/server/ikanalyzer-2012_u6.jar' to lib directory
search.text_analyzer_modesmartSpecify the mode for the text analyzer, the available mode of analyzer are {ansj: [BaseAnalysis, IndexAnalysis, ToAnalysis, NlpAnalysis], hanlp: [standard, nlp, index, nShort, shortest, speed], smartcn: [], jieba: [SEARCH, INDEX], jcseg: [Simple, Complex], mmseg4j: [Simple, Complex, MaxWord], ikanalyzer: [smart, max_word]}.
snowflake.datacenter_id0The datacenter id of snowflake id generator.
snowflake.force_stringfalseWhether to force the snowflake long id to be a string.
snowflake.worker_id0The worker id of snowflake id generator.
memory.modeoff-heapThe memory mode used for query in HugeGraph.
memory.max_capacity1073741824The maximum memory capacity in bytes that can be managed for all queries in HugeGraph.
memory.one_query_max_capacity104857600The maximum memory capacity in bytes that can be managed for a query in HugeGraph.
memory.alignment8The alignment used for round memory size.
Raft 配置项 (已废弃)

发行包中的图配置模板已将这些配置项标注为废弃。它们仅在 raft.mode=true 时生效, 且 raft.group_peers 从 rest-server.properties 读取,而不是图配置文件。

config optiondefault valuedescription
raft.modefalseWhether the backend storage works in raft mode.
raft.safe_readfalseWhether to use linearly consistent read.
raft.path./raftlogThe log path of current raft node.
raft.use_replicator_pipelinetrueWhether to use replicator line, when turned on it multiple logs can be sent in parallel, and the next log doesn’t have to wait for the ack message of the current log to be sent.
raft.election_timeout10000Timeout in milliseconds to launch a round of election.
raft.snapshot_interval3600The interval in seconds to trigger snapshot save.
raft.snapshot_threads4The thread number used to do snapshot.
raft.snapshot_parallel_compressfalseWhether to enable parallel compress.
raft.snapshot_compress_threads4The thread number used to do snapshot compress.
raft.snapshot_decompress_threads4The thread number used to do snapshot decompress.
raft.backend_threadsCPUsThe thread number used to apply task to backend.
raft.read_index_threads8The thread number used to execute reading index.
raft.read_strategyReadOnlyLeaseBasedThe linearizability of read strategy, allowed values are [ReadOnlyLeaseBased, ReadOnlySafe].
raft.apply_batch1The apply batch size to trigger disruptor event handler.
raft.queue_size16384The disruptor buffers size for jraft RaftNode, StateMachine and LogManager.
raft.queue_publish_timeout60The timeout in second when publish event into disruptor.
raft.rpc_threadsmax(CPUs * 2, 80)The rpc threads for jraft RPC layer.
raft.rpc_connect_timeout5000The rpc connect timeout in milliseconds for jraft rpc.
raft.rpc_timeout60The general rpc timeout in seconds for jraft rpc.
raft.install_snapshot_rpc_timeout36000The install snapshot rpc timeout in seconds for jraft rpc.
raft.rpc_buf_low_water_mark10485760The ChannelOutboundBuffer’s low water mark of netty, when buffer size less than this size, the method ChannelOutboundBuffer.isWritable() will return true, it means that low downstream pressure or good network.
raft.rpc_buf_high_water_mark20971520The ChannelOutboundBuffer’s high water mark of netty, only when buffer size exceed this size, the method ChannelOutboundBuffer.isWritable() will return false, it means that the downstream pressure is too great to process the request or network is very congestion, upstream needs to limit rate at this time.

RocksDB 后端配置项

config optiondefault valuedescription
backendMust be set to rocksdb.
serializerMust be set to binary.
rocksdb.data_pathrocksdb-data/dataThe path for storing data of RocksDB.
rocksdb.wal_pathrocksdb-data/walThe path for storing WAL of RocksDB.
rocksdb.sst_pathThe path for ingesting SST file into RocksDB.
rocksdb.data_disks[]The optimized disks for storing data of RocksDB. The format of each element: STORE/TABLE: /path/disk.Allowed keys are [g/vertex, g/edge_out, g/edge_in, g/vertex_label_index, g/edge_label_index, g/range_int_index, g/range_float_index, g/range_long_index, g/range_double_index, g/secondary_index, g/search_index, g/shard_index, g/unique_index, g/olap]
rocksdb.log_levelINFOThe info log level of RocksDB.
rocksdb.num_levels7Set the number of levels for this database.
rocksdb.compaction_styleLEVELSet compaction style for RocksDB: LEVEL/UNIVERSAL/FIFO.
rocksdb.optimize_modetrueOptimize for heavy workloads and big datasets.
rocksdb.bulkload_modefalseSwitch to the mode to bulk load data into RocksDB.
rocksdb.compression_per_level[none, none, snappy, snappy, snappy, snappy, snappy]The compression algorithms for different levels of RocksDB, allowed values are none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd.
rocksdb.bottommost_compressionnoneThe compression algorithm for the bottommost level of RocksDB, allowed values are none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd.
rocksdb.compressionsnappyThe compression algorithm for compressing blocks of RocksDB, allowed values are none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd.
rocksdb.max_background_jobs8Maximum number of concurrent background jobs, including flushes and compactions.
rocksdb.max_subcompactions4The value represents the maximum number of threads per compaction job.
rocksdb.delayed_write_rate16777216The rate limit in bytes/s of user write requests when need to slow down if the compaction gets behind.
rocksdb.max_open_files-1The maximum number of open files that can be cached by RocksDB, -1 means no limit.
rocksdb.max_manifest_file_size104857600The max size of manifest file in bytes.
rocksdb.skip_stats_update_on_db_openfalseWhether to skip statistics update when opening the database, setting this flag true allows us to not update statistics.
rocksdb.skip_check_sst_size_on_db_openfalseWhether to skip checking sizes of all sst files when opening the database.
rocksdb.max_file_opening_threads16The max number of threads used to open files.
rocksdb.max_total_wal_size0Total size of WAL files in bytes. Once WALs exceed this size, we will start forcing the flush of column families related, 0 means no limit.
rocksdb.bytes_per_sync0Allows OS to incrementally sync SST files to disk while they are being written, asynchronously in the background. Issue one request for every bytes_per_sync written. 0 turns it off.
rocksdb.wal_bytes_per_sync0Allows OS to incrementally sync WAL files to disk while they are being written, asynchronously in the background. Issue one request for every bytes_per_sync written. 0 turns it off.
rocksdb.strict_bytes_per_syncfalseWhen true, guarantees SST/WAL files have at most bytes_per_sync/wal_bytes_per_sync bytes submitted for writeback at any given time. This can be used to handle cases where processing speed exceeds I/O speed.
rocksdb.db_write_buffer_size0Total size of write buffers in bytes across all column families, 0 means no limit.
rocksdb.log_readahead_size0The number of bytes to prefetch when reading the log. 0 means the prefetching is disabled.
rocksdb.compaction_readahead_size0The number of bytes to perform bigger reads when doing compaction. If running RocksDB on spinning disks, you should set this to at least 2MB. 0 means the prefetching is disabled.
rocksdb.row_cache_capacity0The capacity in bytes of global cache for table-level rows. 0 means the row_cache is disabled.
rocksdb.delete_obsolete_files_period21600The periodicity in seconds when obsolete files get deleted, 0 means always do full purge.
rocksdb.write_buffer_size134217728Amount of data in bytes to build up in memory.
rocksdb.max_write_buffer_number6The maximum number of write buffers that are built up in memory.
rocksdb.min_write_buffer_number_to_merge2The minimum number of write buffers that will be merged together.
rocksdb.max_write_buffer_number_to_maintain0The total maximum number of write buffers to maintain in memory for conflict checking when transactions are used.
rocksdb.memtable_bloom_size_ratio0.0If prefix-extractor is set and memtable_bloom_size_ratio is not 0, or if memtable_whole_key_filtering is set true, create bloom filter for memtable with the size of write_buffer_size * memtable_bloom_size_ratio. If it is larger than 0.25, it is santinized to 0.25.
rocksdb.memtable_whole_key_filteringfalseEnable whole key bloom filter in memtable, it can potentially reduce CPU usage for point-look-ups. Note this will only take effect if memtable_bloom_size_ratio > 0.
rocksdb.memtable_huge_page_size0The page size for huge page TLB for bloom in memtable. If <= 0, not allocate from huge page TLB but from malloc.
rocksdb.inplace_update_supportfalseAllows thread-safe inplace updates if a put key exists in current memtable and sizeof new value is smaller.
rocksdb.level_compaction_dynamic_level_bytesfalseWhether to enable level_compaction_dynamic_level_bytes, if it’s enabled we give max_bytes_for_level_multiplier a priority against max_bytes_for_level_base, the bytes of base level is dynamic for a more predictable LSM tree, it is useful to limit worse case space amplification. Turning this feature on/off for an existing DB can cause unexpected LSM tree structure so it’s not recommended.
rocksdb.max_bytes_for_level_base536870912The upper-bound of the total size of level-1 files in bytes.
rocksdb.max_bytes_for_level_multiplier10.0The ratio between the total size of level (L+1) files and the total size of level L files for all L.
rocksdb.target_file_size_base67108864The target file size for compaction in bytes.
rocksdb.target_file_size_multiplier1The size ratio between a level L file and a level (L+1) file.
rocksdb.level0_file_num_compaction_trigger2Number of files to trigger level-0 compaction.
rocksdb.level0_slowdown_writes_trigger20Soft limit on number of level-0 files for slowing down writes.
rocksdb.level0_stop_writes_trigger36Hard limit on number of level-0 files for stopping writes.
rocksdb.soft_pending_compaction_bytes_limit68719476736The soft limit to impose on pending compaction in bytes.
rocksdb.hard_pending_compaction_bytes_limit274877906944The hard limit to impose on pending compaction in bytes.
rocksdb.allow_mmap_writesfalseAllow the OS to mmap file for writing.
rocksdb.allow_mmap_readsfalseAllow the OS to mmap file for reading sst tables.
rocksdb.use_direct_readsfalseEnable the OS to use direct I/O for reading sst tables.
rocksdb.use_direct_io_for_flush_and_compactionfalseEnable the OS to use direct read/writes in flush and compaction.
rocksdb.use_fsyncfalseIf true, then every store to stable storage will issue a fsync.
rocksdb.atomic_flushfalseIf true, flushing multiple column families and committing their results atomically to MANIFEST. Note that it’s not necessary to set atomic_flush=true if WAL is always enabled.
rocksdb.format_version5The format version of BlockBasedTable, allowed values are 0~5.
rocksdb.index_typekBinarySearchThe index type used to lookup between data blocks with the sst table, allowed values are [kBinarySearch,kHashSearch,kTwoLevelIndexSearch,kBinarySearchWithFirstKey].
rocksdb.data_block_index_typekDataBlockBinarySearchThe search type used to point lookup in data block with the sst table, allowed values are [kDataBlockBinarySearch,kDataBlockBinaryAndHash].
rocksdb.data_block_hash_table_util_ratio0.75The hash table utilization ratio value of entries/buckets. It is valid only when data_block_index_type=kDataBlockBinaryAndHash.
rocksdb.block_size4096Approximate size of user data packed per block, Note that it corresponds to uncompressed data.
rocksdb.block_size_deviation10The percentage of free space used to close a block.
rocksdb.block_restart_interval16The block restart interval for delta encoding in blocks.
rocksdb.block_cache_capacity8388608The amount of block cache in bytes that will be used by RocksDB, 0 means no block cache.
rocksdb.cache_index_and_filter_blockstrueSet this option true if we’d put index/filter blocks to the block cache.
rocksdb.pin_l0_filter_and_index_blocks_in_cachetrueSet this option true if we’d pin L0 index/filter blocks to the block cache.
rocksdb.bloom_filter_bits_per_key-1The bits per key in bloom filter, a good value is 10, which yields a filter with ~ 1% false positive rate. Set bloom_filter_bits_per_key > 0 to enable bloom filter, -1 means no bloom filter (0~0.5 round down to no filter).
rocksdb.bloom_filter_block_based_modefalseIf bloom filter is enabled, set this option true to use block based filter rather than full filter.
rocksdb.bloom_filter_whole_key_filteringtrueIf bloom filter is enabled, set this option true to place whole keys in the bloom filter, else place the prefix of keys when prefix-extractor is set.
rocksdb.optimize_filters_for_hitstrueIf bloom filter is enabled, this flag allows us to not store filters for the last level. set this option true to optimize the filters mainly for cases where keys are found rather than also optimize for keys missed.
rocksdb.partition_filters_and_indexesfalseIf bloom filter is enabled, set this option true to use partitioned full filters and indexes for each sst file. This option is incompatible with block-based filters.
rocksdb.pin_top_level_index_and_filtertrueIf partition_filters_and_indexes is set true, set this option true if we’d pin top-level index of partitioned filter and index blocks to the block cache.
rocksdb.prefix_extractor_n_bytes0The prefix-extractor uses the first N bytes of a key as its prefix, it will use the full key when a key is shorter than the N. 0 means unset prefix-extractor.
K8s 配置项 (可选)

对应配置文件rest-server.properties

config optiondefault valuedescription
server.use_k8sfalseWhether to use k8s to support multiple tenancy.
server.deploy_in_k8sfalseWhether to deploy server in k8s.
server.urls_to_pdhttp://0.0.0.0:8080Used as the server address reserved for PD and provided to clients, only used when starting the server in k8s.
server.k8s_urlhttps://127.0.0.1:8888The url of k8s.
server.k8s_use_cafalseWhether to use ca to k8s api server.
server.k8s_caThe ca file of k8s api server.
server.k8s_client_caThe client ca file of k8s api server.
server.k8s_client_keyThe client key file of k8s api server.
k8s.apifalseThe k8s api start status when the computer service is enabled.
k8s.namespacehugegraph-computer-systemThe namespace used for k8s work when the computer service is enabled.
k8s.kubeconfigThe k8s kube config file when the computer service is enabled.
k8s.hugegraph_urlThe hugegraph url for k8s work when the computer service is enabled.
k8s.enable_internal_algorithmtrueWhether to open k8s internal algorithm.
service.access_pd_namehgService name for server to access pd service.
service.access_pd_tokenService token for server to access pd service.
server.k8s_oltp_image127.0.0.1/kgs_bd/hugegraphserver:3.0.0The oltp server image of k8s.
server.k8s_olap_imagehugegraph/hugegraph-server:v1The olap server image of k8s.
server.k8s_storage_imagehugegraph/hugegraph-server:v1The storage server image of k8s.
server.default_oltp_k8s_namespacehugegraph-serverThe default oltp namespace for HugeGraph default graph space.
server.default_olap_k8s_namespacehugegraph-computer-systemThe default olap namespace for HugeGraph default graph space.
k8s.internal_algorithm[page-rank, degree-centrality, wcc, triangle-count, rings, rings-with-filter, betweenness-centrality, closeness-centrality, lpa, links, kcore, louvain, clustering-coefficient, ppr, subgraph-match]The names of the built-in k8s algorithms.
k8s.algorithmsSee ServerOptions.K8S_ALGORITHMSThe name:paramsClass mapping of the built-in k8s algorithms.
Arthas 诊断配置项 (可选)

对应配置文件rest-server.properties

config optiondefault valuedescription
arthas.telnetPort8562Arthas telnet port.
arthas.httpPort8561Arthas HTTP port.
arthas.ip0.0.0.0Arthas bind IP.
arthas.disabledCommandsjadDisabled Arthas commands, separated by commas.
RPC Server 配置项

对应配置文件rest-server.properties

config optiondefault valuedescription
rpc.server_hostThe hosts/ips bound by rpc server to provide services, empty value means not enabled.
rpc.server_port8090The port bound by rpc server to provide services.
rpc.server_adaptive_portfalseWhether the bound port is adaptive, if it’s enabled, when the port is in use, automatically +1 to detect the next available port. Note that this process is not atomic, so there may still be port conflicts.
rpc.server_timeout30The timeout(in seconds) of rpc server execution.
rpc.remote_urlThe remote urls of rpc peers, it can be set to multiple addresses, which are concat by ‘,’, empty value means not enabled.
rpc.client_connect_timeout20The timeout(in seconds) of rpc client connect to rpc server.
rpc.client_reconnect_period10The period(in seconds) of rpc client reconnect to rpc server.
rpc.client_read_timeout40The timeout(in seconds) of rpc client read from rpc server.
rpc.client_retries3Failed retry number of rpc client calls to rpc server.
rpc.client_load_balancerconsistentHashThe rpc client uses a load-balancing algorithm to access multiple rpc servers in one cluster. Default value is ‘consistentHash’, means forwarding by request parameters.
rpc.protocolboltRpc communication protocol, client and server need to be specified the same value.
rpc.serializationhessian2Rpc serialization type, client and server must set the same value. Note: If you choose ‘protobuf’, you need to add the relative IDL file. (Could refer PD/Store *.proto)
rpc.config_order999Sofa-RPC configuration file loading order, the larger the more later loading.
rpc.logger_implcom.alipay.sofa.rpc.log.SLF4JLoggerImplSofa-RPC log implementation class.
HBase 后端配置项
config optiondefault valuedescription
backendMust be set to hbase.
serializerMust be set to hbase.
hbase.hostslocalhostThe hostnames or ip addresses of HBase zookeeper, separated with commas.
hbase.port2181The port address of HBase zookeeper.
hbase.threads_max64The max threads num of hbase connections.
hbase.znode_parent/hbaseThe znode parent path of HBase zookeeper.
hbase.zk_retry3The recovery retry times of HBase zookeeper.
hbase.truncate_timeout30The timeout in seconds of waiting for store truncate.
hbase.aggregation_timeout43200The timeout in seconds of waiting for aggregation.
hbase.kerberos_enablefalseIs Kerberos authentication enabled for HBase.
hbase.kerberos_keytabThe HBase’s key tab file for kerberos authentication.
hbase.kerberos_principalThe HBase’s principal for kerberos authentication.
hbase.krb5_conf/etc/krb5.confKerberos configuration file, including KDC IP, default realm, etc.
hbase.hbase_site/etc/hbase/conf/hbase-site.xmlThe HBase’s configuration file
hbase.enable_partitiontrueIs pre-split partitions enabled for HBase.
hbase.vertex_partitions10The number of partitions of the HBase vertex table.
hbase.edge_partitions30The number of partitions of the HBase edge table.

≤ 1.5 版本配置 (Legacy)

以下后端存储在 1.7.0+ 版本中不再支持,仅在 1.5.x 及更早版本中可用:

Cassandra 后端配置项
config optiondefault valuedescription
backendMust be set to cassandra.
serializerMust be set to cassandra.
cassandra.hostlocalhostThe seeds hostname or ip address of cassandra cluster.
cassandra.port9042The seeds port address of cassandra cluster.
cassandra.connect_timeout5The cassandra driver connect server timeout(seconds).
cassandra.read_timeout20The cassandra driver read from server timeout(seconds).
cassandra.keyspace.strategySimpleStrategyThe replication strategy of keyspace, valid value is SimpleStrategy or NetworkTopologyStrategy.
cassandra.keyspace.replication[3]The keyspace replication factor of SimpleStrategy, like ‘[3]’.Or replicas in each datacenter of NetworkTopologyStrategy, like ‘[dc1:2,dc2:1]’.
cassandra.usernameThe username to use to login to cassandra cluster.
cassandra.passwordThe password corresponding to cassandra.username.
cassandra.compression_typenoneThe compression algorithm of cassandra transport: none/snappy/lz4.
cassandra.jmx_port=71997199The port of JMX API service for cassandra.
cassandra.aggregation_timeout43200The timeout in seconds of waiting for aggregation.
ScyllaDB 后端配置项
config optiondefault valuedescription
backendMust be set to scylladb.
serializerMust be set to scylladb.

其它与 Cassandra 后端一致。

MySQL & PostgreSQL 后端配置项
config optiondefault valuedescription
backendMust be set to mysql.
serializerMust be set to mysql.
jdbc.drivercom.mysql.jdbc.DriverThe JDBC driver class to connect database.
jdbc.urljdbc:mysql://127.0.0.1:3306The url of database in JDBC format.
jdbc.usernamerootThe username to login database.
jdbc.password******The password corresponding to jdbc.username.
jdbc.ssl_modefalseThe SSL mode of connections with database.
jdbc.reconnect_interval3The interval(seconds) between reconnections when the database connection fails.
jdbc.reconnect_max_times3The reconnect times when the database connection fails.
jdbc.storage_engineInnoDBThe storage engine of backend store database, like InnoDB/MyISAM/RocksDB for MySQL.
jdbc.postgresql.connect_databasetemplate1The database used to connect when init store, drop store or check store exist.
PostgreSQL 后端配置项
config optiondefault valuedescription
backendMust be set to postgresql.
serializerMust be set to postgresql.

其它与 MySQL 后端一致。

PostgreSQL 后端的 driver 和 url 应该设置为:

  • jdbc.driver=org.postgresql.Driver
  • jdbc.url=jdbc:postgresql://localhost:5432/

3 - HugeGraph 内置用户权限与扩展权限配置及使用

概述

HugeGraph 内置 StandardAuthenticator,支持多用户认证和基于“用户、用户组、操作、资源”的权限控制。

StandardAuthenticator 模式的几个核心设计:

  • 初始化时创建超级管理员 (admin) 用户,后续通过超级管理员创建其它用户,新创建的用户被分配足够权限后,可以创建或管理更多的用户
  • 支持动态创建用户、用户组、资源,支持动态分配或取消权限
  • 用户可以属于一个或多个用户组,每个用户组可以拥有对任意个资源的操作权限,操作类型包括:读、写、删除、执行等种类
  • “资源” 描述了图数据库中的数据,比如符合某一类条件的顶点,每一个资源包括 type、label、properties三个要素,共有 18 种类型、任意 label、任意 properties 可组合形成的资源,一个资源的内部条件是且关系,多个资源之间的条件是或关系

举例说明:

// 场景:某用户只有北京地区的数据读取权限
user(name=xx) -belong-> group(name=xx) -access(read)-> target(graph=graph1, resource={label: person, city: Beijing})

配置用户认证

HugeGraph 目前默认未启用用户认证功能,需通过修改配置文件来启用该功能。

⚠️ SEC 提醒:图查询语言 (Gremlin/Cypher) 的安全性

鉴于图查询语言的灵活性可能带来的潜在系统安全隐患,不要把 Gremlin、Cypher 等查询接口直接暴露到公网。生产环境应同时启用鉴权、IP 白名单和审计日志,并通过 Docker 或 Kubernetes 隔离 Server 进程。

StandardAuthenticator 支持多用户认证和细粒度权限控制。也可以实现 HugeAuthenticator 接口来接入已有的用户系统。

用户认证使用 HTTP Basic Authentication。Basic 后面的值是 用户名:密码 的 Base64 编码。使用 curl 时可直接通过 -u 传入凭据:

curl -u 'admin:<password>' \
  http://localhost:8080/graphspaces/DEFAULT/graphs/hugegraph/schema/vertexlabels

警告:在 1.5.0 之前版本的 HugeGraph-Server 在鉴权模式下存在 JWT 相关的安全隐患,请务必使用新版本或自行修改 JWT token 的 secretKey。

修改方式为在配置文件rest-server.properties中重写auth.token_secret信息:(1.5.0 后会默认生成随机值则无需配置)

auth.token_secret=XXXX   #这里为 32 位 String,由 a-z,A-Z 和 0-9 组成

也可以通过下面的命令实现:

RANDOM_STRING=$(head /dev/urandom | tr -dc A-Za-z0-9 | head -c 32)
echo "auth.token_secret=${RANDOM_STRING}" >> rest-server.properties

由于默认值在每次启动时随机生成,当 token 需要在重启后继续有效、或者需要被多个服务节点接受时,必须显式配置该项。token 的有效期由 auth.token_expire 决定,默认为 86400 秒。

StandardAuthenticator 模式

StandardAuthenticator模式是通过在数据库后端存储用户信息来支持用户认证和权限控制,该实现基于数据库存储的用户的名称与密码进行认证(密码已被加密),基于用户的角色来细粒度控制用户权限。下面是具体的配置流程(重启服务生效):

在配置文件gremlin-server.yaml中配置authenticator及其rest-server文件路径:

authentication: {
  authenticator: org.apache.hugegraph.auth.StandardAuthenticator,
  authenticationHandler: org.apache.hugegraph.auth.WsAndHttpBasicAuthHandler,
  config: {tokens: conf/rest-server.properties}
}

在 rest-server.properties 中配置认证器和权限数据存储图:

auth.authenticator=org.apache.hugegraph.auth.StandardAuthenticator
auth.graph_store=hugegraph
# 内置 admin 账号的密码,默认为 pa,在首次启动时生效
#auth.admin_pa=<your-admin-password>

# auth client config
# 如果是分开部署 GraphServer 和 AuthServer,还需要指定下面的配置,地址填写 AuthServer 的 IP:RPC 端口
#auth.remote_url=127.0.0.1:8899,127.0.0.1:8898,127.0.0.1:8897

其中,graph_store配置项是指使用哪一个图来存储用户信息,如果存在多个图的话,选取任意一个均可。

在配置文件hugegraph{n}.properties中配置gremlin.graph信息:

gremlin.graph=org.apache.hugegraph.auth.HugeFactoryAuthProxy

权限 API 的调用方式见 Authentication API 文档。

自定义用户认证系统

如果需要支持更加灵活的用户系统,可自定义 authenticator 进行扩展,自定义 authenticator 实现接口org.apache.hugegraph.auth.HugeAuthenticator即可,然后修改配置文件中authenticator配置项指向该实现。

基于鉴权模式启动

首次执行 init-store.sh 时,如果尚未创建 admin 用户,命令会要求输入管理员密码。对于已经初始化的持久化后端,init-store.sh 会补充认证所需的系统信息,无需删除原有图数据。

# stop the hugeGraph firstly
bin/stop-hugegraph.sh

# 初始化认证系统信息;已有后端数据会被保留
bin/init-store.sh

# start hugeGraph again
bin/start-hugegraph.sh

使用 Docker 时开启鉴权模式

对于镜像 hugegraph/hugegraph 大于等于 1.2.0 的版本,我们可以在启动 docker 镜像的同时开启鉴权模式

具体做法如下:

1. 采用 docker run

在 docker run 中添加环境变量 PASSWORD=xxx(密码可以自由设置)即可开启鉴权模式::

docker run -itd -e PASSWORD=xxx --name=server -p 8080:8080 hugegraph/hugegraph:1.7.0

2. 采用 docker-compose

使用 docker-compose 在环境变量中设置 PASSWORD=xxx即可

version: '3'
services:
  server:
    image: hugegraph/hugegraph:1.7.0
    container_name: server
    ports:
      - 8080:8080
    environment:
      - PASSWORD=xxx

3. 进入容器后重新开启鉴权模式

首先进入容器:

docker exec -it server bash
# 用于快速修改配置, 修改前的文件被保存在conf-bak文件夹下
bin/enable-auth.sh

之后参照 基于鉴权模式启动 即可

4 - 配置 HugeGraphServer 使用 https 协议

概述

HugeGraphServer 默认使用的是 http 协议,如果用户对请求的安全性有要求,可以配置成 https。

服务端配置

修改 conf/rest-server.properties 配置文件,将 restserver.url 的 schema 部分改为 https。

# 将协议设置为 https
restserver.url=https://127.0.0.1:8080
# 服务端 keystore 文件路径,当协议为 https 时该默认值自动生效,可按需修改此项
ssl.keystore_file=conf/hugegraph-server.keystore
# 服务端 keystore 文件密码,当协议为 https 时该默认值自动生效,可按需修改此项
ssl.keystore_password=******

由于 keystore 文件没有声明许可证,发行包中并不包含它。当 restserver.url 以 https 开头而 conf/hugegraph-server.keystore 不存在时,bin/start-hugegraph.sh 会在启动前从 hugegraph-doc 仓库的 binary-1.5 分支下载该文件,其密码为 hugegraph。 这两项都是 ssl.keystore_file 和 ssl.keystore_password 的默认值,用户可以生成自己的 keystore 文件及密码,然后修改这两个配置项。

客户端配置

在 HugeGraph-Client 中使用 https

在构造 HugeClient 时传入 https 相关的配置,代码示例:

String url = "https://localhost:8080";
String graphName = "hugegraph";
HugeClientBuilder builder = HugeClient.builder(url, graphName);
// 客户端 keystore 文件路径
String trustStoreFilePath = "hugegraph.truststore";
// 客户端 keystore 密码
String trustStorePassword = "******";
builder.configSSL(trustStoreFilePath, trustStorePassword);
HugeClient hugeClient = builder.build();

注意:HugeGraph-Client 在 1.9.0 版本以前是直接以 new 的方式创建,并且不支持 https 协议,在 1.9.0 版本以后改成以 builder 的方式创建,并支持配置 https 协议。

在 HugeGraph-Loader 中使用 https

启动导入任务时,在命令行中添加如下选项:

# https
--protocol https
# 客户端证书文件路径,当指定 --protocol 为 https 时,默认值 conf/hugegraph.truststore 自动生效,可按需修改
--trust-store-file {file}
# 客户端证书文件密码,当指定 --protocol 为 https 时,默认值 hugegraph 自动生效,可按需修改
--trust-store-password {password}

hugegraph-loader 的 conf 目录下已经放了一个默认的客户端证书文件 hugegraph.truststore,其密码是 hugegraph。

在 HugeGraph-Tools 中使用 https

执行命令时,在命令行中添加如下选项:

# 客户端证书文件路径,当 url 中使用 https 协议时,默认值 conf/hugegraph.truststore 自动生效,可按需修改
--trust-store-file {file}
# 客户端证书文件密码,当 url 中使用 https 协议时,默认值 hugegraph 自动生效,可按需修改
--trust-store-password {password}
# 执行迁移命令时,当 --target-url 中使用 https 协议时,默认值 conf/hugegraph.truststore 自动生效,可按需修改
--target-trust-store-file {target-file}
# 执行迁移命令时,当 --target-url 中使用 https 协议时,默认值 hugegraph 自动生效,可按需修改
--target-trust-store-password {target-password}

hugegraph-tools 的 conf 目录下已经放了一个默认的客户端证书文件 hugegraph.truststore,其密码是 hugegraph。

如何生成证书文件

本部分给出生成证书的示例,如果默认的证书已经够用,或者已经知晓如何生成,可跳过。

服务端

  1. ⽣成服务端私钥,并且导⼊到服务端 keystore ⽂件中,server.keystore 是给服务端⽤的,其中保存着⾃⼰的私钥
keytool -genkey -alias serverkey -keyalg RSA -keystore server.keystore

过程中根据需求填写描述信息,默认证书的描述信息如下:

名字和姓⽒:hugegraph
组织单位名称:hugegraph
组织名称:hugegraph
城市或区域名称:BJ
州或省份名称:BJ
国家代码:CN
  1. 根据服务端私钥,导出服务端证书
keytool -export -alias serverkey -keystore server.keystore -file server.crt

server.crt 就是服务端的证书

客户端

keytool -import -alias serverkey -file server.crt -keystore client.truststore

client.truststore 是给客户端⽤的,其中保存着受信任的证书

5 - 配置 RocksDB 后端

概述

RocksDB 是一个嵌入式的 LSM-tree 键值存储。使用 rocksdb 后端时,HugeGraph-Server 把全部图数据保存在 服务进程内部的 RocksDB 实例中,不需要额外部署存储服务。发布包中的 conf/graphs/hugegraph.properties 默认使用的就是这个后端。

从 1.7.0 版本开始,服务端只接受 memory、rocksdb、hbase 和 hstore 作为后端。rocksdb 后端把 数据写在单台服务器的本地磁盘上,不支持共享存储,因此多个服务无法基于同一个数据目录提供同一个图。 分布式部署请使用 hstore 后端,配合 PD 与 Store。

RocksDB 的 JNI 依赖在 hugegraph-rocksdb/pom.xml 中固定为 8.10.2 版本,因此磁盘格式与各配置项的 语义都以 RocksDB 8.10 为准。

该后端上报的驱动版本是 1.11,在初始化图时会写入 system store 的 meta 表中。

选择后端

在图配置文件(conf/graphs/<graph>.properties)中设置后端与序列化器:

gremlin.graph=org.apache.hugegraph.HugeFactory

backend=rocksdb
serializer=binary

store=hugegraph

# rocksdb backend config
#rocksdb.data_path=/path/to/disk
#rocksdb.wal_path=/path/to/disk
  • backend=rocksdb 选择 RocksDB 存储实现。
  • serializer=binary 是发布包模板为该后端使用的序列化器。内置的序列化器为 binary、binaryscatter 和 text。
  • store 是该图在后端中的库名,同时也是存储实现拿到的图名的一部分。

首次启动前执行一次 bin/init-store.sh 创建各个 store,然后再启动服务。bin/init-store.sh 与 bin/hugegraph-server.sh 都会加载 RocksDB 库,数据目录在执行这些脚本的机器上创建。

发布包会为打包时 backend.properties 中列出的每个后端注册配置项空间和存储实现,该文件的取值来自 hugegraph.backends 构建属性。默认构建会注册 rocksdb, hbase, hstore;使用 -Drocksdb-only 构建会 激活 rocksdb-only profile,产出的发布包只注册 rocksdb。未注册的后端在启动时会报 Not exists BackendStoreProvider。

注册过程还会额外注册一个名字 rocksdbsst,它对应的实现写出 SST 文件而不是打开一个可用的数据库。 该名字不在允许的后端列表中,因此 backend=rocksdbsst 会被拒绝并报 backend is illegal: rocksdbsst。 如果要把 SST 文件导入普通的 rocksdb 图,请使用下面介绍的 rocksdb.sst_path。

数据目录结构

有两个目录需要关注:rocksdb.data_path(默认 rocksdb-data/data)和 rocksdb.wal_path (默认 rocksdb-data/wal)。相对路径基于服务的工作目录解析,也就是安装目录。

每个图会打开三个 store:m 存放 schema,g 存放图数据,s 是 system store。store 名会拼接到上面 两个路径之后,因此一个默认的单图安装目录如下:

rocksdb-data/
  data/
    m/    # schema store:属性键、顶点/边/索引标签、计数器
    g/    # graph store:顶点、边、索引表、olap 表
    s/    # system store:任务、服务信息、后端 meta(驱动版本)
  wal/
    m/
    g/
    s/

后端的每张表在所属 store 中对应一个 RocksDB 列族,名字形如 <database>+<table>,其中 database 由图名 推导得到。已有数据目录中的列族总是会被重新打开,因此旧版本创建的表仍然可读。

还需要注意:

  • 两个图不能共用同一个数据路径。通过 API 基于已有配置克隆创建图时,存储实现会在 rocksdb.data_path 和 rocksdb.wal_path 后面追加 _<newGraph>。删除这样的图会同时删除这两个目录。
  • 快照创建在数据目录旁边:数据路径的最后两段会加上前缀重写,因此在默认路径下 graph store 的快照位于 rocksdb-data/<prefix>_data/g。恢复快照时会先关闭实例,删除数据目录,再把快照移动到原位置。
  • 设置了 rocksdb.data_disks 时,其中列出的表会在指定路径下作为独立的 RocksDB 实例打开,而不再放在 rocksdb.data_path 下。服务最多并发打开 8 个实例,打开最多等待 600 秒,会话关闭最多等待 30 秒。

路径与日志配置项

config optiondefault valuedescription
rocksdb.data_pathrocksdb-data/dataRocksDB 数据存储路径,不允许为空。
rocksdb.data_disks[]为部分表指定独立磁盘,每个元素格式为 STORE/TABLE: /path/disk。允许的键为 [g/vertex, g/edge_out, g/edge_in, g/vertex_label_index, g/edge_label_index, g/range_int_index, g/range_float_index, g/range_long_index, g/range_double_index, g/secondary_index, g/search_index, g/shard_index, g/unique_index, g/olap]。磁盘路径不能与 rocksdb.data_path 相同。
rocksdb.wal_pathrocksdb-data/walRocksDB WAL 存储路径,不允许为空。
rocksdb.sst_path(空)待导入 RocksDB 的 SST 文件所在路径,为空表示不导入。
rocksdb.log_levelINFORocksDB 的日志级别,可选值:DEBUG、INFO、WARN、ERROR、FATAL、HEADER。

压缩与合并配置项

config optiondefault valuedescription
rocksdb.num_levels7数据库的层数,取值范围 1 到 2^31-1。
rocksdb.compaction_styleLEVELRocksDB 的 compaction 策略:LEVEL/UNIVERSAL/FIFO。
rocksdb.optimize_modetrue针对高负载和大数据量做优化,具体行为见下文的配置项生效方式一节。
rocksdb.bulkload_modefalse切换到批量导入数据的模式。
rocksdb.compression_per_level[none, none, snappy, snappy, snappy, snappy, snappy]各层使用的压缩算法,可选值为 none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd。列表必须为空,或者元素个数正好等于 rocksdb.num_levels。
rocksdb.bottommost_compressionnone最底层使用的压缩算法,可选值为 none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd。
rocksdb.compressionsnappy压缩数据块使用的压缩算法,可选值为 none/snappy/z/bzip2/lz4/lz4hc/xpress/zstd。

数据库级配置项

config optiondefault valuedescription
rocksdb.max_background_jobs8后台任务(包括 flush 和 compaction)的最大并发数,取值范围 1 到 2^31-1。
rocksdb.max_subcompactions4单个 compaction 任务使用的最大线程数,取值范围 1 到 2^31-1。
rocksdb.delayed_write_rate16777216(16 MB/s)当 compaction 落后需要降速时,用户写请求的限速值,单位字节每秒。
rocksdb.max_open_files-1RocksDB 可缓存的最大打开文件数,-1 表示不限制。
rocksdb.max_manifest_file_size104857600(100 MB)manifest 文件的最大字节数。
rocksdb.skip_stats_update_on_db_openfalse打开数据库时是否跳过统计信息更新,设为 true 表示不更新统计信息。
rocksdb.skip_check_sst_size_on_db_openfalse打开数据库时是否跳过检查所有 sst 文件的大小。
rocksdb.max_file_opening_threads16打开文件使用的最大线程数,取值范围 1 到 2^31-1。
rocksdb.max_total_wal_size0WAL 文件的总大小上限,单位字节。超过后会强制 flush 相关的列族,0 表示不限制。
rocksdb.bytes_per_sync0允许操作系统在后台异步地增量同步正在写入的 SST 文件,每写入这么多字节发起一次请求,0 表示关闭。
rocksdb.wal_bytes_per_sync0同上,作用于 WAL 文件,0 表示关闭。
rocksdb.strict_bytes_per_syncfalse为 true 时保证任意时刻提交回写的 SST/WAL 数据不超过 bytes_per_sync/wal_bytes_per_sync 字节,可用于处理写入速度超过 I/O 速度的场景。
rocksdb.db_write_buffer_size0所有列族 write buffer 的总大小上限,单位字节,0 表示不限制。
rocksdb.log_readahead_size0读取日志时预取的字节数,0 表示关闭预取。
rocksdb.compaction_readahead_size0compaction 时批量读取的字节数。如果 RocksDB 跑在机械盘上,建议至少设为 2MB,0 表示关闭预取。
rocksdb.row_cache_capacity0表级行缓存的全局容量,单位字节,0 表示关闭 row_cache。
rocksdb.delete_obsolete_files_period21600(6 小时)删除废弃文件的周期,单位秒,0 表示每次都做完整清理。该值传给 RocksDB 前会换算成微秒。

Memtable 配置项

config optiondefault valuedescription
rocksdb.write_buffer_size134217728(128 MB)在内存中累积的数据量,单位字节,最小 1 MB。该值针对单个列族生效。
rocksdb.max_write_buffer_number6内存中累积的 write buffer 的最大个数,取值范围 1 到 2^31-1。
rocksdb.min_write_buffer_number_to_merge2会被合并到一起的 write buffer 的最小个数,取值范围 1 到 2^31-1。
rocksdb.max_write_buffer_number_to_maintain0使用事务时,为冲突检查在内存中保留的 write buffer 总数上限。
rocksdb.memtable_bloom_size_ratio0.0设置了 prefix-extractor 且该值不为 0,或者 memtable_whole_key_filtering 为 true 时,为 memtable 创建大小为 write_buffer_size * memtable_bloom_size_ratio 的布隆过滤器。大于 0.25 的取值会被收敛为 0.25,取值范围 0.0 到 1.0。
rocksdb.memtable_whole_key_filteringfalse在 memtable 中启用全键布隆过滤器,可以降低点查的 CPU 开销。只有 memtable_bloom_size_ratio > 0 时才生效。
rocksdb.memtable_huge_page_size0memtable 中布隆过滤器使用的 huge page TLB 的页大小,小于等于 0 时不从 huge page TLB 分配,改用 malloc。
rocksdb.inplace_update_supportfalse当写入的键已存在于当前 memtable 且新值更小时,允许线程安全的原地更新。

层级大小与写入限流配置项

config optiondefault valuedescription
rocksdb.level_compaction_dynamic_level_bytesfalse是否启用 level_compaction_dynamic_level_bytes。启用后 max_bytes_for_level_multiplier 的优先级高于 max_bytes_for_level_base,基准层的大小是动态的,LSM tree 更可预测,有助于限制最差情况下的空间放大。对已有数据库开关该特性可能造成异常的 LSM tree 结构,因此不推荐改动。
rocksdb.max_bytes_for_level_base536870912(512 MB)level-1 文件总大小的上限,单位字节,最小 1 MB。
rocksdb.max_bytes_for_level_multiplier10.0对所有层 L,第 L+1 层文件总大小与第 L 层文件总大小的比值,最小 1.0。
rocksdb.target_file_size_base67108864(64 MB)compaction 的目标文件大小,单位字节,最小 1 MB。
rocksdb.target_file_size_multiplier1第 L 层文件与第 L+1 层文件的大小比值。
rocksdb.level0_file_num_compaction_trigger2触发 level-0 compaction 的文件数。
rocksdb.level0_slowdown_writes_trigger20触发写入降速的 level-0 文件数软上限。
rocksdb.level0_stop_writes_trigger36触发停写的 level-0 文件数硬上限。
rocksdb.soft_pending_compaction_bytes_limit68719476736(64 GB)待 compaction 数据量的软上限,单位字节,最小 1 GB。
rocksdb.hard_pending_compaction_bytes_limit274877906944(256 GB)待 compaction 数据量的硬上限,单位字节,最小 1 GB。

文件 I/O 配置项

config optiondefault valuedescription
rocksdb.allow_mmap_writesfalse允许操作系统以 mmap 方式写文件。
rocksdb.allow_mmap_readsfalse允许操作系统以 mmap 方式读取 sst 文件。
rocksdb.use_direct_readsfalse读取 sst 文件时使用直接 I/O。
rocksdb.use_direct_io_for_flush_and_compactionfalseflush 和 compaction 时使用直接读写。
rocksdb.use_fsyncfalse为 true 时每次持久化都会执行 fsync。
rocksdb.atomic_flushfalse为 true 时多个列族的 flush 结果会原子地提交到 MANIFEST。WAL 一直开启的情况下不必设置该项。

SST 表格式与块缓存配置项

config optiondefault valuedescription
rocksdb.format_version5BlockBasedTable 的格式版本,可选值为 0~5。
rocksdb.index_typekBinarySearchsst 文件中数据块之间查找使用的索引类型,可选值为 [kBinarySearch, kHashSearch, kTwoLevelIndexSearch, kBinarySearchWithFirstKey]。
rocksdb.data_block_index_typekDataBlockBinarySearchsst 文件数据块内点查使用的查找类型,可选值为 [kDataBlockBinarySearch, kDataBlockBinaryAndHash]。
rocksdb.data_block_hash_table_util_ratio0.75哈希表 entries/buckets 的使用率,仅在 data_block_index_type=kDataBlockBinaryAndHash 时有效,取值范围 0.0 到 1.0。
rocksdb.block_size4096(4 KB)每个块中打包的用户数据的近似大小,注意对应的是未压缩的数据。
rocksdb.block_size_deviation10用于结束一个块的空闲空间百分比,取值范围 0 到 100。
rocksdb.block_restart_interval16块内增量编码的 restart 间隔。
rocksdb.block_cache_capacity8388608(8 MB)RocksDB 使用的块缓存大小,单位字节,0 表示不使用块缓存。每个列族都会创建一个该大小的独立缓存。

布隆过滤器配置项

只有 rocksdb.bloom_filter_bits_per_key 大于等于 0 时,本组的配置项才会被读取。默认值 -1 表示不启用 布隆过滤器,此时本表中其他配置项都不生效,包括索引和过滤块的缓存相关项。

config optiondefault valuedescription
rocksdb.bloom_filter_bits_per_key-1布隆过滤器中每个键占用的位数,10 是一个不错的取值,对应约 1% 的误判率。大于 0 表示启用布隆过滤器,-1 表示不启用(0~0.5 向下取整为不启用)。
rocksdb.bloom_filter_block_based_modefalse启用布隆过滤器时,设为 true 表示使用 block based filter 而不是 full filter。
rocksdb.bloom_filter_whole_key_filteringtrue启用布隆过滤器时,设为 true 表示把完整的键放入布隆过滤器,否则在设置了 prefix-extractor 时放入键的前缀。
rocksdb.cache_index_and_filter_blockstrue设为 true 表示把索引块和过滤块放入块缓存。
rocksdb.pin_l0_filter_and_index_blocks_in_cachetrue设为 true 表示把 L0 的索引块和过滤块固定在块缓存中。
rocksdb.optimize_filters_for_hitstrue启用布隆过滤器时,该开关允许不为最后一层存储过滤器,设为 true 表示主要针对命中的场景优化过滤器,而不是同时优化未命中的场景。该项在过滤器关闭时也会生效。
rocksdb.partition_filters_and_indexesfalse启用布隆过滤器时,设为 true 表示每个 sst 文件使用分区的 full filter 和索引。该项与 block based filter 不兼容。开启后索引类型会被强制为 kTwoLevelIndexSearch,元数据块大小取 rocksdb.block_size。
rocksdb.pin_top_level_index_and_filtertrue当 partition_filters_and_indexes 为 true 时,设为 true 表示把分区过滤块和索引块的顶层索引固定在块缓存中。
rocksdb.prefix_extractor_n_bytes0prefix-extractor 取键的前 N 个字节作为前缀,键长度小于 N 时使用完整键,0 表示不设置 prefix-extractor。

配置项的生效方式

服务为每个 store 和每个列族构建一次 RocksDB 的选项对象,因此上面任何配置项的修改都在下次启动服务时 生效。

  • rocksdb.optimize_mode=true 会在应用上表中的取值之前先套用一组预设:数据库层面把并行度提高到可用 处理器数的一半(至少 1),允许 memtable 并发写入,并开启写线程自适应让出;列族层面调用 RocksDB 的 level-style 和 universal-style compaction 预设。显式配置项在之后应用,因此配置文件中写明的取值会覆盖 预设。
  • rocksdb.bulkload_mode=true 会关闭自动 compaction,把三个 level-0 触发阈值提高到 int 最大值,把两个 待 compaction 上限提高到 long 最大值。导入结束后要关闭它并重启,否则 compaction 不会运行。
  • rocksdb.block_cache_capacity=0 表示彻底关闭块缓存,而不是不限制大小。
  • rocksdb.prefix_extractor_n_bytes 大于 0 时会安装一个该长度的 capped prefix extractor。
  • 所有列族都使用 uint64add merge 操作符,计数器表依赖它。
  • 数据库不存在时会自动创建,avoid_unnecessary_blocking_io 和 write_dbid_to_manifest 始终开启。

内存说明

RocksDB 的缓存和 write buffer 都是本地内存分配,不属于 bin/hugegraph-server.sh 中设置的 JVM 堆。 GET /metrics/backend 接口会返回存储的使用量:内存数值是所有已打开列族的块缓存用量、固定在块缓存中的 用量、预估的 table reader 内存(索引块和过滤块)以及全部 memtable 大小之和,取自 RocksDB 的属性。

有两个配置项的实际占用会随列族数量成倍增长:

  • rocksdb.block_cache_capacity 为每个列族创建一个缓存实例,因此一台服务的块缓存总量大致等于该值乘以 所有图的 m、g、s 三个 store 中已打开表的数量,再加上 rocksdb.data_disks 额外打开的实例。
  • rocksdb.write_buffer_size 乘以 rocksdb.max_write_buffer_number 限定的是单个列族的 memtable 内存。 rocksdb.db_write_buffer_size 限制一个 store 内所有列族的总量,默认值 0 表示没有这个限制。

rocksdb.row_cache_capacity 不同:它是每个 store 一个缓存,0 表示关闭。

导入 SST 文件

设置 rocksdb.sst_path 即开启导入。打开 store 时以及每次创建表时,服务会遍历 <sst_path>/<column family>/ 目录,收集其中所有非空的 *.sst 文件,导入到对应的列族。导入采用移动 文件的方式而不是复制,因此源目录会被导入过程消耗掉。

raft 模式

RocksDB 后端仍然可以运行在 raft 状态机之后:raft.mode=true 时,本地后端的存储实现会被 raft 实现包装。 包装层会拒绝共享存储的后端,因此 rocksdb 可用而 hbase 不可用。raft 模式下 RocksDB 会话写入时关闭 WAL 且不做 sync,因为状态机可以通过快照加 raft 日志恢复,而该后端支持快照。

使用时需要注意:

  • bin/init-store.sh 在初始化后端时会强制把 raft.mode 置为 false,因此初始化过程不会走 raft。
  • 发布包中的 conf/graphs/hugegraph.properties 已把 raft 相关配置标记为废弃。1.7.0 及之后版本的分布式 部署改用 hstore 后端,配合 PD 与 Store。
  • raft 成员管理接口位于 graphspaces/{graphspace}/graphs/{graph}/raft/ 之下,包括 list_peers、 get_leader、set_leader、transfer_leader、add_peer 和 remove_peer。bin/raft-tools.sh 封装了 同样的操作,但它拼接的 URL 中仍然没有 graphspace 段,在 1.7.0 的服务上需要调整路径才能使用。
  • 其余 raft.* 配置项见 Server 配置选项。

后端能力

该后端的特性开关决定了哪些操作可以下推给存储:

  • 支持按键前缀扫描、按键范围扫描、分页查询、范围条件和 order-by。
  • RocksDB 内部没有索引,因此按名字查询 schema、按标签查询以及按标签删除边由服务端完成,而不是由存储 完成。
  • 通过 RocksDB 的 write batch 支持事务。
  • 支持快照,raft 模式和备份依赖该能力。
  • 不支持共享存储,一个数据目录属于一台服务。
  • 支持 olap 属性,对应的表会作为额外的列族创建。
  • 存储本身不会让数据过期,因此服务端在读取时过滤掉 TTL 已到期的元素。
  • 存储层不支持 in、contains、contains_key 条件,不支持聚合属性,也不支持原地更新顶点或边的属性。

riscv64 平台说明

在 Linux riscv64 上,RocksDB 的 JNI 库需要 libatomic.so.1。bin/util.sh 会查找该库并在 bin/hugegraph-server.sh、bin/init-store.sh 和 bin/dump-store.sh 启动 JVM 之前把它加入 LD_PRELOAD。如果找不到,这些脚本会以 RISC-V RocksDB requires libatomic.so.1; install libatomic1 退出,安装 libatomic1 包即可解决。

6 - 配置 HStore 分布式后端

1 概述

hstore 是 HugeGraph 的分布式存储后端。图使用该后端时,HugeGraph-Server 本地磁盘上不保存任何图数据, 数据由另外两个进程负责:

  • HugeGraph-PD(Placement Driver)保存集群元数据:已注册的 Store 列表、每个图的分区布局、分区到 Store 的映射关系、图的 Schema 以及 Schema 的 id 计数器。
  • HugeGraph-Store 保存实际的键值数据,并通过 Raft 在多个 Store 节点之间复制。

Server 进程内嵌了 PD 客户端和 Store 客户端。每次读写时,它先向 PD 查询该 key 属于哪个分区、当前哪个 Store 节点是这个分区的 leader,然后把请求直接发给这个 Store 节点。

服务端的适配层是 hugegraph-hstore 模块,它以后端名 hstore 注册,驱动版本为 1.13。

选择 hstore 影响的不只是数据写到哪里,Server 还会根据后端类型切换下列行为:

方面使用 hstore使用本地后端
Schema 存储通过 PD 元数据驱动读写 SchemaSchema 保存在 m store 中
Schema id通过 PD 客户端由 PD 分配由 schema store 分配
System store没有独立的 system store,系统数据写入 graph store独立的 s store
任务调度器distributedlocal
权限管理器StandardAuthManagerV2StandardAuthManager
后端版本校验读取 graph store读取 system store
init-store.sh跳过该图,元数据由 PD 和 Store 负责创建本地 store

2 前置条件

hstore 不能独立工作。在 Server 打开 hstore 图之前,PD 集群和至少一个 Store 节点必须已经运行, 并且启动顺序如下:

  1. PD,先启动以便组成 Raft 组。
  2. Store,通过 gRPC 向 PD 注册。gRPC 地址出现在 PD 自身 pd.initial-store-list 中的 Store 会直接进入 Up 状态;不在该列表中、并且 PD 从未见过它处于 Up 或 Offline 的 Store 会注册为 Pending, 需要先激活才能提供数据服务。
  3. Server,随后从 PD 读回 Store 列表。

服务端需要关注的默认端口:

进程gRPC 端口REST 端口
PD86868620
Store85008520

服务端的 pd.peers 指向 PD 的 gRPC 端口,而不是 REST 端口。

另外两个进程的安装与配置方式,参见 安装/构建 HugeGraph-PD 和 安装/构建 HugeGraph-Store。

3 选择 hstore 后端

3.1 图配置文件

在图的属性文件(例如 conf/graphs/hugegraph.properties)中设置后端:

backend=hstore
serializer=binary
store=hugegraph
pd.peers=127.0.0.1:8686

关于这四个配置项:

  • backend=hstore 选择该适配层。自 1.7.0 起允许的取值为 memory、rocksdb、hbase 和 hstore。 发行包中做校验的位置对该值不区分大小写。
  • serializer=binary 是必需的。注册 hstore 后端时只注册了配置空间和存储 provider,并没有注册自己的 序列化器,适配层就是按二进制序列化器编写的。serializer 的内置默认值是 text,因此必须显式写出该项。
  • store=hugegraph 是 PD 看到的图名中的命名空间部分。Server 以 <graphspace>/<store> 打开 provider, 每个底层 store 再追加自己的后缀,因此 PD 中每个 store 对应一个图条目:图数据是 DEFAULT/hugegraph/g,schema store 位是 DEFAULT/hugegraph/m。graphspace 默认为 DEFAULT, g 和 m 是固定的。
  • pd.peers 是以逗号分隔的 PD gRPC 地址列表。适配层从图配置中读取该项,而不是从 rest-server.properties 中读取,图级别的元数据连接也使用同一个值。

如果图配置文件中没有 pd.peers,那么在加载图时,只要 usePD 为 true 或者后端是 hstore, Server 会把 rest-server.properties 中的值复制到图配置里。不过在图配置文件中显式写出该项更清晰。

3.2 rest-server.properties

# use pd
usePD=true
pd.peers=127.0.0.1:8686

usePD=true 让 Server 在启动时从 PD 加载元数据。在这条路径上,它会把元数据管理器连接到 PD,创建内置的 admin 账号和默认图空间,加载图空间与服务,创建内部的系统图(后端固定为 hstore),并加载 PD 中保存的 图配置。

它和图级别的 backend=hstore 是两个独立的开关:一个图可以使用 hstore 而 usePD 保持默认的 false, 此时 Server 不会走基于 PD 的元数据路径。发行包自带的测试启动脚本在后端为 hstore 时会设置该项。

3.3 发行包中的模板文件

发行包在 conf/graphs/hstore.properties.template 中提供了一份该后端的现成图配置文件。它与 hugegraph.properties 的差别是:把 backend 设为 hstore、不注释 pd.peers=127.0.0.1:8686、 并且不包含内存管理配置段。

hstore 的 Docker 镜像会自动套用这份模板:它删除 conf/graphs/hugegraph.properties,再把模板重命名过去, 因此容器启动时就已经选好了 hstore 后端。

本地构建的发行包默认编译了 hstore provider。rocksdb-only 这个 Maven profile 会把编译进去的后端列表 收窄为只有 rocksdb,用这种方式构建出来的发行包会以 Unsupported backend type 拒绝 backend=hstore。

4 hstore 配置项

hstore 配置空间中只有下面两个配置项,它们写在图的属性文件里。

配置项默认值说明
hstore.partition_count0分区数量,PD 依据该值控制分区(Number of partitions)。
hstore.shard_count0副本数量,PD 依据该值控制分区副本(Number of copies)。

4.1 hstore.partition_count

每个 graph store 第一次被打开时,Server 会把这个数字连同图名一起发给 PD。取负值会在此处被拒绝, 报错信息为 The value of hstore.partition_count cannot be less than 0.

PD 对该值的处理方式:

  • 0,也就是默认值,表示交给 PD 决定。对图数据 store,PD 使用自身集群级别的分区总数,该总数由 pd.initial-store-list 中的条目数量、partition.store-max-shard-count 和 partition.default-shard-count 推算得出;对 /m 和 /s store 固定使用 1。
  • 取值在 1 到该总数之间时,按原值使用。
  • 取值大于该总数时,会被下调到该总数。

该数字在 store 首次向 PD 注册时生效,之后再修改属性文件不会让已有的图重新分区。

4.2 hstore.shard_count

hstore.shard_count 声明在 hstore 配置空间中,属性文件里也接受该项,但当前版本服务端没有任何代码读取它: 适配层读取的只有 hstore.partition_count 一项。实际生效的副本数由 PD 的配置决定,即 PD application.yml 中的 partition.default-shard-count。

5 只在 hstore 模式下生效的其他配置项

下列配置项位于公共的 rest-server.properties 和图属性文件中,但只有在使用 PD 和 hstore 后端时才生效, 或者才会改变行为。source 列给出该配置项在 HugeGraph master 分支上的声明位置(文件与行号)。

配置项文件默认值在 hstore 模式下的作用source
pd.peersrest-server.properties127.0.0.1:8686用于元数据、服务发现和系统图的 PD 地址ServerOptions.java:195-201
pd.peers{graph}.properties127.0.0.1:8686后端适配层自身使用的 PD 地址CoreOptions.java:649-654
usePDrest-server.propertiesfalseServer 启动时是否从 PD 加载元数据ServerOptions.java:390-396
clusterrest-server.propertieshg-test集群名,作为所有 PD 元数据 key 的前缀ServerOptions.java:187-193
init_store.enabledrest-server.propertiestruePD/Store 部署下应设为 false,元数据已由存储侧负责ServerOptions.java:371-380
graph.load_from_local_configrest-server.propertiesfalse启动时是否在 PD 中的图配置之外,额外扫描 conf/graphsServerOptions.java:355-361
auth.graph_storerest-server.propertieshugegraph保存权限数据的图,关闭 init-store 时会校验它使用 hstore 后端ServerOptions.java:591-598
graphspace{graph}.propertiesDEFAULTPD 看到的图名的第一段CoreOptions.java:679-685

init-store.sh 从不初始化 hstore 图。在开启的路径上,它扫描 conf/graphs 并跳过后端为 hstore 的每一个 图。如果用 init_store.enabled=false 整体关闭这一步,它会改为校验 admin 账号仍然能在 PD 启动路径上被创建: usePD 必须为 true、权限图必须存在于本地配置中且后端为 hstore、auth.admin_pa 必须显式设置为非空值。 否则启动会直接失败,而不是使用公开的默认密码创建账号。

6 Server 如何通过 PD 发现 Store

适配层在进程中第一次打开 hstore 图时,一次性构建这些客户端:

  1. 用 pd.peers 构建 PD 客户端配置,带上 PD 的鉴权凭据,并开启客户端侧的分区缓存。
  2. 创建进程级的 PD 客户端。
  3. 用该 PD 客户端创建进程级的 Store 客户端。

创建 Store 客户端时,会把一个基于 PD 的分区器同时注册为 Store 客户端节点管理器的 node provider、 partitioner 和 notifier。路由逻辑全部在这个分区器中:

  • 单点和前缀请求:向 PD 查询拥有该 key 的分区,取该分区的 leader 副本,把请求发到对应的 store id。
  • 按 code 的范围扫描:按 code 逐个遍历分区直到覆盖整个范围,每个分区产生一个目标 Store。
  • 全图扫描:向 PD 查询该图的活跃 Store,并向全部 Store 扇出请求。
  • Store 地址解析:通过 PD 把 store id 解析成主机和端口。
  • 缓存失效:当某个 Store 返回分区 leader 已迁移时,notifier 会更新 PD 客户端缓存中的分区 leader 并使过期的分区条目失效,之后的请求就会跟随新的 leader。

由于 Store 列表来自 PD 而不是配置文件,增删 Store 节点只需要针对同一个 PD 集群启动或停止它, 服务端不需要改任何配置。

7 后端能力

hstore 并不支持本地后端的所有查询形式。对用户可见的差异如下:

特性是否支持
按 key 前缀扫描支持
按 key 范围扫描支持
带范围条件的查询支持
带 order by 的查询支持
分页查询支持
OLAP 属性支持
Task 和 Server 顶点支持
Scan token不支持
按名称查询 Schema不支持
按 label 查询不支持
带 in 条件的查询不支持
带 contains 的查询不支持
带 contains key 的查询不支持
按输入 id 顺序排序不支持
按 label 删除边不支持
更新顶点属性不支持
更新边属性不支持
事务不支持
Number 类型不支持
聚合属性不支持
TTL不支持

不支持按输入 id 顺序排序,是因为多节点批量扫描会按 Store 对输入 key 分组,从而丢失全局顺序; 不支持更新顶点和边属性,是因为属性被存放在单个 cell 中。

8 验证

Server 启动后,后端指标接口会返回 PD 当前认为处于活跃状态的 Store 数量:

curl http://localhost:8080/metrics/backend

响应中的 nodes 就是 PD 返回的活跃 Store 数量。nodes 为 0 说明 Server 连上了 PD,但 PD 中没有状态为 Up 的 Store,通常是 Store 节点还没注册,或者因为不在 PD 的 pd.initial-store-list 中而注册成了 Pending。

7 - 配置 HBase 后端

概述

HBase 后端将图数据存储在 Apache HBase 表中。HugeGraph 仅作为 HBase 客户端:它通过 HBase 的 ZooKeeper 集群连接,为每个图创建一个 HBase namespace,并在其中创建该图的 schema 表、数据表和索引表。计数查询由 HBase 的 AggregateImplementation 协处理器完成,HugeGraph 在创建每张表时都会挂载该协处理器。

注意:HBase 后端已废弃,计划在 HugeGraph 2.0 中移除。新部署请使用 hstore(分布式)或 rocksdb(内嵌,默认值),已有的 HBase 部署请规划迁移。

自 1.7.0 起,发行包内置的后端只有 hstore、rocksdb、hbase 和 memory。HBase provider 上报的后端驱动版本为 1.12。

支持的 HBase 版本

客户端 jar 固定为 HBase 2.6.5(hbase-endpoint 加 hbase-shaded-client)。服务端要求 HBase 2.x:当检测到的 HBase 版本低于 2.0 时,scan 逻辑会把 inclusive stop row 改写为 exclusive 并追加一个 0 字节,因为该版本之前 inclusive stop row 不生效。CI 任务和本地 Docker 镜像都使用 HBase 2.6.5,后端也是针对这个版本做测试的。

选择该后端

修改需要使用 HBase 的图的 conf/graphs/hugegraph.properties:

backend=hbase
serializer=hbase

# namespace 名称由该值推导得出
store=hugegraph

hbase.hosts=localhost
hbase.port=2181
hbase.znode_parent=/hbase

注意:serializer 必须设置为 hbase,而不是 binary。HBase 序列化器是 BinarySerializer 的子类,它不在 rowkey 中写入 id 前缀,并写入预分区的顶点表和边表所需要的分区前缀。使用 serializer=binary 时这两点都不生效。

然后初始化后端并启动服务:

./bin/init-store.sh
./bin/start-hugegraph.sh

默认发行包构建时包含 rocksdb, hbase, hstore 三个后端,无需额外引入 jar。使用 rocksdb-only Maven profile 构建的发行包不包含 HBase 后端,此时 backend=hbase 会以 Not exists BackendStoreProvider: hbase 打开失败。

下面所有配置项都位于图配置文件(conf/graphs/hugegraph.properties)中,而不是 rest-server.properties。只有当发行包包含 hbase 后端时,这些配置项才会被注册。

连接配置项

配置项默认值说明
hbase.hostslocalhostHBase ZooKeeper 的主机名或 IP 地址,多个以逗号分隔,不允许为空。对应 hbase.zookeeper.quorum。
hbase.port2181HBase ZooKeeper 的端口,取值范围 1 到 65535。对应 hbase.zookeeper.property.clientPort。
hbase.znode_parent/hbaseHBase ZooKeeper 的 znode 父路径,不允许为空。对应 zookeeper.znode.parent。
hbase.zk_retry3HBase ZooKeeper 的恢复重试次数,取值范围 0 到 1000。对应 zookeeper.recovery.retry。
hbase.threads_max64HBase 连接的最大线程数,取值范围 1 到 1000。对应 hbase.hconnection.threads.max,HBase 自身默认值为 256,这里取更小的值以避免内存溢出。

超时配置项

配置项默认值说明
hbase.truncate_timeout30等待后端 truncate 的超时时间,单位秒,必须为正数。该超时按 store 计算,而一个图有三个 store,因此一次 truncate 最多耗时该值的三倍。
hbase.aggregation_timeout43200(12 小时)等待聚合的超时时间,单位秒,必须为正数。它会设置计数查询所用聚合客户端的 hbase.rpc.timeout。

Kerberos 与 HBase 配置文件配置项

配置项默认值说明
hbase.kerberos_enablefalse是否为 HBase 启用 Kerberos 认证。
hbase.krb5_conf/etc/krb5.confKerberos 配置文件,包含 KDC IP、默认 realm 等。会被设置为 java.security.krb5.conf 系统属性。
hbase.hbase_site/etc/hbase/conf/hbase-site.xmlHBase 的配置文件。无论是否启用 Kerberos,每次建立连接时都会把它作为配置资源加载。
hbase.kerberos_principal(空)Kerberos 认证使用的 HBase principal。
hbase.kerberos_keytab(空)Kerberos 认证使用的 HBase keytab 文件。

当 hbase.kerberos_enable=true 时,HugeGraph 会在连接上把 hadoop.security.authentication 和 hbase.security.authentication 设置为 kerberos,然后在打开连接之前用配置的 principal 从 keytab 登录。因此 Kerberos 环境下 hbase.krb5_conf、hbase.hbase_site、hbase.kerberos_principal 和 hbase.kerberos_keytab 四项都必须有效:

hbase.kerberos_enable=true
hbase.krb5_conf=/etc/krb5.conf
hbase.hbase_site=/etc/hbase/conf/hbase-site.xml
hbase.kerberos_principal=hugegraph/host@EXAMPLE.COM
hbase.kerberos_keytab=/etc/security/keytabs/hugegraph.keytab

即使关闭 Kerberos,hbase.hbase_site 也会被读取,路径不存在时相当于加载了一个空资源。当需要上述配置项之外的 HBase 设置时,把它指向集群自身的 hbase-site.xml。

预分区配置项

配置项默认值说明
hbase.enable_partitiontrue是否为 HBase 启用预分区。它同时决定后端是否声明支持前缀扫描和范围扫描。
hbase.vertex_partitions10HBase 顶点表的分区数,不允许为负数。
hbase.edge_partitions30HBase 边表的分区数,不允许为负数。

启用预分区后,顶点表按 hbase.vertex_partitions 个 region 创建,两张边表各按 hbase.edge_partitions 个 region 创建,序列化器会在 rowkey 前面加上 id 哈希得到的分区前缀。

注意:请在初始化后端之前,按实际数据量和 region server 数量调整分区数。它对导入速度影响很大,并且只在建表时生效。

关闭 hbase.enable_partition 会恢复不带前缀的原始 rowkey。作为交换,后端此时会声明支持前缀扫描和范围扫描,这两类扫描在预分区 rowkey 下无法工作。

Namespace 与表结构

每个图对应一个 HBase namespace,名称为 <graphspace>/<store> 转小写,并把 / 替换为 _,因为 HBase namespace 名称只允许字母数字和 _ 字符。在默认配置 graphspace=DEFAULT、store=hugegraph 下,namespace 为 default_hugegraph。

在该 namespace 内,一个图包含三个 store:schema store m、graph store g 和 system store s:

Store表
schema (m)VL、EL、PK、IL、C、m_si
graph (g)g_v、g_oe、g_ie、g_si、g_vi、g_ei、g_ii、g_fi、g_li、g_di、g_ai、g_hi、g_ui
system (s)s_v、s_oe、s_ie、s_si、s_vi、s_ei、s_ii、s_fi、s_li、s_di、s_ai、s_hi、s_ui、M

g_v 是顶点表,g_oe 和 g_ie 分别是出边表和入边表,其余 g_* 表依次是二级索引、顶点标签索引、边标签索引、范围索引(int、float、long、double)、全文索引、shard 索引和唯一索引表。所有表都只有一个名为 f 的列族,并且都在建表时挂载了 org.apache.hadoop.hbase.coprocessor.AggregateImplementation 协处理器。只有 g_v、g_oe 和 g_ie 会预分区,system store 中同名的那几张表按单个 region 创建。

system store 中的 M 表保存 init-store.sh 写入的后端版本。truncate 图时会排除该表,因为丢失它会导致下次启动的版本校验失败。清空图会删除这些表;连同存储空间一起清空则会删除整个 namespace。

GET /metrics/backend 会返回 HBase 集群状态:cluster_id、master_name、average_load、hbase_version、region_count、leaving_servers、nodes、region_servers,以及一个 servers map,其中包含每个 region server 的堆内存、磁盘、请求数和 region 明细。PUT /graphspaces/{graphspace}/graphs/{name}/compact 会请求 HBase 对该图的所有表做 compaction。

使用 Docker 做本地测试

Server 仓库中的 docker/hbase 会构建一个 HBase 2.6.5 单机镜像(hugegraph/hbase:2.6.5,容器名 hg-hbase-test),用于本地开发和测试。以下命令都在仓库根目录执行。

为运行在宿主机上的 HugeGraph 启动 HBase:

docker compose -p hg-hbase -f docker/hbase/docker-compose.hbase.yml build --no-cache hbase
HBASE_MASTER_HOSTNAME=localhost HBASE_REGIONSERVER_HOSTNAME=localhost \
docker compose -p hg-hbase -f docker/hbase/docker-compose.hbase.yml up -d
until docker exec hg-hbase-test nc -z localhost 2181 >/dev/null 2>&1; do sleep 2; done

为运行在同一个 Docker 网络中的容器化 HugeGraph 启动 HBase:

HBASE_HOSTNAME=hbase docker compose -p hg-hbase -f docker/hbase/docker-compose.hbase.yml up -d

对外公布的主机名很重要:容器启动时会把 HBASE_MASTER_HOSTNAME 和 HBASE_REGIONSERVER_HOSTNAME 写入自己的 hbase-site.xml,未设置时回退到 HBASE_HOSTNAME(默认 hbase)。如果客户端无法解析这个主机名,即使 ZooKeeper 可用,也会报 UnknownHostException: hbase:16000。

映射到宿主机的端口:

端口服务
2181ZooKeeper,与 hbase.port 默认值一致
16000HBase Master RPC
16010HBase Master Web UI,http://localhost:16010
16020HBase RegionServer RPC
16030HBase RegionServer Web UI,http://localhost:16030

针对它运行后端测试:

mvn test -pl hugegraph-server/hugegraph-test -am -P core-test,hbase

停止并删除数据卷:

docker compose -p hg-hbase -f docker/hbase/docker-compose.hbase.yml down -v

该镜像会分别启动 ZooKeeper、master 和 region server 三个守护进程,并等到 master 上报有存活的 server 之后才开始 tail 日志,因此首次启动会比较慢。请给 Docker 分配至少 4 GB 内存。compose 的健康检查也因此设置了 90 秒的 start period。

限制

HBase 后端不支持以下特性:

  • 事务。rollback 只会丢弃尚未提交的批次,而 commit 是逐表写入的,因此跨表不是原子的。
  • 原地更新单个顶点或边属性,以及合并顶点属性。属性存放在一个 cell 中,因此会重写整个属性列。
  • 按名称查询 schema,以及仅按标签查询顶点或边。这两者都需要 HBase 二级索引。
  • 按标签删除边。
  • 带 in 条件、contains 条件或 contains_key 条件的查询。
  • 聚合属性和 OLAP 属性。
  • 原生数值类型(后端特性 supportsNumberType 为关闭状态)。
  • scan token。
  • hbase.enable_partition 为 true 时的前缀扫描和范围扫描。
  • 除 count 以外的聚合函数,其它聚合函数会被拒绝。
  • 快照。创建或恢复后端快照会抛出 UnsupportedOperationException。

已支持的特性包括顶点和边的 TTL、分页查询、order by 查询、范围条件,以及按输入 id 排序。

继续后将加载第三方服务 Kapa,并将你的问题发送给 Kapa 生成回答。

隐私政策