Skip to content

This is the multi-page printable view of this section. .

Return to the regular view of this page.

HugeGraph ToolChain

HugeGraph Toolchain includes the Java and Go clients, Loader, Hubble, Tools, Spark Connector, and SeaTunnel Sink/Source. Choose an entry by the task you need to complete, then open the component guide for its configuration and commands.

TaskStart hereBest for
Visualize graphsHubbleViewing and managing graphs in a Web UI
Import graph dataLoader, SeaTunnel Sink, Spark ConnectorImporting data directly or connecting an existing pipeline
Export or migrate graph dataTools, SeaTunnel SourceBackup, export, cross-graph migration, and continuous reads

Testing Guide: For running toolchain tests locally, please refer to HugeGraph Toolchain Local Testing Guide

DeepWiki provides real-time updated project documentation with more comprehensive and accurate content, suitable for quickly understanding the latest project information.

📖 https://deepwiki.com/apache/hugegraph-toolchain

Source repository: apache/hugegraph-toolchain

1 - Graph visualization

Hubble provides a Web interface for HugeGraph data, schema, queries, and imports. Start with a standalone deployment without PD, then read the distributed differences when needed.

1.1 - Manage an HStore Cluster with Hubble

Connect HStore and PD to Hubble and understand GraphSpaces, Schema templates, cluster topology, and node metrics.

This guide covers the differences between HStore + PD and standalone RocksDB. For modeling, importing data, and querying, see the Hubble standalone guide. This guide follows Toolchain master.

The main repository’s docker/docker-compose-hstore.yml already combines PD, Store, Server, and Hubble. Follow the adjacent Docker README to prepare .env and generated Hubble local configuration, then start it from docker/; do not maintain another deployment YAML. The settings below explain connection differences and do not replace the README’s PD credential and service-readiness requirements.

The screenshots use Hubble built from Toolchain 1.8.0 with Server, PD, and Store 1.7.0. Hubble currently returns the static value 3.0.0 from /about, which does not identify its build version. This pairing describes the screenshot environment; latest is mutable. Metrics and permission APIs vary by version.

Connect to a distributed cluster

Hubble still manages graph data through the Server graph API. In distributed mode, it discovers Servers through PD and collects cluster information from PD and Store. Hubble does not read or write graph data directly in Store.

For the official Compose deployment, prepare .env in the main repository’s docker/ directory following its README, then generate the Hubble configuration:

set -a
. ./.env
set +a
./set-hubble-pd-password.sh hstore

The script reads the loaded HG_PD_AUTH_SECRET_KEY and writes the PD operations password to the host file docker/conf/hubble/hstore.local.properties. Edit this generated file to customize connection or operations settings, preserving operations.pd.password so it matches PD’s secret. Do not replace it with the tracked .example template. Compose mounts it read-only at /hubble/conf/hugegraph-hubble.properties inside the container; do not edit it there. Do not commit the generated file or .env. Running the script again overwrites the generated file, so reapply custom settings afterward.

Start the services from the same docker/ directory and confirm Server registration and Store readiness:

docker compose -f docker-compose-hstore.yml up -d --wait

For separately deployed services, follow the PD deployment guide and HStore deployment guide. A source or binary Hubble deployment instead uses the package’s conf/hugegraph-hubble.properties. These settings match the official minimal topology; use backend-reachable addresses for other deployments:

pd.enabled=true
cluster=hg
pd.peers=pd:8686
pd.server=pd:8620
SettingPurposeBundled value
pd.enabledExplicitly enable PD mode; server.direct_url is not used in this mode.false
clusterCluster name used for Server discovery; it must match the registration.hg
pd.peersPD gRPC addresses, separated by commas.127.0.0.1:8686
pd.serverPD REST address for cluster operations, not a list of gRPC peers.127.0.0.1:8620

Do not interchange ports 8686 and 8620, or retain 127.0.0.1 for connections between containers. The bundled file explicitly sets pd.enabled=false, while the Java fallback for a missing key is true; set it explicitly in either deployment. Restart source or binary Hubble deployments after configuration changes. For Compose, recreate the Hubble container from docker/ after editing or regenerating the host file so the read-only bind mount loads it again; retain the original project name and all -f arguments:

docker compose -f docker-compose-hstore.yml up -d --force-recreate hubble

You do not enter a Server host and port for each graph in the UI.

Organize graphs and permissions with GraphSpaces

A GraphSpace groups graphs, Schema templates, and access permissions. For example, create sales and research spaces so each team can work with its own graphs. Modeling, importing, and querying within a space follow the standalone guide.

Create and adjust a GraphSpace

With Server authentication enabled, creating, editing, and deleting GraphSpaces require super administrator access. Start with a globally unique GraphSpace name, an optional display alias, and the maximum graph count. For example, use research as the API identifier and “Research_graph” as the display alias. The name cannot change after creation. The alias and description can change; aliases do not participate in URLs or permission matching. Choose GraphSpace administrators from existing accounts.

“Advanced deployment and resource limits” includes CPU and memory limits for graph query/write services and asynchronous compute tasks, plus the storage capacity limit. These are deployment and quota settings, not current usage or a promise to resize Docker containers when the form is saved. Defaults are 100 graphs, 64 CPU cores and 128 GB of memory for each of the graph and compute services, and 1000000 GB of storage. Keep the defaults for a container trial. Configure Kubernetes namespaces, Operator images, and algorithm images only for the corresponding deployment or compute use.

Select a space before opening a graph. Check the current graph after switching spaces, especially when spaces contain identically named graphs. The list reflects account access. If a space is missing, check membership permissions before creating more graphs.

GraphSpace creation and advanced resource limits

Reuse Schema templates

User-defined Schema templates require PD mode and persist reusable Groovy Schema within a GraphSpace. When several business graphs share vertex labels, edge labels, and indexes, save the model as a template and select it when creating subsequent graphs. For example, a user template in research can be reused for new graphs in that space. After switching spaces, select a template belonging to the new space.

User templates differ from the built-in sample templates in the main guide: built-in templates help explore preset models, while user templates preserve your own models for future use. A template is not an import of data; loading sample data is a separate graph creation choice. With Server authentication enabled, creating a user template requires write access to its space. Updating or deleting it also requires ownership or the corresponding administrative permission. Anonymous mode does not enforce per-user permissions or template ownership checks. Standalone mode does not provide user template management.

Assign access to a GraphSpace

For account creation, login, and personal details, see the standalone guide. Distributed mode adds “Manage GraphSpace members” to assign permissions between existing accounts and selected spaces. Servers supporting permission presets provide these common choices:

PresetIntended userScope
GraphSpace read-onlyUsers who inspect and query graph dataSelected spaces.
GraphSpace read-writeUsers who model, import, and maintain graph dataSelected spaces.
GraphSpace administratorOwners managing a space’s members and graph resourcesSelected spaces; does not grant GraphSpace creation/editing or cluster operations access.
Super administratorOperators managing accounts, GraphSpaces, and the clusterGlobal; assign according to responsibilities.

An account may have different access in different spaces, such as read-write in research and read-only in sales. Accounts and space memberships are managed separately; removing a member does not delete the global account. After creating an account, continue directly to assigning space access. When changing an existing membership, check the selected space and preset and preserve permissions still needed in other spaces. Enabling pd.enabled grants no permissions; legacy custom roles on older Servers are not interchangeable with these presets. See Server authentication and authorization.

GraphSpace members and access presets

Locate problems through the cluster overview

The overview presents Server → PD → Store in one topology. Switch to the node list to filter by type, status, or name. PD Leader, online Store count, graphs, partitions, replicas, and data size help identify cluster scale and nodes requiring attention. Partitions and replicas describe data distribution. Data size is observed usage, distinct from a GraphSpace’s configured storage limit.

Start with cluster and source status, then inspect an affected node. UP indicates a successful collection from the source; DEGRADED indicates a cluster or partial source issue, and DOWN indicates that the corresponding probe failed. Graphs may remain queryable when topology is available but some metrics are missing. The UI distinguishes unsupported, unavailable, stale, and failed collections and includes observation times. An empty value is not zero, and an older value is not a current observation.

Cluster topology, node status, and capacity

Read node details

NodeMain observationsUse
ServerAvailability, JVM/CPU/memory, and backend metricsCheck the query/write entrypoint and resource pressure.
PDLeader/Follower role, status, and runtime metrics supplied upstreamCheck the metadata/scheduling entrypoint; do not infer a role when Leader information is absent.
StorePartition and Leader partition counts, system/disk metrics, Raft group and enabled group countsInspect storage nodes and replica service and compare distribution across nodes.
Store system, drive, Raft, and partition metrics

The PD Leader and Store Leader partitions serve different purposes: PD coordinates cluster metadata, while a Store leads the respective data partition’s Raft group. Hubble displays upstream observations and does not replace a full Raft replica consistency check. Available metrics vary by component version.

With Server authentication enabled, cluster operations require a super administrator (ADMIN level). GraphSpace administrators and ordinary members do not inherit cluster access. Use container isolation, a trusted HTTPS entrypoint, Server authentication, and network allowlists; avoid publishing Hubble or component ports directly.

Configure operations access

Hubble’s backend uses the PD/Store operations credentials, separately from the Server account used to log in through the browser. For Compose, edit the generated host file docker/conf/hubble/hstore.local.properties; other deployments use the Hubble package configuration. The script-generated PD password must match the secret in .env. Configure the Store service account when Store authentication is enabled. Do not include passwords in documentation, screenshots, or committed configuration:

SettingBundled defaultConfiguration
operations.pd.username / operations.pd.passwordUsername hubble, empty passwordMatch PD operations REST authentication.
operations.store.username / operations.store.passwordUsername hubble, empty passwordSet the service account when upstream Store REST authentication is enabled.
operations.store.allowed_targets[http://127.0.0.1:8520,http://[::1]:8520]List trusted Store metric origins.

For example, when a containerized Store advertises store:8520, set:

operations.store.allowed_targets=[http://store:8520]

For multiple nodes, list each trusted origin. Every entry must use http or https with an explicit port and no path, credentials, or wildcard, and must match the Store metric target returned by PD. Adding an origin to this list does not register or discover a node.

If topology is available but metrics are incomplete, check backend connectivity and authentication to PD REST, the Store REST/metric targets returned by PD, and their match with the allowlist. A partially available overview means some sources could not be collected; it does not imply that all graph APIs are unavailable.

2 - Graph Visualization with Hubble: Standalone Quick Start

Visualize graph data with Hubble and RocksDB Server: start with Docker, explore schema, import CSV, and run Gremlin queries.

Hubble is the HugeGraph Web management and graph visualization interface. Use one workspace to manage schema, import data, run queries, and switch between graph, table, and JSON results. This guide uses standalone RocksDB Server + Hubble, without PD or Store.

For HStore, read the shared operations here first, then follow the distributed supplement. This guide follows Toolchain master (currently 1.8.0); Docker latest is mutable, so check the actual running versions.

Warning

Do not expose Hubble or Server directly to the public network. In production, use HTTPS, containers, authentication and authorization, and an access allowlist.

Start the standalone pair

Use the main repository’s docker/docker-compose.yml instead of writing another Compose file. It already combines RocksDB Server and Hubble, with networking, health checks, and data volumes. See the adjacent README for deployment details.

git clone --branch master --single-branch --depth 1 https://github.com/apache/hugegraph.git
cd hugegraph/docker

If you already have the main repository, enter its docker/ directory. Compose mounts conf/hubble/standalone.properties from that directory. It sets pd.enabled=false and server.direct_url=http://server:8080; both services communicate over one Docker network, without a Server address configured per graph. Do not download only the YAML and start it from another directory: relative configuration files may be missing.

Hubble defaults to host loopback port 8088; Server publishes 8080. For a trial on your machine only, change Server’s ports entry to 127.0.0.1:8080:8080 to avoid exposing the anonymous API to other machines.

Choose an unused project name for this trial and keep these variables in the same terminal. Restore this project name if you use another terminal.

export HUGEGRAPH_VERSION=latest
export HUBBLE_IMAGE=hugegraph/hubble:latest
export HUBBLE_DEMO_PROJECT="hubble-demo-$(date +%Y%m%d-%H%M%S)"
docker compose ls
docker compose -p "$HUBBLE_DEMO_PROJECT" -f docker-compose.yml pull
docker compose -p "$HUBBLE_DEMO_PROJECT" -f docker-compose.yml up -d --wait
docker compose -p "$HUBBLE_DEMO_PROJECT" -f docker-compose.yml ps
curl -fsS http://127.0.0.1:8080/versions

Once services are healthy, open http://127.0.0.1:8088. In a fresh directory without HUGEGRAPH_ADMIN_PASSWORD, Server allows anonymous access and Hubble opens the home page directly. For authentication, follow the Docker README to configure an administrator password and JWT secret in .env, then sign in with a Server account. Hubble has no separate account database. Personal and account-management pages depend on the authentication mode and your permissions. Do not overwrite an existing .env.

Use latest to try current features, and pin a published image version or digest for production. Images are convenience distributions; official release archives are on the download page. The algorithm and account screenshots use Hubble built from Toolchain 1.8.0 with Server 1.7.0. Hubble currently returns the static value 3.0.0 from /about, which does not identify the build version. This pairing describes those screenshots, not a fixed meaning of latest. Available controls depend on Server capabilities. Compose’s server-data and hubble-data retain graph data and Hubble metadata respectively. For further persistence and production settings, see the Server deployment guide.

Home: find the right starting point

Home groups the workspace into graph overview, data preparation, and graph queries. Use it to understand the workflow, then return to any section through the sidebar. The top graph selector determines the target of queries, schema operations, and async tasks; check the current graph after changing pages. Standalone mode has only the DEFAULT GraphSpace and needs no PD configuration.

SectionPurpose
Graph Overview and detailsSelect a graph, load samples, inspect its size, then model or query it
Schema configurationDefine properties, vertex/edge labels, and indexes
GQL TraversalWrite queries and explore graph, table, and JSON results
Built-in AlgorithmsExplore neighbors, paths, and similarity with parameter forms
Async TasksTrack background queries, schema changes, and index operations
Data Source ManagementUpload files or configure external readers
Data ImportMap source fields to the graph model and run or schedule ingestion
Profile and Account ManagementUpdate personal details/passwords and manage accounts or space members according to permissions
OperationsInspect Server nodes; PD mode also includes cluster overview and PD/Store nodes

Graph Overview and details: get to know a graph

Select the default graph hugegraph in Graph Overview. The overview provides graph entry points and action menus. Graph details show schema and data statistics, with routes to modeling, data preparation, and queries. Statistics describe overall size; update them after importing or modifying data.

Load the People & Software Demo Graph from the graph’s More actions menu. Samples add their schema and missing elements without clearing existing data. Start with an empty graph to avoid conflicting schema names. The remaining examples explore the people and software in this graph.

Graph overview with the person/software sample

Graph creation depends on Server capabilities. Its form accepts a name, optional alias, and schema or sample; it does not configure a Server host or account per graph. The connection comes from Hubble configuration. User-defined schema templates require PD mode; see the distributed supplement.

Schema modeling: define the shape of your data

Schema determines valid properties and relationships, vertex ID generation, and query indexes. Open the graph’s schema configuration. List view is useful for maintaining definitions; graph view helps explain how labels connect.

DefinitionWhat to decide
PropertiesData type and cardinality; distinguish numbers from text
Vertex labelsProperties, nullable properties, ID strategy, and primary keys
Edge labelsSource/target labels, properties, frequency, and sort keys
Vertex / edge indexesIndex type and fields that match filtering and range queries

In the person/software sample, person generates IDs from the primary key name, with nullable age and city. software has custom numeric IDs, and created connects people to software. This matches the Loader example.

Vertex labels with primary-key and numeric ID strategies

For a new model, define properties, vertex labels, edge labels, then indexes. Associated-property and index information help inspect dependencies. Schema deletion and index creation/rebuild may submit background tasks. Acceptance of an operation is only the first step; confirm its final status in Async Tasks.

GQL workspace: query and explore relationships

Open GQL Traversal and confirm hugegraph is selected. The workspace puts the editor and results together, with immediate or async execution, query favorites, and reusable execution history. Start with this Gremlin query:

g.V().hasLabel('person').valueMap()

Inspect person properties in table or JSON view. To visualize the people-to-software relationships, run:

g.V().hasLabel('person').outE('created').inV().path()
Gremlin path query and graph result

Graph results support 2D / 3D. Click a vertex or edge to inspect its ID, label, and properties; double-click a vertex to expand its neighbors. Layout, styling, and filtering help highlight relevant relationships, while export helps share results. Use New to create elements, or edit existing data when authorized. Layout, colors, and display limits only change presentation; adding/editing elements and Gremlin writes change Server data.

Use Ctrl / Command + Enter to execute. Immediate mode suits small explorations; submit long queries asynchronously and avoid returning an entire large graph. Cypher is available only when Server supports it. Text2GQL is currently a UI preview with no model or query service connected; it cannot generate executable queries.

Built-in algorithms: explore with parameter forms

Use Built-in Algorithms when you prefer a form to writing traversal code. Search for an algorithm and supply its parameters. Neighbor exploration answers what surrounds a vertex, path algorithms connect two vertices, and similarity/ranking algorithms compare or select vertices. Start with a known vertex ID and limit direction, edge labels, depth, and result size before attempting broader computation.

Forms provide parameter guidance and documentation links, and restore common parameters when navigating away and back. Results use graph or algorithm-specific panels. Consult the help and documentation links beside the algorithm title for definitions and parameters. While the parameter form is focused, Ctrl / Command + Enter runs the current algorithm.

OLAP batch algorithms require external compute services such as Computer or Vermeer. The two containers in this example provide online graph operations, not those compute services.

For example, select K-neighbor (GET) with source=1:marko, max_depth=1, and limit=20. After parameter validation, use the run button on the card or the form shortcut to explore one-hop neighbors.

Built-in neighbor algorithm parameters and graph result

Async Tasks: confirm background results

Async Tasks lists background work for the current graph, including async queries and some schema/index operations. Filter by task type and status, inspect IDs, creation times, and execution states, open successful query results, or expand failure information. Completed task records can be deleted where the interface permits. These tasks are separate from import execution history.

For example, submit g.V().count() asynchronously, confirm success in the list, then inspect the returned count. A successful submission means the request was accepted, not that computation or indexing has finished. Use task errors and Server logs together when diagnosing failures.

Data Source Management: prepare the input

A data source defines where data comes from and how to parse it, and can be referenced by import tasks. Hubble supports FILE, HDFS, JDBC, and Kafka, each with its own path, connection, or subscription settings. FILE is an easy starting point: upload a file, configure its format, delimiter, encoding, and header, and check column names before mapping. Hubble configuration controls upload limits and permitted extensions.

Save this UTF-8 file as people.csv for the person example:

name,age,city
docs_alice,28,Beijing
docs_bob,32,Shanghai

Create a FILE data source and upload it. Select CSV (comma separation and UTF-8 by default), with column names name,age,city. Header, delimiter, and encoding belong to the data source, not mapping settings. The source fields must match the file.

Data Import: turn fields into a queryable graph

Data Import converts source rows into vertices and edges. Its four configuration sections identify the target, select fields, map them to the graph, and choose execution timing. Use the preceding data source to create a person import into DEFAULT / hugegraph:

SectionSettings for this example
Basic InformationTarget graph, new data source, and a recognizable task name
Source FieldsSelect name, age, and city, moving them to the selected field list
Mapping FieldsAdd a person vertex mapping; use Auto Match for same-name properties, then verify types
ScheduleChoose one-time execution; confirmation submits the task immediately

person uses PRIMARY_KEY, so do not select a separate ID column. Custom ID strategies require an ID column; AUTOMATIC lets Server generate IDs, while PRIMARY_KEY derives them from mapped primary-key properties. Edge mappings need source/target fields that follow the corresponding vertex ID rules.

The task list manages configuration and execution entry points; execution history in task details shows each instance’s state, count, and errors. Periodic schedules and real-time Kafka tasks are also available; choose an execution mode compatible with the source. After completion, verify the result in the GQL workspace:

g.V().hasLabel('person').has('name', within('docs_alice', 'docs_bob')).valueMap()

The result should include both new people. Import counts are not necessarily counts of newly created vertices: reruns may update existing elements, and header processing can affect reader counts. If the import fails, check source fields, numeric types, nullable properties, and target schema. Use Hubble for small trials and HugeGraph Loader for production bulk ingestion.

Profile and account permissions

Hubble uses Server authentication and accounts, with no separate user database. Anonymous mode hides Profile and Account Management. After authentication is enabled, Profile shows the current account’s details and permissions and allows changing your password. Editing details such as a nickname requires Server support for the personal-profile API. Changing a password ends the current session; sign in again with the new password.

Standalone account management

In standalone mode, the administrator can create, inspect, edit, delete, and batch-create accounts. Ordinary standalone accounts created through Hubble receive read, write, delete, and execute permissions across all graphs; they are not read-only or isolated to one graph. Configure finer resource permissions through Server authentication and authorization rather than relying on GraphSpace presets.

GraphSpace permissions in PD mode

With a Server supporting default-role APIs, global accounts and space access can be managed separately. The administrator manages global accounts; a space administrator manages members only within authorized spaces. Ordinary members do not receive account-management or operations entry points. The following presets are for PD mode, and are not universally editable on older or standalone Servers:

PresetScope and purpose
SUPER_ADMINGlobal account, GraphSpace, and operations management; grant or revoke super-administrator access
GS_ADMINManage authorized spaces and their members, without granting other-space or global super-administrator access
GS_READ_WRITERead and write graph data within authorized spaces, without managing global accounts
GS_READ_ONLYRead graph data within authorized spaces, without writes

An account may have different permissions in different spaces. Select a space before adding an existing account or changing member permissions. Before replacing custom permissions with a preset, inspect the grants that need to be retained; complex permissions may not match a single preset. After changes, refresh permission context or sign in again and verify menus and space selection. Server still validates every request. Older Servers may hide or disable unsupported operations; a visible button alone does not establish resource authorization. See the distributed supplement for space management.

Authenticated account list in PD mode

Operations: inspect the standalone Server

Standalone Operations provides Node Information for Server only, with no PD/Store nodes or cluster overview. Search nodes, filter health status, and open node details to inspect available version, system, JVM, and Server-backend metrics. When metrics are unavailable or stale, use collection state and the last successful observation time; a missing value is neither zero nor proof of health.

Anonymous mode can read operations information. With authentication enabled, operations are available only to administrators with the capability. This is an observation and diagnosis interface, not a start/stop or scaling console. The distributed supplement covers cluster overview and the PD/Store node hierarchy.

Keyboard shortcuts and graph interactions

Use the topbar shortcut-help button to see key bindings. Their scope differs: typing ? in an input does not trigger global help.

ActionKey or gestureScope
Open / close shortcut help?Outside inputs and editors
Execute queryCtrl / Command + EnterQuery editor
Run current algorithmCtrl / Command + EnterAlgorithm parameter form
Toggle graph fullscreenFClick to focus the graph canvas first; not a global binding
Inspect element detailsClick a vertex or edgeGraph result
Expand neighboring relationshipsDouble-click a vertexGraph result

Diagnose connection and result issues

SymptomCheck first
The Hubble page does not openCheck container status and docker compose -p "$HUBBLE_DEMO_PROJECT" -f docker-compose.yml logs hubble
The page opens but graphs are unavailableServer health, a server.direct_url reachable from Hubble, and a shared network
Login appears or write actions are missingServer authentication and account permissions; Hubble has no independent authentication switch
Expected data is missingCurrent graph, sample load result, matching labels/properties; distinguish canvas display from stored data
No cluster overviewThis example has no PD; see the distributed supplement

Configuration can affect displayed query size. gremlin.suffix_limit defaults to 250 and supplies .limit(N) appended to applicable Gremlin queries; it is not a universal hard limit. gremlin.vertex_degree_limit (100) and gremlin.edges_total_limit (500) constrain expansion. FILE uploads allow csv,txt by default, with 1 GB per file and 10 GB total. Override upload_file.* settings when needed.

Stop the trial or build from source

When finished, run this in the main repository’s docker/ directory:

docker compose -p "${HUBBLE_DEMO_PROJECT:?}" -f docker-compose.yml down --volumes

This removes the project’s containers, network, named volumes, and anonymous volume, losing the sample data and Hubble import tasks. To retain data, omit --volumes when taking the deployment down and reuse the same project name when starting it again.

To obtain an exact master build, use JDK 11 and Maven. The Maven plugin installs the required Node/Yarn; you do not need to install them separately. These commands skip tests:

git clone --branch master --single-branch https://github.com/apache/hugegraph-toolchain.git
cd hugegraph-toolchain
mvn install -pl hugegraph-client,hugegraph-loader -am -Dmaven.javadoc.skip=true -DskipTests -ntp
cd hugegraph-hubble
mvn package -Dmaven.javadoc.skip=true -DskipTests -ntp
cd apache-hugegraph-hubble-*
# Edit conf/hugegraph-hubble.properties with the correct Server URL
bin/start-hubble.sh

The bundled configuration binds to localhost:8088. bin/stop-hubble.sh requests graceful shutdown before forcing termination on timeout. For development and testing, see the Toolchain local test guide.

3 - Graph import

Choose an import tool when you need to write file, database, or message data into HugeGraph. Use Loader for a direct import, SeaTunnel Sink when an existing Source, Transform, and Sink pipeline should be reused, or Spark Connector from a Spark job.

3.1 - Import Graph Data with SeaTunnel Sink

SeaTunnel connects data sources such as databases and Kafka to HugeGraph. The connector has two parts: Source reads data and Sink writes data[1][2], with SeaTunnel transform components available between them. To export or migrate data from HugeGraph, see the SeaTunnel Source export and migration guide.

Version requirement: This guide targets SeaTunnel 3.0+. All examples use the mappings configuration available in SeaTunnel 3.0+.

Loader imports data directly with graph mappings; SeaTunnel 3.0+ combines Source, Transform, and Sink, and both support JDBC, Kafka, and graph data

Click a diagram to view the original size.

1 Loader, Tools, and SeaTunnel

HugeGraph-Loader is suited to direct imports from common data sources. HugeGraph-Tools focuses on standalone graph management, backup, and export. SeaTunnel organizes a job as Source → Transform → Sink, so you can reuse existing connectors, transforms, and data pipelines.

Table legend

✅ Supported natively; ⚠️ conditional support or requires an extra component/external platform; ❌ not provided

ComparisonLoaderToolsSeaTunnel
Task coverage✅ Direct graph imports✅ Backup, restore, and export✅ Import, export, and migration with composable Source, Transform, and Sink stages
Job configurationJSON mapping file describing the source, vertices, and edgesCommand-line options and operationsHOCON job file[3] combining Source, Transform, and Sink
Default deployment✅ Standalone CLI; ⚠️ Spark Loader can extend it✅ Standalone CLI✅ Standalone; ✅ distributed
Execution engine⚠️ Mainly CLI; Spark Loader is a separate extension❌ Does not provide a Spark/Flink execution engine✅ HugeGraph Source and Sink support Zeta, Spark, and Flink[1][2][7][8][9][10]
Frontend and observability❌ No built-in frontend; inspect CLI logs❌ No built-in frontend; inspect CLI logs✅ Built-in Web UI job panel for task status and runtime information
Input and output⚠️ Focused on graph imports and common files, JDBC, Kafka, and similar sources⚠️ Focused on graph data and backup files in common storage✅ Dozens of connectors, including JDBC, Kafka, and SQL-CDC
Scheduling and resource management❌ No unified cross-task scheduling or resource allocation❌ No unified cross-task scheduling or resource allocation⚠️ Can integrate with DolphinScheduler for scheduling and task management
Simplicity✅ Focused and simple; a future binary CLI will make quick use easier✅ Direct commands for standalone operations⚠️ More runtime components, suited to long-lived data pipelines
High-throughput import✅ Supports bypass-server and other optimizations; measured peaks can reach 1-2 million records/s with specific backends and hardware, so benchmark the actual setup⚠️ Focuses on backup and export rather than bulk-import throughput✅ Scales throughput through parallelism, distributed engines, and connectors

SeaTunnel covers Loader’s graph-import and Tools’ export and migration scenarios in one expandable pipeline, and it also supports SQL-CDC and dozens of input and output types. Loader and Tools normally run on one machine, while SeaTunnel supports both standalone and distributed deployments and scales with data and task volume. Tools’ schedule-backup can create a crontab entry, but it does not provide unified workflow orchestration and resource management.

Existing Spark/Flink daily jobs

Both HugeGraph Source and Sink list SeaTunnel Engine (Zeta), Spark, and Flink as supported engines in SeaTunnel 3.0+. If you express the daily job as a SeaTunnel job and submit it to that engine, records can move directly from Source to Transform to Sink without an intermediate file. If you keep the existing Spark/Flink DAG, SeaTunnel does not automatically take over its in-memory DataFrame or stream. Adapt it into a SeaTunnel job or expose the data through a Source connector

Loader and Tools are focused, direct, and quick to start. Use Loader for a direct graph import; use Tools for backup, restore, export, or daily operations. If a SeaTunnel job already exists, adding HugeGraph to that pipeline is usually simpler. For higher import throughput, Loader’s bypass-server path and other import optimizations are a better fit; measured peaks of 1-2 million records/s require a specific backend, data set, and hardware configuration and are not a general performance guarantee. For new SeaTunnel jobs, use 3.0+ and mappings. Recheck the connector configuration when using another version.

2 Prepare the environment

2.1 Get SeaTunnel 3.0+

Use JDK 11 and set JAVA_HOME; the official deployment guide lists Java 8 and 11 as prerequisites. Choose a 3.0+ binary distribution from the official release page[4] and verify its checksum or signature using the files linked there. The commands below use 3.0.0 as an example; set SEATUNNEL_VERSION to the release you downloaded.

export SEATUNNEL_VERSION=3.0.0
tar -xzf "apache-seatunnel-${SEATUNNEL_VERSION}-bin.tar.gz"
cd "apache-seatunnel-${SEATUNNEL_VERSION}"

The binary distribution does not include connector plugins. In config/plugin_config, select the connectors needed by the examples:

--seatunnel-connectors--
connector-hugegraph
connector-jdbc
connector-kafka
--end--

Install the matching released plugins using the upstream plugin installation instructions[5]:

sh bin/install-plugin.sh "${SEATUNNEL_VERSION}"

Run the remaining commands from this installation directory and keep the engine and connector plugins at the same version. This guide uses the bundled Zeta engine in local mode[6][7]. Check that connectors/ contains HugeGraph, JDBC, and Kafka as needed[11][12]. The JDBC examples also require the MySQL driver JAR in lib/, with driver class com.mysql.cj.jdbc.Driver.

2.2 Prepare HugeGraph and data sources

Start HugeGraph Server and create a graph for testing. The examples use the hugegraph graph in the DEFAULT graph space. Adjust these names to match the server configuration; graph space names are case-sensitive. If authentication is enabled, provide username and password in the HugeGraph Source and Sink configurations.

The following graph model is shared by the JDBC and Kafka examples. mappings creates missing PropertyKey, VertexLabel, and EdgeLabel definitions by default; existing schema definitions must be compatible.

Graph elementName and properties
Propertiesname is Text; age and since are Int
Vertexperson, primary key name, properties name and age
Edgeknows, from person to person, property since

The mysql, kafka, and hugegraph host names in the examples are placeholders. Replace them with addresses reachable from the SeaTunnel runtime. Inside a container, 127.0.0.1 points to that container; services on the same Docker network can use their service names. Set host to a host name or IP address, and set the port separately.

3 Import from a relational database (sql2graph)

Use two jobs for this import: write the person table as vertices first, then write the knows table as edges. Both edge endpoints will already exist when the edge job runs.

The person table creates marko and vadas vertices; the knows table creates a directed edge with since 2010 through endpoint fields

3.1 Import vertices

Prepare the sample data in the MySQL demo database and grant the configured account read access:

CREATE TABLE person (
  name VARCHAR(64) PRIMARY KEY,
  age INT NOT NULL
);
INSERT INTO person VALUES ('marko', 29), ('vadas', 27);

Save the following as config/sql2graph-person.conf and replace the database user name and password:

env {
  job.mode = "BATCH"
}

source {
  Jdbc {
    url = "jdbc:mysql://mysql:3306/demo?useSSL=false&serverTimezone=UTC"
    driver = "com.mysql.cj.jdbc.Driver"
    username = "seatunnel"
    password = "change_me"
    query = "SELECT name, age FROM person ORDER BY name"
  }
}

sink {
  HugeGraph {
    host = "hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    batch_failure_fallback = false
    mappings = [
      {
        type = "VERTEX"
        label = "person"
        idStrategy = "PRIMARY_KEY"
        idFields = ["name"]
        properties = ["name", "age"]
      }
    ]
  }
}
./bin/seatunnel.sh --config ./config/sql2graph-person.conf -m local

Check the result in Hubble or Gremlin. You should find marko and vadas with their ages:

g.V().hasLabel('person').valueMap('name', 'age')

idFields = ["name"] uses the name to generate the primary key. Importing the same name again writes to the same vertex. properties lists the source fields to write.

3.2 Import edges

Prepare the relation table. Its two endpoint fields correspond to person.name from the vertex job:

CREATE TABLE knows (
  source_name VARCHAR(64) NOT NULL,
  target_name VARCHAR(64) NOT NULL,
  since INT NOT NULL
);
INSERT INTO knows VALUES ('marko', 'vadas', 2010);
Expand the configuration and save it as config/sql2graph-knows.conf
env {
  job.mode = "BATCH"
}

source {
  Jdbc {
    url = "jdbc:mysql://mysql:3306/demo?useSSL=false&serverTimezone=UTC"
    driver = "com.mysql.cj.jdbc.Driver"
    username = "seatunnel"
    password = "change_me"
    query = "SELECT source_name, target_name, since FROM knows ORDER BY source_name, target_name"
  }
}

sink {
  HugeGraph {
    host = "hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    batch_failure_fallback = false
    check_vertex = true
    mappings = [
      {
        type = "EDGE"
        label = "knows"
        sourceConfig = {
          label = "person"
          idFields = ["source_name"]
        }
        targetConfig = {
          label = "person"
          idFields = ["target_name"]
        }
        fieldMapping = {
          source_name = "name"
          target_name = "name"
        }
        properties = ["since"]
      }
    ]
  }
}

After the vertex job succeeds, run the edge job:

./bin/seatunnel.sh --config ./config/sql2graph-knows.conf -m local

The following query should return a knows edge from marko to vadas with since set to 2010:

g.V().has('person', 'name', 'marko').outE('knows').where(inV().has('name', 'vadas')).valueMap()

sourceConfig and targetConfig identify the endpoint fields. fieldMapping maps them to the vertex primary key name, and properties = ["since"] writes only the edge property. The example enables check_vertex = true and disables per-record fallback after a batch failure (batch_failure_fallback = false), so a missing endpoint or write failure causes the job to fail.

If the relation table has only numeric foreign keys while the graph uses names as primary keys, join the names in SQL before passing the records to the Sink. See MySQL CDC Source[13] for MySQL CDC integration.

4 Import from Kafka (kafka2graph)

Kafka is useful for a continuous stream of events. Create the user-events topic and publish the following JSON message. Each message becomes one person vertex:

{"name":"marko","age":29}

Save the following as config/kafka2graph.conf:

env {
  job.mode = "STREAMING"
  checkpoint.interval = 10000
  sink.flush.interval = 5000
}

source {
  Kafka {
    bootstrap.servers = "kafka:9092"
    topic = "user-events"
    consumer.group = "hugegraph-import"
    start_mode = "earliest"
    format = "json"
    schema = {
      fields = {
        name = "string"
        age = "int"
      }
    }
  }
}

sink {
  HugeGraph {
    host = "hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    batch_failure_fallback = false
    mappings = [
      {
        type = "VERTEX"
        label = "person"
        idStrategy = "PRIMARY_KEY"
        idFields = ["name"]
        properties = ["name", "age"]
      }
    ]
  }
}
./bin/seatunnel.sh --config ./config/kafka2graph.conf -m local

Use the Gremlin query from section 3.1 to check the data. The streaming job keeps running. checkpoint.interval saves job state every 10 seconds, while sink.flush.interval asks Zeta to flush every 5 seconds so a small number of messages does not wait for a full batch.

HugeGraph Sink writes with at-least-once semantics, so recovery can replay records. PRIMARY_KEY sends the same name to the same vertex, but it does not make every update exactly-once. Scheduled flushing is provided by Zeta and does not apply to Spark or Flink engines.

5 Common configuration and troubleshooting

The following table applies to the SeaTunnel 3.0+ version used by this guide:

ConfigurationPurpose
host, portSet the HugeGraph host and port
graph_name, graph_spaceSelect an existing graph and graph space
mappingsDefine how input fields become vertices or edges
propertiesList the source fields written by each mapping
schema_save_modemappings creates missing schema by default; existing schema must still be compatible
batch_sizeNumber of records per batch; default 500
env.sink.flush.intervalZeta scheduled flush interval in milliseconds
check_vertexCheck edge endpoints; the edge job in this guide sets it to true
batch_failure_fallback[2]Defaults to true, so a failed batch falls back to record-by-record retries, capped by max_insert_errors; the examples explicitly set false so a batch failure stops the job
max_insert_errorsNumber of failed records that record-by-record fallback may skip; default 500, -1 for unlimited, and only applies when batch_failure_fallback is enabled

Use these checks when a job fails:

  • mappings is unknown or HugeGraph Source is missing: Check that the engine and HugeGraph connector come from the same SeaTunnel 3.0+ build.
  • Connection failure: Check the host, port, graph space, authentication details, and whether the SeaTunnel runtime can reach the service.
  • Schema incompatibility: Check the ID strategy, property types, and edge endpoints. Automatic creation does not change an existing PRIMARY_KEY label into CUSTOMIZE_STRING.
  • Small Kafka batches do not appear promptly: Confirm that the job uses Zeta and set sink.flush.interval in env. In this version, batch_interval_ms is retained only for compatibility and cannot replace it.

6 Choosing a tool

Choose a tool based on the work to complete. Use Tools for graph management, Gremlin, backup, or cloning. Use Loader for a direct graph import. Choose SeaTunnel when you need to reuse a Source, Transform, and Sink pipeline. For SeaTunnel graph reads and migrations, prepare the environment using the SeaTunnel 3.0+ version used by this guide.

Choosing a tool: Tools for graph management, Loader for direct imports, and SeaTunnel for reusable data pipelines

7 References

HugeGraph connectors

[1] HugeGraph Source
[2] HugeGraph Sink

Configuration and deployment

[3] HOCON job configuration
[4] SeaTunnel 3.0+ release download
[5] Connector plugin installation
[6] SeaTunnel local deployment

Execution engines

[7] SeaTunnel Engine Overview
[8] SeaTunnel Spark Engine
[9] SeaTunnel Flink Engine
[10] Connector V2 multi-engine support

Data source connectors

[11] JDBC Source
[12] Kafka Source
[13] MySQL CDC Source

4 - HugeGraph-Loader Quick Start

Bulk import graph data into HugeGraph with Loader from files, HDFS, relational databases, Kafka, and other HugeGraph graphs.

This guide follows Toolchain master (currently 1.8.0). Released packages and Docker latest may differ from source; check the version you use and build from source for unreleased functionality.

1 HugeGraph-Loader Overview

HugeGraph-Loader is the data import component of HugeGraph, which can convert data from various data sources into graph vertices and edges and import them into the graph database in batches.

Currently supported data sources include:

  • Local disk file or directory, supports TEXT, CSV and JSON format files, supports compressed files
  • HDFS file or directory supports compressed files
  • Mainstream relational databases, such as MySQL, PostgreSQL, Oracle, SQL Server
  • Kafka topic
  • An existing HugeGraph graph, used to copy data from one graph into another

Local disk files and HDFS files support resumable uploads.

It will be explained in detail below.

Note: HugeGraph-Loader requires HugeGraph Server service, please refer to HugeGraph-Server Quick Start to download and start Server

Testing Guide: For running HugeGraph-Loader tests locally, please refer to HugeGraph Toolchain Local Testing Guide

2 Get HugeGraph-Loader

HugeGraph-Loader is available in the following three ways:

  • Use docker image (Convenient for Test/Dev)
  • Download the compiled tarball
  • Clone source code then compile and install

2.1 Use Docker image (Convenient for Test/Dev)

We can deploy the loader service using docker run -itd --name loader hugegraph/loader:latest. For the data that needs to be loaded, it can be copied into the loader container either by mounting -v /path/to/data/file:/loader/file or by using docker cp.

Alternatively, to start the loader using docker-compose, the command is docker-compose up -d. An example of the docker-compose.yml is as follows:

This combination is for local anonymous testing; the Server API is published only on the local host.

version: '3'

services:
  server:
    image: hugegraph/hugegraph:latest
    container_name: server
    ports:
      - 127.0.0.1:8080:8080

  loader:
    image: hugegraph/loader:latest
    container_name: loader
    # mount your own data here
    # volumes:
      # - /path/to/data/file:/loader/file

The specific data loading process can be referenced under 4.5 User Docker to load data

Note:

  1. The docker image of hugegraph-loader is a convenience release to start hugegraph-loader quickly, but not official distribution artifacts. You can find more details from ASF Release Distribution Policy.

  2. Pin a published version tag or image digest in production. latest is mutable and does not guarantee alignment with master.

2.2 Download the compiled archive

Choose a published Toolchain archive from the download page and extract it. A release may not include all master functionality; use the source build below for the implementation described here.

2.3 Clone source code to compile and install

Clone the master branch:

git clone --branch master --single-branch https://github.com/apache/hugegraph-toolchain.git

Compile and generate tar package:

cd hugegraph-toolchain
mvn clean package -pl hugegraph-loader -am -DskipTests -ntp

For Oracle input, obtain a JDBC driver JAR compatible with your Oracle/JDK versions and place it in the extracted Loader lib/ directory. The startup script loads JARs there; installing a driver only in the local Maven repository does not put it on Loader’s runtime classpath.

3 How to use

The basic process of using HugeGraph-Loader is divided into the following steps:

  • Write graph schema
  • Prepare data files
  • Write input source map files
  • Execute command import

3.1 Construct graph schema

This step is the modeling process. Users need to have a clear idea of ​​their existing data and the graph model they want to create, and then write the schema to build the graph model.

For example, if you want to create a graph with two types of vertices and two types of edges, the vertices are “people” and “software”, the edges are “people know people” and “people create software”, and these vertices and edges have some attributes, For example, the vertex “person” has: “name”, “age” and other attributes, “Software” includes: “name”, “sale price” and other attributes; side “knowledge” includes: “date” attribute and so on.

Example graph with person and software vertices connected by knows and created edges

graph model example

After designing the graph model, we can use groovy to write the definition of schema and save it to a file, here named schema.groovy.

// Create some properties
schema.propertyKey("name").asText().ifNotExist().create();
schema.propertyKey("age").asInt().ifNotExist().create();
schema.propertyKey("city").asText().ifNotExist().create();
schema.propertyKey("date").asText().ifNotExist().create();
schema.propertyKey("price").asDouble().ifNotExist().create();

// Create the person vertex type, which has three attributes: name, age, city, and the primary key is name
schema.vertexLabel("person").properties("name", "age", "city").primaryKeys("name").ifNotExist().create();
// Create a software vertex type, which has two properties: name, price, the primary key is name
schema.vertexLabel("software").properties("name", "price").primaryKeys("name").ifNotExist().create();

// Create the knows edge type, which goes from person to person
schema.edgeLabel("knows").sourceLabel("person").targetLabel("person").properties("date").ifNotExist().create();
// Create the created edge type, which points from person to software
schema.edgeLabel("created").sourceLabel("person").targetLabel("software").ifNotExist().create();

Please refer to the corresponding section in hugegraph-client for the detailed description of the schema.

3.2 Prepare data

The data sources currently supported by HugeGraph-Loader include:

  • local disk file or directory
  • HDFS file or directory
  • Partial relational database
  • Kafka topic
  • An existing HugeGraph graph
3.2.1 Data source structure
3.2.1.1 Local disk file or directory

The user can specify a local disk file as the data source. If the data is scattered in multiple files, a certain directory is also supported as the data source, but multiple directories are not supported as the data source for the time being.

For example, my data is scattered in multiple files, part-0, part-1 … part-n. To perform the import, it must be ensured that they are placed in one directory. Then in the loader’s mapping file, specify path as the directory.

Supported file formats include:

  • TEXT
  • CSV
  • JSON

TEXT is a text file with custom delimiters, the first line is usually the header, and the name of each column is recorded, and no header line is allowed (specified in the mapping file). Each remaining row represents a record, which will be converted into a vertex/edge; each column of the row corresponds to a field, which will be converted into the id, label or attribute of the vertex/edge;

An example is as follows:

id|name|lang|price|ISBN
1|lop|java|328|ISBN978-7-107-18618-5
2|ripple|java|199|ISBN978-7-100-13678-5

CSV is a TEXT file with commas , as delimiters. When a column value itself contains a comma, the column value needs to be enclosed in double quotes, for example:

marko,29,Beijing
"li,nary",26,"Wu,han"

The JSON file requires that each line is a JSON string, and the format of each line needs to be consistent.

{"source_name": "marko", "target_name": "vadas", "date": "20160110", "weight": 0.5}
{"source_name": "marko", "target_name": "josh", "date": "20130220", "weight": 1.0}
3.2.1.2 HDFS file or directory

Users can also specify HDFS files or directories as data sources, all of the above requirements for local disk files or directories apply here. In addition, since HDFS usually stores compressed files, loader also provides support for compressed files, and local disk file or directory also supports compressed files.

Currently supported compressed file types include: GZIP, BZ2, XZ, LZMA, SNAPPY_RAW, SNAPPY_FRAMED, Z, DEFLATE, LZ4_BLOCK, LZ4_FRAMED, ORC, and PARQUET.

3.2.1.3 Mainstream relational database

The loader also supports some relational databases as data sources, and currently supports MySQL, PostgreSQL, Oracle, and SQL Server.

However, the requirements for the table structure are relatively strict at present. If association query needs to be done during the import process, such a table structure is not allowed. The associated query means: after reading a row of the table, it is found that the value of a certain column cannot be used directly (such as a foreign key), and you need to do another query to determine the true value of the column.

For example, Suppose there are three tables, person, software and created

// person schema
id | name | age | city
// software schema
id | name | lang | price
// created schema
id | p_id | s_id | date

If the id strategy of person or software is specified as PRIMARY_KEY when modeling (schema), choose name as the primary key (note: this is the concept of vertex-label in hugegraph), when importing edge data, the source vertex and target need to be spliced ​​out. For the id of the vertex, you must go to the person/software table with p_id/s_id to find the corresponding name. In the case of the schema that requires additional query, the loader does not support it temporarily. In this case, the following two methods can be used instead:

  1. The id strategy of person and software is still specified as PRIMARY_KEY, but the id column of the person table and software table is used as the primary key attribute of the vertex, so that the id can be generated by directly splicing p_id and s_id with the label of the vertex when importing an edge;
  2. Specify the id policy of person and software as CUSTOMIZE, and then directly use the id column of the person table and the software table as the vertex id, so that p_id and s_id can be used directly when importing edges;

The key point is to make the edge use p_id and s_id directly, don’t check it again.

3.2.2 Prepare vertex and edge data
3.2.2.1 Vertex Data

The vertex data file consists of data line by line. Generally, each line is used as a vertex, and each column is used as a vertex attribute. The following description uses CSV format as an example.

  • person vertex data (the data itself does not contain a header)
Tom,48,Beijing
Jerry,36,Shanghai
  • software vertex data (the data itself contains the header)
name,price
Photoshop,999
Office,388
3.2.2.2 Edge data

The edge data file consists of data line by line. Generally, each line is used as an edge. Some columns are used as the IDs of the source and target vertices, and other columns are used as edge attributes. The following uses JSON format as an example.

  • knows edge data
{"source_name": "Tom", "target_name": "Jerry", "date": "2008-12-12"}
  • created edge data
{"source_name": "Tom", "target_name": "Photoshop"}
{"source_name": "Tom", "target_name": "Office"}
{"source_name": "Jerry", "target_name": "Office"}

3.3 Write data source mapping file

3.3.1 Mapping file overview

The mapping file of the input source is used to describe how to establish the mapping relationship between the input source data and the vertex type/edge type of the graph. It is organized in JSON format and consists of multiple mapping blocks, each of which is responsible for mapping an input source. Mapped to vertices and edges.

Specifically, each mapping block contains an input source and multiple vertex mapping and edge mapping blocks, and the input source block corresponds to the local disk file or directory, HDFS file or directory and relational database are responsible for describing the basic information of the data source, such as where the data is, what format, what is the delimiter, etc. The vertex map/edge map is bound to the input source, which columns of the input source can be selected, which columns are used as ids, which columns are used as attributes, and what attributes are mapped to each column, the values ​​of the columns are mapped to what values ​​of attributes, and so on.

In the simplest terms, each mapping block describes: where is the file to be imported, which type of vertices/edges each line of the file is to be used as which columns of the file need to be imported, and the corresponding vertices/edges of these columns. what properties, etc.

Note: The format of the mapping file before version 0.11.0 and the format after 0.11.0 has changed greatly. For the convenience of expression, the mapping file (format) before 0.11.0 is called version 1.0, and the version after 0.11.0 is version 2.0. And unless otherwise specified, the “map file” refers to version 2.0.

Click to expand/collapse the skeleton of the map file for version 2.0
{
  "version": "2.0",
  "structs": [
    {
      "id": "1",
      "input": {
      },
      "vertices": [
        {},
        {}
      ],
      "edges": [
        {},
        {}
      ]
    }
  ]
}

Two versions of the mapping file are given directly here (the above graph model and data file are described)

Click to expand/collapse the mapping file for version 2.0
{
  "version": "2.0",
  "structs": [
    {
      "id": "1",
      "skip": false,
      "input": {
        "type": "FILE",
        "path": "vertex_person.csv",
        "file_filter": {
          "extensions": [
            "*"
          ]
        },
        "format": "CSV",
        "delimiter": ",",
        "date_format": "yyyy-MM-dd HH:mm:ss",
        "time_zone": "GMT+8",
        "skipped_line": {
          "regex": "(^#|^//).*|"
        },
        "compression": "NONE",
        "header": [
          "name",
          "age",
          "city"
        ],
        "charset": "UTF-8",
        "list_format": {
          "start_symbol": "[",
          "elem_delimiter": "|",
          "end_symbol": "]"
        }
      },
      "vertices": [
        {
          "label": "person",
          "skip": false,
          "id": null,
          "unfold": false,
          "field_mapping": {},
          "value_mapping": {},
          "selected": [],
          "ignored": [],
          "null_values": [
            ""
          ],
          "update_strategies": {}
        }
      ],
      "edges": []
    },
    {
      "id": "2",
      "skip": false,
      "input": {
        "type": "FILE",
        "path": "vertex_software.csv",
        "file_filter": {
          "extensions": [
            "*"
          ]
        },
        "format": "CSV",
        "delimiter": ",",
        "date_format": "yyyy-MM-dd HH:mm:ss",
        "time_zone": "GMT+8",
        "skipped_line": {
          "regex": "(^#|^//).*|"
        },
        "compression": "NONE",
        "header": null,
        "charset": "UTF-8",
        "list_format": {
          "start_symbol": "",
          "elem_delimiter": ",",
          "end_symbol": ""
        }
      },
      "vertices": [
        {
          "label": "software",
          "skip": false,
          "id": null,
          "unfold": false,
          "field_mapping": {},
          "value_mapping": {},
          "selected": [],
          "ignored": [],
          "null_values": [
            ""
          ],
          "update_strategies": {}
        }
      ],
      "edges": []
    },
    {
      "id": "3",
      "skip": false,
      "input": {
        "type": "FILE",
        "path": "edge_knows.json",
        "file_filter": {
          "extensions": [
            "*"
          ]
        },
        "format": "JSON",
        "delimiter": null,
        "date_format": "yyyy-MM-dd HH:mm:ss",
        "time_zone": "GMT+8",
        "skipped_line": {
          "regex": "(^#|^//).*|"
        },
        "compression": "NONE",
        "header": null,
        "charset": "UTF-8",
        "list_format": null
      },
      "vertices": [],
      "edges": [
        {
          "label": "knows",
          "skip": false,
          "source": [
            "source_name"
          ],
          "unfold_source": false,
          "target": [
            "target_name"
          ],
          "unfold_target": false,
          "field_mapping": {
            "source_name": "name",
            "target_name": "name"
          },
          "value_mapping": {},
          "selected": [],
          "ignored": [],
          "null_values": [
            ""
          ],
          "update_strategies": {}
        }
      ]
    },
    {
      "id": "4",
      "skip": false,
      "input": {
        "type": "FILE",
        "path": "edge_created.json",
        "file_filter": {
          "extensions": [
            "*"
          ]
        },
        "format": "JSON",
        "delimiter": null,
        "date_format": "yyyy-MM-dd HH:mm:ss",
        "time_zone": "GMT+8",
        "skipped_line": {
          "regex": "(^#|^//).*|"
        },
        "compression": "NONE",
        "header": null,
        "charset": "UTF-8",
        "list_format": null
      },
      "vertices": [],
      "edges": [
        {
          "label": "created",
          "skip": false,
          "source": [
            "source_name"
          ],
          "unfold_source": false,
          "target": [
            "target_name"
          ],
          "unfold_target": false,
          "field_mapping": {
            "source_name": "name",
            "target_name": "name"
          },
          "value_mapping": {},
          "selected": [],
          "ignored": [],
          "null_values": [
            ""
          ],
          "update_strategies": {}
        }
      ]
    }
  ]
}

Click to expand/collapse the mapping file for version 1.0
{
  "vertices": [
    {
      "label": "person",
      "input": {
        "type": "file",
        "path": "vertex_person.csv",
        "format": "CSV",
        "header": ["name", "age", "city"],
        "charset": "UTF-8"
      }
    },
    {
      "label": "software",
      "input": {
        "type": "file",
        "path": "vertex_software.csv",
        "format": "CSV"
      }
    }
  ],
  "edges": [
    {
      "label": "knows",
      "source": ["source_name"],
      "target": ["target_name"],
      "input": {
        "type": "file",
        "path": "edge_knows.json",
        "format": "JSON"
      },
      "field_mapping": {
        "source_name": "name",
        "target_name": "name"
      }
    },
    {
      "label": "created",
      "source": ["source_name"],
      "target": ["target_name"],
      "input": {
        "type": "file",
        "path": "edge_created.json",
        "format": "JSON"
      },
      "field_mapping": {
        "source_name": "name",
        "target_name": "name"
      }
    }
  ]
}

The 1.0 version of the mapping file is centered on the vertex and edge, and sets the input source; while the 2.0 version is centered on the input source, and sets the vertex and edge mapping. Some input sources (such as a file) can generate both vertices and edges. If you write in the 1.0 format, you need to write an input block in each of the vertex and edge mapping blocks. The two input blocks are exactly the same; and the 2.0 version only needs to write input once. Therefore, compared with version 1.0, version 2.0 can save some repetitive writing of input.

In the bin directory of hugegraph-loader-{version}, there is a script tool mapping-convert.sh that can directly convert the mapping file of version 1.0 to version 2.0. The usage is as follows:

bin/mapping-convert.sh struct.json

A struct-v2.json will be generated in the same directory as struct.json.

The bin directory also ships utf8-bom-to-utf8.sh, which strips the UTF-8 BOM from a single data file, or from every file under a directory. It is useful when a CSV or TEXT file exported by a Windows tool fails to parse because its first header column carries an invisible BOM:

bin/utf8-bom-to-utf8.sh /path/to/file-or-dir
3.3.2 Input Source

Input sources are currently divided into five categories: FILE, HDFS, JDBC, KAFKA and GRAPH, which are distinguished by the type node. We call them local file input sources, HDFS input sources, JDBC input sources, KAFKA input sources and GRAPH input source, which are described below.

3.3.2.1 Local file input source
  • id: The id of the input source. This field is used to support some internal functions. It is not required (it will be automatically generated if it is not filled in). It is strongly recommended to write it, which is very helpful for debugging;
  • skip: whether to skip the input source, because the JSON file cannot add comments, if you do not want to import an input source during a certain import, but do not want to delete the configuration of the input source, you can set it to true to skip it, the default is false, not required;
  • input: input source map block, composite structure
    • type: an input source type, file or FILE must be filled;
    • path: the path of the local file or directory, an absolute path or a path relative to the Loader process working directory (not the mapping file). An absolute path is recommended; required;
    • file_filter: filter files with compound conditions from path, compound structure, currently only supports configuration extensions, represented by child node extensions, the default is “*”, which means to keep all files;
    • format: the format of the local file, the optional values are CSV, TEXT and JSON, which must be uppercase, the default is CSV, optional;
    • header: column names; if omitted, the first data line supplies them. With an explicit header, an identical first line is still skipped by default (see has_header). JSON does not require a header; optional;
    • has_header: for CSV and TEXT, the first line of every file is dropped when it is identical to the header, so a header repeated in each part file of a directory is not imported as data. Set this to false to turn that check off when the first line of a file is real data that happens to equal the header, optional;
    • delimiter: The column delimiter of the file line. The default depends on format: a comma "," for CSV and a tab "\t" for TEXT; a CSV file accepts no other delimiter. The JSON file does not need to be specified, optional;
    • charset: the encoded character set of the file, the default is UTF-8, optional;
    • date_format: custom date format, the default value is yyyy-MM-dd HH:mm:ss, optional; if the date is presented in the form of a timestamp, this item must be written as timestamp (fixed writing);
    • extra_date_formats: a list of fallback date formats, tried when a value does not match date_format, empty by default, optional;
    • time_zone: Set which time zone the date data is in, the default value is GMT+8, optional;
    • skipped_line: The line to be skipped, compound structure, currently only the regular expression of the line to be skipped can be configured, described by the child node regex. The default regex is (^#|^//).*|, which skips lines starting with # or // and empty lines; to keep such lines, set regex to a pattern that matches nothing, optional;
    • compression: The compression format of the file, the optional values ​​are NONE, GZIP, BZ2, XZ, LZMA, SNAPPY_RAW, SNAPPY_FRAMED, Z, DEFLATE, LZ4_BLOCK, LZ4_FRAMED, ORC and PARQUET, the default is NONE, which means a non-compressed file, optional; for ORC and PARQUET the header is matched case-insensitively;
    • list_format: When a column of the file (non-JSON) is a collection structure (the Cardinality of the PropertyKey in the corresponding figure is Set or List), you can use this item to set the start character, separator, and end character of the column, compound structure :
      • start_symbol: The start character of the collection structure column (the default value is the empty string "", JSON format currently does not support specification)
      • elem_delimiter: the delimiter of the collection structure column (the default value is |, and it must differ from delimiter; JSON format currently only supports native , delimiter)
      • end_symbol: the end character of the collection structure column (the default value is the empty string "", the JSON format does not currently support specification)
      • ignored_elems: the elements dropped after the column is split, the default value is [""], so empty elements are ignored
3.3.2.2 HDFS input source

The nodes and meanings of the above local file input source are basically applicable here. Only the different and unique nodes of the HDFS input source are listed below.

  • type: input source type, must fill in hdfs or HDFS, required;
  • path: the path of the HDFS file or directory, it must be the absolute path of HDFS, required;
  • core_site_path: the path of the core-site.xml file of the HDFS cluster, the key point is to specify the address of the NameNode (fs.default.name) and the implementation of the file system (fs.hdfs.impl), required;
  • hdfs_site_path: the path of the hdfs-site.xml file of the HDFS cluster, optional;
  • dir_filter: when path is a directory, decides which sub-directories are walked into, compound structure, optional:
    • include_regex: only directories whose name matches this regular expression are read, empty by default, which places no restriction;
    • exclude_regex: directories whose name matches this regular expression are skipped, empty by default;
  • kerberos_config: how to authenticate against a Kerberos-secured HDFS cluster, compound structure, optional:
    • enable: whether to authenticate with Kerberos, the default is false;
    • krb5_conf: the path of the krb5.conf file, required when enable is true;
    • principal: the Kerberos principal, required when enable is true;
    • keytab: the path of the keytab file, required when enable is true;
3.3.2.3 JDBC input source

As mentioned above, it supports multiple relational databases, but because their mapping structures are very similar, they are collectively referred to as JDBC input sources, and then use the vendor node to distinguish different databases.

  • type: input source type, must fill in jdbc or JDBC, required;
  • vendor: database type, optional options are [MySQL, PostgreSQL, Oracle, SQLServer], case-insensitive, required;
  • driver: the JDBC driver class, optional; when it is left out, the default driver of the vendor listed in the tables below is used;
  • url: the url of the database that jdbc wants to connect to, required;
  • database: the name of the database to be connected, required;
  • schema: The name of the schema to be connected, different databases have different requirements, and the details are explained below;
  • table: the name of the table to be connected, at least one of table or custom_sql is required;
  • custom_sql: custom SQL statement, at least one of table or custom_sql is required;
  • username: username to connect to the database, required;
  • password: password for connecting to the database, required;
  • where: an extra condition appended to the generated select statement, written without the where keyword, optional;
  • batch_size: The size of one page when obtaining table data by page, the default is 500, optional;

MYSQL

NodeFixed value or common value
vendorMYSQL
drivercom.mysql.cj.jdbc.Driver
urljdbc:mysql://127.0.0.1:3306

schema: nullable, if filled in, it must be the same as the value of database

POSTGRESQL

NodeFixed value or common value
vendorPOSTGRESQL
driverorg.postgresql.Driver
urljdbc:postgresql://127.0.0.1:5432

schema: nullable, default is “public”

ORACLE

NodeFixed value or common value
vendorORACLE
driveroracle.jdbc.driver.OracleDriver
urljdbc:oracle:thin:@127.0.0.1:1521

schema: nullable, the default value is the username in upper case

SQLSERVER

NodeFixed value or common value
vendorSQLSERVER
drivercom.microsoft.sqlserver.jdbc.SQLServerDriver
urljdbc:sqlserver://127.0.0.1:1433

schema: required

3.3.2.4 Kafka input source
  • type: input source type, kafka or KAFKA, required;
  • bootstrap_server: the list of kafka bootstrap servers, required;
  • topic: the topic to subscribe to, required;
  • group: group of Kafka consumers, required;
  • from_beginning: sets auto.offset.reset to earliest when true, or latest when false (default). It applies only when no valid committed offset exists; otherwise consumption resumes from the committed offset. Optional;
  • format: format of each message, options are CSV, TEXT and JSON, must be uppercase, required;
  • header: column name of each column of a message; no header line is read from the topic, so it has to be given for CSV and TEXT, while JSON messages do not need it;
  • delimiter: delimiter of the message columns, used by TEXT only, since CSV always splits on ,, optional;
  • charset: encoding charset of the messages, default is UTF-8, optional;
  • date_format: customized date format, default value is yyyy-MM-dd HH:mm:ss, optional; if the date is presented in the form of timestamp, this item must be written as timestamp (fixed);
  • extra_date_formats: a customized list of another date formats, empty by default, optional; each item in the list is an alternate date format to the date_format specified date format;
  • time_zone: set which time zone the date data is in, default is GMT+8, optional;
  • skipped_line: current master KafkaReader does not implement this filter. Kafka messages go directly to the parser, so configuring this field does not skip blank or comment messages; preprocess them upstream or in the consumer to avoid parse failures;
  • batch_size: the maximum number of records fetched in one poll (max.poll.records), default is 500, optional;
  • early_stop: the record pulled from Kafka broker at a certain time is empty, stop the task, default is false, only for debugging, optional;
3.3.2.5 GRAPH input Source

The GRAPH input source reads vertices and edges out of another HugeGraph graph, reached through HugeGraph-PD, and writes them into the target graph. When a mapping file contains a GRAPH input source, every input source in it that is not skipped has to be a GRAPH input source as well, and the loader puts the target graph into RESTORING mode for the duration of the import.

  • type: Data source type; must be filled in as graph or GRAPH (required);
  • graphspace: Source graphSpace name (required);
  • graph: Source graph name (required);
  • username: HugeGraph username; the --username command-line option is used when this is empty;
  • password: HugeGraph password; the --password command-line option is used when this is empty;
  • selected_vertices: the vertex labels to copy, each item written as {"label": "...", "properties": [...], "query": {...}}, where properties narrows the copied properties and query is an optional filter passed to the source graph;
  • ignored_vertices: the vertex labels to skip, each item written as {"label": "...", "properties": [...]};
  • selected_edges: the edge labels to copy, items have the same shape as in selected_vertices;
  • ignored_edges: the edge labels to skip, items have the same shape as in ignored_vertices;
  • pd-peers: HugeGraph-PD node addresses of the source cluster; the --pd-peers option is used when this is empty;
  • meta-endpoints: Meta service endpoints of the source cluster; the --meta-endpoints option is used when this is empty;
  • cluster: Source cluster name; the --cluster option is used when this is empty;
  • batch_size: Batch size for reading data from the source graph; default is 500;
3.3.3 Vertex and Edge Mapping

The nodes of vertex and edge mapping (a key in the JSON file) have a lot of the same parts. The same parts are introduced first, and then the unique nodes of vertex map and edge map are introduced respectively.

Nodes of the same section

  • label: label to which the vertex/edge data to be imported belongs, required;
  • skip: whether to skip this vertex/edge mapping while the input source and the other mappings stay active, the default is false, optional;
  • field_mapping: Map the column name of the input source column to the attribute name of the vertex/edge, optional;
  • value_mapping: map the data value of the input source to the attribute value of the vertex/edge, optional;
  • selected: select some columns to insert, other unselected ones are not inserted, cannot exist at the same time as ignored, optional;
  • ignored: ignore some columns so that they do not participate in insertion, cannot exist at the same time as selected, optional;
  • null_values: You can specify some strings to represent null values, such as “NULL”. If the vertex/edge attribute corresponding to this column is also a nullable attribute, the value of this attribute will not be set when constructing the vertex/edge, optional ;
  • update_strategies: If the data needs to be updated in batches in a specific way, you can specify a specific update strategy for each attribute (see below for details), optional;
  • unfold: Whether to unfold the column, each unfolded column will form a row with other columns, which is equivalent to unfolding into multiple rows; for example, the value of a certain column (id column) of the file is [1,2,3], The values ​​of other columns are 18,Beijing. When unfold is set, this row will become 3 rows, namely: 1,18,Beijing, 2,18,Beijing and 3,18, Beijing. Note that this will only expand the column selected as id. Default false, optional;

Update strategy supports 8 types: (requires all uppercase)

  1. Value accumulation: SUM
  2. Take the greater of the two numbers/dates: BIGGER
  3. Take the smaller of two numbers/dates: SMALLER
  4. Set property takes union: UNION
  5. Set attribute intersection: INTERSECTION
  6. List attribute append element: APPEND
  7. List/Set attribute delete element: ELIMINATE
  8. Override an existing property: OVERRIDE

Note: If the newly imported attribute value is empty, the existing old data will be used instead of the empty value. For the effect, please refer to the following example

// The update strategy is specified in the JSON file as follows
{
  "vertices": [
    {
      "label": "person",
      "update_strategies": {
        "age": "SMALLER",
        "set": "UNION"
      },
      "input": {
        "type": "file",
        "path": "vertex_person.txt",
        "format": "TEXT",
        "header": ["name", "age", "set"]
      }
    }
  ]
}

// 1. Write a line of data with the OVERRIDE update strategy (null means empty here)
'a b null null'

// 2. Write another line
'null null c d'

// 3. Finally we can get
'a b c d'   

// If there is no update strategy, you will get
'null null c d'

Note : After adopting the batch update strategy, the number of disk read requests will increase significantly, and the import speed will be several times slower than that of pure write coverage (at this time HDD disk [IOPS](https://en.wikipedia .org/wiki/IOPS) will be the bottleneck, SSD is recommended for speed)

Unique Nodes for Vertex Maps

  • id: Specify a column as the id column of the vertex. When the vertex id policy is CUSTOMIZE, it is required; when the id policy is PRIMARY_KEY, it must be empty;

Unique Nodes for Edge Maps

  • source: Select certain columns of the input source as the id column of source vertex. When the id policy of the source vertex is CUSTOMIZE, a certain column must be specified as the id column of the vertex; when the id policy of the source vertex is When PRIMARY_KEY, one or more columns must be specified for splicing the id of the generated vertex, that is, no matter which id strategy is used, this item is required;
  • target: Specify certain columns as the id columns of target vertex, similar to source, so I won’t repeat them;
  • unfold_source: Whether to unfold the source column of the file, the effect is similar to that in the vertex map, and will not be repeated;
  • unfold_target: Whether to unfold the target column of the file, the effect is similar to that in the vertex mapping, and will not be repeated;

3.4 Execute command import

After preparing the graph model, data file, and input source mapping relationship file, the data file can be imported into the graph database.

The import process is controlled by commands submitted by the user, and the user can control the specific process of execution through different parameters.

3.4.1 Parameter description
ParameterDefault valueRequired or notDescription
-f or --fileYPath to configure script
-g or --graphhugegraphGraph name
--graphspaceDEFAULTGraph space name
-s or --schemaSchema file path; optional when the Schema already exists
-h or --host or -ilocalhostAddress of HugeGraphServer
-p or --port8080Port number of HugeGraphServer
--usernamenullWhen HugeGraphServer enables permission authentication, the username of the current graph
--passwordnullWhen HugeGraphServer enables permission authentication, the password of the current graph
--create-graphfalseWhether to automatically create the graph if it does not exist
--tokennullWhen HugeGraphServer has enabled authorization authentication, the token of the current graph
--protocolhttpProtocol for sending requests to the server, optional http or https
--pd-peersPD service node addresses
--pd-tokenToken for accessing PD service
--meta-endpointsMeta information storage service addresses
--directfalseStore direct mode is disabled in current master; imports still use Server API. Keep the default
--route-typeNODE_PORTRoute selection method (optional values: NODE_PORT / DDS / BOTH)
--clusterhgCluster name
--trust-store-fileWhen the request protocol is https, the client’s certificate file path
--trust-store-passwordWhen the request protocol is https, the client certificate password
--clear-all-datafalseWhether to clear the original data on the server before importing data
--clear-timeout240Timeout for clearing the original data on the server before importing data
--incremental-modefalseWhether to use the breakpoint resume mode; only input sources FILE and HDFS support this mode. Enabling this mode allows starting the import from where the last import stopped
--failure-modefalseWhen failure mode is true, previously failed data will be imported. Generally, the failed data file needs to be manually corrected and edited before re-importing
--batch-insert-threadsCPUsBatch insert thread pool size (CPUs is the number of logical cores available to the current OS)
--single-insert-threads8Size of single insert thread pool
--max-conn4 * CPUsThe maximum number of HTTP connections between HugeClient and HugeGraphServer; while it is left at its default, it is raised automatically to 4 * --batch-insert-threads
--max-conn-per-route2 * CPUsThe maximum number of HTTP connections for each route between HugeClient and HugeGraphServer; while it is left at its default, it is raised automatically to 2 * --batch-insert-threads
--batch-size500The number of data items in each batch when importing data
--max-parse-errors1The maximum number of data parsing errors allowed (per line); the program exits when this value is reached
--max-insert-errors500The maximum number of data insertion errors allowed (per row); the program exits when this value is reached
--timeout60Timeout (seconds) for insert result return
--shutdown-timeout10Waiting time for multithreading to stop (seconds)
--retry-times3Maximum number of retries after a timeout
--retry-interval10Interval before retry (seconds)
--check-vertexfalseWhether to check if the vertices connected by the edge exist when inserting the edge
--print-progresstrueWhether to print the number of imported items in real time on the console
--dry-runfalseEnable this mode to only parse data without importing; usually used for testing
--help or -helpfalsePrint help information
--parser-threads or --parallel-countmax(2,CPUs/2)Number of parallel read pipelines; --parallel-count is deprecated
--start-file0Start file index for partial loading
--end-file-1End file index for partial loading
--scatter-sourcesfalseScatter multiple sources for I/O optimization
--cdc-flush-interval30000The flush interval for Flink CDC
--cdc-sink-parallelism1The sink parallelism for Flink CDC
--max-read-errors1The maximum number of read error lines before exiting
--max-read-lines-1LThe maximum number of read lines, task stops when reached
--test-modefalseWhether the loader works in test mode
--use-prefilterfalseWhether to filter vertex in advance
--short-idMap a customized vertex ID to a shorter generated ID, written as label:field:type, where type is one of boolean, byte, int, long, float, double, text, blob, date and uuid; repeat the option to cover several labels
--vertex-edge-limit-1LThe maximum number of vertex’s edges
--sink-typetruespark-loader only: true writes through the HugeGraph server API, false generates HFiles and bulk-loads them into HBase
--vertex-partitions64The number of partitions of the HBase vertex table, used with --sink-type false
--edge-partitions64The number of partitions of the HBase edge table, used with --sink-type false
--vertex-table-nameHBase vertex table name, used with --sink-type false
--edge-table-nameHBase edge table name, used with --sink-type false
--hbase-zk-quorumHBase ZooKeeper quorum, used with --sink-type false
--hbase-zk-portHBase ZooKeeper port, used with --sink-type false
--hbase-zk-parentHBase ZooKeeper parent, used with --sink-type false
--restorefalseSet graph mode to RESTORING
--backendhstoreThe backend store type when creating graph if not exists
--serializerbinaryThe serializer type when creating graph if not exists
--scheduler-typedistributedThe task scheduler type when creating graph if not exists
--batch-failure-fallbacktrueWhether to fallback to single insert when batch insert fails

The loader prints its usage and exits when it is given fewer than three arguments, so -f struct.json on its own is not enough.

3.4.2 Breakpoint Continuation Mode

Usually, the Loader task takes a long time to execute. If the import interrupt process exits for some reason, and next time you want to continue the import from the interrupted point, this is the scenario of using breakpoint continuation.

The user sets the command line parameter –incremental-mode to true to open the breakpoint resume mode. The key to breakpoint continuation lies in the progress file. When the import process exits, the import progress at the time of exit will be recorded. Recorded in the progress file, the progress file is located in the ${struct} directory, the file name is like load-progress_${timestamp}, ${struct} is the prefix of the mapping file, and ${timestamp} is the start of the import moment, formatted as yyyyMMdd-HHmmss. For example, for an import task started at 2019-10-10 12:30:30, the mapping file used is struct-example.json, then the path of the progress file is the same as struct-example.json Sibling struct-example/load-progress_20191010-123030. When the directory holds several progress files, the resumed import reads the last one in name order, which is the most recent one.

Note: The generation of progress files is independent of whether –incremental-mode is turned on or not, and a progress file is generated at the end of each import.

If the data file formats are all legal and the import task is stopped by the user (CTRL + C or kill, kill -9 is not supported), that is to say, if there is no error record, the next import only needs to be set to Continue for the breakpoint.

But if the limit of –max-read-errors, –max-parse-errors or –max-insert-errors is reached because too much data is invalid or network abnormality is reached, Loader will record these original rows that failed into the failure file, after the user modifies the data lines in the failure file, set –failure-mode to true to import these “failure files” as input sources (does not affect the normal file import), Of course, if there is still a problem with the modified data line, it will be logged again to the failure file (don’t worry about duplicate lines, they are dropped when the file is closed). Failure mode lifts the three error limits above, so the whole failure file is scanned.

Each input source, that is each item of structs in the mapping file, gets its own failure file. The file is named after the id of that input source with the suffix .error and is stored in the ${struct}/failure-data directory. Every failed line is written as a pair of lines: a tip line starting with #### READ ERROR:, #### PARSE ERROR: or #### INSERT ERROR:, followed by the original data line. When the input source has a header, that header is written as JSON to a sibling ${id}.header file, so the failure file can be read back with the right columns. For example, if the mapping file has an input source with id 1 holding a vertex mapping person and an input source with id 3 holding an edge mapping knows, each of which has some error lines, you will see the following files in the ${struct}/failure-data directory when the Loader exits:

  • 1.error: the failed lines of input source 1, each preceded by its tip line
  • 1.header: the header of input source 1, written only when that input source has a header
  • 3.error: the failed lines of input source 3
  • 3.header: the header of input source 3

A .error file that turns out to be empty is deleted when the Loader exits, so only the input sources that really had failed lines leave a file behind. In incremental mode new failures are appended to the existing file, otherwise the file is rewritten from scratch.

3.4.3 logs directory file description

The log and error data during program execution will be written into the hugegraph-loader.log file.

3.4.4 Execute command

Run bin/hugegraph-loader.sh and pass in parameters

bin/hugegraph-loader.sh -g {GRAPH_NAME} -f ${INPUT_DESC_FILE} -s ${SCHEMA_FILE} -h {HOST} -p {PORT}

The script runs the JVM under JAVA_HOME when that variable is set, and java from the PATH otherwise. It passes the contents of the JVM_OPTS environment variable, then -Xmx10g and the class path built from lib/, to that JVM, so JVM_OPTS is the place to add JVM flags. Logging is configured by conf/log4j2.xml.

3.4.5 Minimal Local-File Import

First start HugeGraph Server and ensure the current user can access hugegraph in graphspace DEFAULT, create schema, and write data. If authentication is enabled, add the appropriate --username, --password, or --token. This example requires only local files, without Kafka, HDFS, or JDBC.

From the extracted Loader installation directory, create data, schema, and a version 2.0 mapping file:

mkdir -p demo-loader

cat > demo-loader/people.csv <<'EOF'
docs-smoke-alice,29
docs-smoke-bob,31
EOF

cat > demo-loader/schema.groovy <<'EOF'
schema.propertyKey("docs_name").asText().ifNotExist().create()
schema.propertyKey("docs_age").asInt().ifNotExist().create()
schema.vertexLabel("docs_smoke_person").properties("docs_name", "docs_age").primaryKeys("docs_name").ifNotExist().create()
EOF

cat > demo-loader/struct.json <<'EOF'
{
  "version": "2.0",
  "structs": [
    {
      "id": "people",
      "input": {
        "type": "FILE",
        "path": "demo-loader/people.csv",
        "format": "CSV",
        "header": ["docs_name", "docs_age"]
      },
      "vertices": [
        {"label": "docs_smoke_person"}
      ],
      "edges": []
    }
  ]
}
EOF

The CSV has no header row; input.header in the mapping supplies column names. input.path is relative to the Loader process working directory, so keep running from the installation directory. If Server runs in another container, set --host to a Server address reachable by Loader:

bin/hugegraph-loader.sh \
  --graphspace DEFAULT --graph hugegraph \
  --file demo-loader/struct.json --schema demo-loader/schema.groovy \
  --host 127.0.0.1 --port 8080

Expected result: two vertices and zero edges imported. This example uses dedicated schema names to avoid conflicts with a preloaded person label. In Hubble, run g.V().hasLabel('docs_smoke_person').values('docs_name') to confirm docs-smoke-alice and docs-smoke-bob.

4 Complete example

Given below is an example in the example directory of the hugegraph-loader package. (GitHub address)

4.1 Prepare data

Vertex file: example/file/vertex_person.csv

marko,29,Beijing
vadas,27,Hongkong
josh,32,Beijing
peter,35,Shanghai
"li,nary",26,"Wu,han"
tom,null,NULL

Vertex file: example/file/vertex_software.txt

id|name|lang|price|ISBN
1|lop|java|328|ISBN978-7-107-18618-5
2|ripple|java|199|ISBN978-7-100-13678-5

Edge file: example/file/edge_knows.json

{"source_name": "marko", "target_name": "vadas", "date": "20160110", "weight": 0.5}
{"source_name": "marko", "target_name": "josh", "date": "20130220", "weight": 1.0}

Edge file: example/file/edge_created.json

{"source_name": "marko", "target_id": 1, "date": "2017-12-10", "weight": 0.4}
{"source_name": "josh", "target_id": 1, "date": "2009-11-11", "weight": 0.4}
{"source_name": "josh", "target_id": 2, "date": "2017-12-10", "weight": 1.0}
{"source_name": "peter", "target_id": 1, "date": "2017-03-24", "weight": 0.2}

4.2 Write schema

Click to expand/collapse the schema file: example/file/schema.groovy
schema.propertyKey("name").asText().ifNotExist().create();
schema.propertyKey("age").asInt().ifNotExist().create();
schema.propertyKey("city").asText().ifNotExist().create();
schema.propertyKey("weight").asDouble().ifNotExist().create();
schema.propertyKey("lang").asText().ifNotExist().create();
schema.propertyKey("date").asText().ifNotExist().create();
schema.propertyKey("price").asDouble().ifNotExist().create();

schema.vertexLabel("person").properties("name", "age", "city").primaryKeys("name").nullableKeys("age", "city").ifNotExist().create();
schema.vertexLabel("software").useCustomizeNumberId().properties("name", "lang", "price").ifNotExist().create();

schema.indexLabel("personByAge").onV("person").by("age").range().ifNotExist().create();
schema.indexLabel("personByCity").onV("person").by("city").secondary().ifNotExist().create();
schema.indexLabel("personByAgeAndCity").onV("person").by("age", "city").secondary().ifNotExist().create();
schema.indexLabel("softwareByPrice").onV("software").by("price").range().ifNotExist().create();

schema.edgeLabel("knows").sourceLabel("person").targetLabel("person").properties("date", "weight").ifNotExist().create();
schema.edgeLabel("created").sourceLabel("person").targetLabel("software").properties("date", "weight").ifNotExist().create();

schema.indexLabel("createdByDate").onE("created").by("date").secondary().ifNotExist().create();
schema.indexLabel("createdByWeight").onE("created").by("weight").range().ifNotExist().create();
schema.indexLabel("knowsByWeight").onE("knows").by("weight").range().ifNotExist().create();

person uses name as its primary key. Tom’s age and city are removed by null_values, so both properties must be nullable. software uses custom numeric IDs: mapping id supplies the vertex ID and created.target_id supplies the edge’s target ID. Run this example in an empty graph without these schema labels; ifNotExist() does not replace incompatible existing definitions.

4.3 Write the input source mapping file example/file/struct.json

Click to expand/collapse the input source mapping file example/file/struct.json
{
  "vertices": [
    {
      "label": "person",
      "input": {
        "type": "file",
        "path": "example/file/vertex_person.csv",
        "format": "CSV",
        "header": ["name", "age", "city"],
        "charset": "UTF-8",
        "skipped_line": {
          "regex": "(^#|^//).*"
        }
      },
      "null_values": ["NULL", "null", ""]
    },
    {
      "label": "software",
      "input": {
        "type": "file",
        "path": "example/file/vertex_software.txt",
        "format": "TEXT",
        "delimiter": "|",
        "charset": "GBK"
      },
      "id": "id",
      "ignored": ["ISBN"]
    }
  ],
  "edges": [
    {
      "label": "knows",
      "source": ["source_name"],
      "target": ["target_name"],
      "input": {
        "type": "file",
        "path": "example/file/edge_knows.json",
        "format": "JSON",
        "date_format": "yyyyMMdd"
      },
      "field_mapping": {
        "source_name": "name",
        "target_name": "name"
      }
    },
    {
      "label": "created",
      "source": ["source_name"],
      "target": ["target_id"],
      "input": {
        "type": "file",
        "path": "example/file/edge_created.json",
        "format": "JSON",
        "date_format": "yyyy-MM-dd"
      },
      "field_mapping": {
        "source_name": "name"
      }
    }
  ]
}

4.4 Command to import

sh bin/hugegraph-loader.sh -g hugegraph -f example/file/struct.json -s example/file/schema.groovy

After the import is complete, statistics similar to the following will appear:

vertices/edges has been loaded this time : 8/6
--------------------------------------------------
count metrics
     input read success            : 14
     input read failure            : 0
     vertex parse success          : 8
     vertex parse failure          : 0
     vertex insert success         : 8
     vertex insert failure         : 0
     edge parse success            : 6
     edge parse failure            : 0
     edge insert success           : 6
     edge insert failure           : 0

4.5 Use Docker to load data

4.5.1 Use docker exec to load data directly
4.5.1.1 Prepare data

If you just want to try out the loader, you can import the built-in example dataset without needing to prepare additional data yourself.

If using custom data, before importing data with the loader, we need to copy the data into the container.

First, following the steps in 4.1–4.3, we can prepare the data and then use docker cp to copy the prepared data into the loader container.

Suppose we’ve prepared the corresponding dataset following the above steps, stored in the hugegraph-dataset folder with the following file structure:

tree -f hugegraph-dataset/

hugegraph-dataset
├── hugegraph-dataset/edge_created.json
├── hugegraph-dataset/edge_knows.json
├── hugegraph-dataset/schema.groovy
├── hugegraph-dataset/struct.json
├── hugegraph-dataset/vertex_person.csv
└── hugegraph-dataset/vertex_software.txt

Copy the files into the container.

docker cp hugegraph-dataset loader:/loader/dataset
docker exec -it loader ls /loader/dataset

edge_created.json  edge_knows.json  schema.groovy  struct.json  vertex_person.csv  vertex_software.txt
4.5.1.2 Data loading

Taking the built-in example dataset as an example, we can use the following command to load the data.

If you need to import your custom dataset, you need to modify the paths for -f (data script) and -s (schema) configurations.

You can refer to 3.4.1-Parameter description for the rest of the parameters.

docker exec -it loader bin/hugegraph-loader.sh -g hugegraph -f example/file/struct.json -s example/file/schema.groovy -h server -p 8080

If loading a custom dataset, following the previous example, you would use:

docker exec -it loader bin/hugegraph-loader.sh -g hugegraph -f /loader/dataset/struct.json -s /loader/dataset/schema.groovy -h server -p 8080

Also update every input.path in the custom struct.json to its actual container path, such as /loader/dataset/vertex_person.csv. Changing only -f and -s does not change where Loader looks for data files.

If loader and server are in the same Docker network, you can specify -h {server_container_name}; otherwise, you need to specify the IP of the server host (in our example, server_container_name is server).

Then we can see the result:

HugeGraphLoader worked in NORMAL MODE
vertices/edges loaded this time : 8/6
--------------------------------------------------
count metrics
    input read success            : 14                  
    input read failure            : 0                   
    vertex parse success          : 8                   
    vertex parse failure          : 0                   
    vertex insert success         : 8                   
    vertex insert failure         : 0                   
    edge parse success            : 6                   
    edge parse failure            : 0                   
    edge insert success           : 6                   
    edge insert failure           : 0                   
--------------------------------------------------
meter metrics
    total time                    : 0.199s              
    read time                     : 0.046s              
    load time                     : 0.153s              
    vertex load time              : 0.077s              
    vertex load rate(vertices/s)  : 103                 
    edge load time                : 0.112s              
    edge load rate(edges/s)       : 53   

You can also use curl or hubble to observe the import result. Here’s an example using curl:

These requests use the GraphSpace-aware Server API. For an older Server, omit /graphspaces/DEFAULT from the path.

curl --compressed "http://localhost:8080/graphspaces/DEFAULT/graphs/hugegraph/graph/vertices"
{"vertices":[{"id":1,"label":"software","type":"vertex","properties":{"name":"lop","lang":"java","price":328.0}},{"id":2,"label":"software","type":"vertex","properties":{"name":"ripple","lang":"java","price":199.0}},{"id":"1:tom","label":"person","type":"vertex","properties":{"name":"tom"}},{"id":"1:josh","label":"person","type":"vertex","properties":{"name":"josh","age":32,"city":"Beijing"}},{"id":"1:marko","label":"person","type":"vertex","properties":{"name":"marko","age":29,"city":"Beijing"}},{"id":"1:peter","label":"person","type":"vertex","properties":{"name":"peter","age":35,"city":"Shanghai"}},{"id":"1:vadas","label":"person","type":"vertex","properties":{"name":"vadas","age":27,"city":"Hongkong"}},{"id":"1:li,nary","label":"person","type":"vertex","properties":{"name":"li,nary","age":26,"city":"Wu,han"}}]}

If you want to check the import result of edges, you can use curl --compressed "http://localhost:8080/graphspaces/DEFAULT/graphs/hugegraph/graph/edges".

4.5.2 Enter the docker container to load data

Besides using docker exec directly for data import, we can also enter the container for data loading. The basic process is similar to 4.5.1.

Enter the container by docker exec -it loader bash and execute the command:

sh bin/hugegraph-loader.sh -g hugegraph -f example/file/struct.json -s example/file/schema.groovy -h server -p 8080

The results of the execution will be similar to those shown in 4.5.1.

4.6 Import data by spark-loader

The current source uses Spark 3.2.2 and Scala 2.12. Other combinations need independent verification.

The parameters of spark-loader are divided into two parts. Note: Because the abbreviations of these two-parameter names have overlapping parts, please use the full name of the parameter. And there is no need to guarantee the order between the two parameters.

Example:

sh bin/hugegraph-spark-loader.sh --master yarn \
--deploy-mode cluster --name spark-hugegraph-loader --file ./hugegraph.json \
--username admin --token admin --host xx.xx.xx.xx --port 8093 \
--graph graph-test --num-executors 6 --executor-cores 16 --executor-memory 15g

bin/hugegraph-spark-loader.sh submits org.apache.hugegraph.loader.spark.HugeGraphSparkLoader through ${SPARK_HOME}/bin/spark-submit, so SPARK_HOME has to point at a Spark installation. Every jar under lib/ is put on the class path. The Spark application name defaults to hugegraph-spark-loader and can be changed through the APP_NAME environment variable.

bin/get-params.sh splits the command line: only the options below are handed to the loader, every other argument is passed to spark-submit unchanged. The splitter matches long option names only, so short forms such as -f and -g are not recognised.

--graph --schema --host --port --username --token --protocol
--trust-store-file --trust-store-password --clear-all-data --clear-timeout
--incremental-mode --failure-mode --batch-insert-threads --single-insert-threads
--max-conn --max-conn-per-route --batch-size --max-parse-errors --max-insert-errors
--timeout --shutdown-timeout --retry-times --retry-interval --check-vertex
--print-progress --dry-run --sink-type --vertex-partitions --edge-partitions --help

--file is treated separately: with --deploy-mode cluster the mapping file is shipped to the executors through --files and the loader receives only its base name, otherwise the path is passed through as written.

In this mode the loader reads FILE, HDFS and JDBC input sources; a KAFKA or GRAPH input source is rejected.

By default (--sink-type true) each Spark partition opens a HugeClient and writes vertices and edges through the HugeGraph server API. With --sink-type false the loader generates HFiles and bulk-loads them into HBase instead, taking the table names and ZooKeeper settings from --vertex-table-name, --edge-table-name, --hbase-zk-quorum, --hbase-zk-port, --hbase-zk-parent, --vertex-partitions and --edge-partitions.

The current source uses Flink 1.13.5 with flink-connector-mysql-cdc 2.2.1 and Scala 2.12. Other combinations need independent verification.

bin/hugegraph-flinkcdc-loader.sh submits org.apache.hugegraph.loader.flink.HugeGraphFlinkCDCLoader through ${FLINK_HOME}/bin/flink run, so FLINK_HOME has to be set. The job captures MySQL change events with Flink CDC and applies them to the graph, which keeps the graph in step with the source tables.

The mapping file uses the same format as for the command-line loader, but every input source has to be a JDBC input source over MySQL: the loader takes url, database, table, username and password from it and parses the host and port out of url. Vertex and edge mappings work as usual. --cdc-flush-interval and --cdc-sink-parallelism in 3.4.1 apply to this mode only.

The command line is split by bin/get-params.sh in the same way as for spark-loader, with the arguments that are not loader options going to flink run.

Example:

sh bin/hugegraph-flinkcdc-loader.sh --file ./mysql-cdc.json \
--host xx.xx.xx.xx --port 8080 --graph hugegraph --username admin --token admin

5 - Graph export and migration

Choose an export or migration tool when you need to back up, export, or move data between graphs. Tools is suited to standalone operations and backups; SeaTunnel Source connects graph reads to an expandable data pipeline.

5.1 - Export and Migrate Graph Data with SeaTunnel Source

If you need to copy data from one HugeGraph graph to another, use graph2graph: HugeGraph Source reads vertices and edges from the source graph (graph A), and HugeGraph Sink writes them to the target graph (graph B), with an optional Transform in between. The data path is graph A → HugeGraph Source → (optional Transform) → HugeGraph Sink → graph B. If you need to export graph data to a file, JDBC, Kafka, or another system, use graph2any, where a downstream Sink receives the records read by HugeGraph Source. This page covers both job types.

Version requirement: This guide targets SeaTunnel 3.0+

SeaTunnel 3.0+ provides HugeGraph Source, with schema auto-discovery and multi-label reads. See the official Source documentation for parallel-scan backend requirements and configuration limits.

Before starting, complete the shared environment and configuration steps on the import page. They cover JDK, HOCON, plugin installation, and the sample graph model.

1 Migrate a HugeGraph graph (graph2graph)

The following example migrates person vertices and knows edges from a source graph. Use a separate target graph. This section uses CUSTOMIZE_STRING to preserve vertex IDs. Do not reuse the person label created earlier with PRIMARY_KEY.

These two jobs migrate only the selected labels and properties. They do not copy every source schema setting, such as indexes and TTLs. Pause writes to the source graph during the migration so both jobs read a consistent point in time. Afterward, compare vertex and edge counts and sample properties.

Regenerating a primary key can change 1:marko to 2:marko; CUSTOMIZE_STRING preserves the original ID so edge endpoints still resolve

1.1 Migrate vertices first

Source adds a ~id column for the original ID, and Sink stores it as a string. Do not declare ~id in schema.fields; manually declaring this reserved column is rejected.

Expand the configuration and save it as config/graph2graph-person.conf
env {
  job.mode = "BATCH"
}

source {
  HugeGraph {
    host = "source-hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    label = "person"
    label_type = "VERTEX"
    schema = {
      fields = {
        name = "string"
        age = "int"
      }
    }
  }
}

sink {
  HugeGraph {
    host = "target-hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    batch_failure_fallback = false
    mappings = [
      {
        type = "VERTEX"
        label = "person"
        idStrategy = "CUSTOMIZE_STRING"
        idFields = ["~id"]
        properties = ["name", "age"]
      }
    ]
  }
}
./bin/seatunnel.sh --config ./config/graph2graph-person.conf -m local

1.2 Migrate edges second

After the vertex job succeeds, use the ~source_id and ~target_id columns added by Source to locate endpoints. Because the previous job preserved the original IDs, these columns can refer directly to vertices in the target graph.

Expand the configuration and save it as config/graph2graph-knows.conf
env {
  job.mode = "BATCH"
}

source {
  HugeGraph {
    host = "source-hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    label = "knows"
    label_type = "EDGE"
    schema = {
      fields = {
        since = "int"
      }
    }
  }
}

sink {
  HugeGraph {
    host = "target-hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    check_vertex = true
    batch_failure_fallback = false
    mappings = [
      {
        type = "EDGE"
        label = "knows"
        sourceConfig = {
          label = "person"
          idFields = ["~source_id"]
        }
        targetConfig = {
          label = "person"
          idFields = ["~target_id"]
        }
        properties = ["since"]
      }
    ]
  }
}
./bin/seatunnel.sh --config ./config/graph2graph-knows.conf -m local

This example checks endpoints and makes write errors fail the job. The default check_vertex = false does not guarantee a consistent result: a missing endpoint can create a dangling edge, so a successful job is not a substitute for checking the migrated graph.

Why preserve IDs? A HugeGraph PRIMARY_KEY ID contains the internal ID of the vertex label, and that internal ID can differ between graphs. For example, a source vertex can be 1:marko, while regenerating the primary key in the target graph can produce 2:marko. Reusing the source edge endpoints after regenerating vertex IDs can connect edges to the wrong vertices. This example stores the original ID as a string, which changes the target graph’s ID strategy

When Source reads every label, omit label to read all labels of label_type (default VERTEX). It produces one output table per label. Bind each Sink mapping to its table with sourceTable, for example sourceTable = "default.person"; use the full table name shown in the Writer log for the exact value. Do not reuse the single-label configuration from this section. See the HugeGraph Source documentation for other limitations.

2 Export to another system (graph2any)

graph2any uses HugeGraph Source[1] to read vertices or edges and sends them to a downstream Sink. The example below exports person vertices to local JSON files; to export to JDBC, Kafka, or another system, replace LocalFile[2] and its options.

env {
  job.mode = "BATCH"
}

source {
  HugeGraph {
    host = "hugegraph"
    port = 8080
    graph_name = "hugegraph"
    graph_space = "DEFAULT"
    label = "person"
    label_type = "VERTEX"
    schema = {
      fields = {
        name = "string"
        age = "int"
      }
    }
  }
}

sink {
  LocalFile {
    path = "/tmp/hugegraph-export/${table_name}"
    file_format_type = "json"
  }
}

Save this as config/graph2file-person.conf and run it from the SeaTunnel installation directory:

./bin/seatunnel.sh --config ./config/graph2file-person.conf -m local

To export edges, change the Source label to an edge label, set label_type = "EDGE", and declare the edge properties in schema.fields. Source also outputs the reserved columns ~source_id, ~source_label, ~target_id, and ~target_label; write them to the file or pass them to downstream transforms as needed.

This page covers row reads and writes. It does not copy source indexes, TTLs, or other schema settings. For the complete Source options and shared environment guidance, return to the SeaTunnel graph import guide[3].

3 References

Connectors

[1] HugeGraph Source
[2] LocalFile Sink

Related guide

[3] SeaTunnel graph import guide

6 - Tools Quick Start

1 HugeGraph-Tools Overview

HugeGraph-Tools is an automated deployment, management and backup/restore component of HugeGraph.

Testing Guide: For running HugeGraph-Tools tests locally, please refer to HugeGraph Toolchain Local Testing Guide

2 Get HugeGraph-Tools

HugeGraph-Tools is included in the Toolchain distribution. You can download the distribution or build it from source.

  • Download the compiled tarball
  • Clone source code then compile and install

2.1 Download the compiled archive

Download the latest version of the HugeGraph-Toolchain package:

export VERSION=1.7.0
export ARCHIVE="apache-hugegraph-toolchain-incubating-${VERSION}"
wget "https://downloads.apache.org/hugegraph/${VERSION}/${ARCHIVE}.tar.gz"
tar zxf "${ARCHIVE}.tar.gz"
# hugegraph-tools ships inside the toolchain package, in a directory
# whose version suffix is the same as the archive's
cd "${ARCHIVE}/apache-hugegraph-tools-incubating-${VERSION}"

2.2 Clone source code to compile and install

Please ensure that the wget command is installed before compiling the source code

Download the latest version of the HugeGraph-Tools source package:

# 1. get from github
git clone https://github.com/apache/hugegraph-toolchain.git

# 2. Download a released source package
export VERSION=1.7.0
export ARCHIVE="apache-hugegraph-toolchain-incubating-${VERSION}"
wget "https://downloads.apache.org/hugegraph/${VERSION}/${ARCHIVE}-src.tar.gz"

Compile and generate tar package:

cd hugegraph-toolchain
mvn package -pl hugegraph-tools -am -DskipTests -ntp

The package is generated as hugegraph-tools/target/apache-hugegraph-tools-${version}.tar.gz, and the unpacked directory hugegraph-tools/apache-hugegraph-tools-${version} (containing bin/ and lib/) is created next to it.

3 How to use

3.1 Function overview

After decompression, enter the apache-hugegraph-tools-${version} directory, you can use bin/hugegraph or bin/hugegraph help to view the usage information, and bin/hugegraph help <sub-command> to view the usage of a single sub-command. mainly divided:

  • Graph management type, graph-mode-set, graph-mode-get, graph-list, graph-get, graph-clear, graph-create, graph-clone and graph-drop
  • Asynchronous task management type, task-list, task-get, task-delete, task-cancel and task-clear
  • Gremlin type, gremlin-execute and gremlin-schedule
  • Backup/Restore type, backup, restore, migrate, schedule-backup and dump
  • Authentication data backup/restore type, auth-backup and auth-restore
  • Install deployment type, deploy, clear, start-all and stop-all
Usage: hugegraph [options] [command] [command options]
3.2 [options]-Global Variable

options is a global variable of HugeGraph-Tools, which can be configured in hugegraph-tools/bin/hugegraph, including:

  • –graph,HugeGraph-Tools The name of the graph to operate on, the default value is hugegraph
  • –url,The service address of HugeGraph-Server, the default is http://127.0.0.1:8080
  • –user,When HugeGraph-Server opens authentication, pass username
  • –password,When HugeGraph-Server opens authentication, pass the user’s password
  • –timeout,Timeout when connecting to HugeGraph-Server, the default is 30s
  • –trust-store-file,The path of the certificate file, when –url uses https, the truststore file used by HugeGraph-Client, the default is empty, which means using the built-in truststore file conf/hugegraph.truststore of hugegraph-tools
  • –trust-store-password,The password of the certificate file, when –url uses https, the password of the truststore used by HugeGraph-Client, the default is empty, representing the password of the built-in truststore file of hugegraph-tools
  • –throw-mode, whether HugeGraph-Tools throws the exception instead of printing the error message and exiting, the default is false (mainly used by tests)

The protocol is taken from the scheme of –url: use https://... to connect over https. –trust-store-file and –trust-store-password can only be set when –url uses https, and both –user and –password must be given together or omitted together.

The above global variables can also be set through environment variables. One way is to use export on the command line to set temporary environment variables, which are valid until the command line is closed

Global VariableEnvironment VariableExample
–urlHUGEGRAPH_URLexport HUGEGRAPH_URL=http://127.0.0.1:8080
–graphHUGEGRAPH_GRAPHexport HUGEGRAPH_GRAPH=hugegraph
–userHUGEGRAPH_USERNAMEexport HUGEGRAPH_USERNAME=admin
–passwordHUGEGRAPH_PASSWORDexport HUGEGRAPH_PASSWORD=test
–timeoutHUGEGRAPH_TIMEOUTexport HUGEGRAPH_TIMEOUT=30
–trust-store-fileHUGEGRAPH_TRUST_STORE_FILEexport HUGEGRAPH_TRUST_STORE_FILE=/tmp/trust-store
–trust-store-passwordHUGEGRAPH_TRUST_STORE_PASSWORDexport HUGEGRAPH_TRUST_STORE_PASSWORD=xxxx

Another way is to set the environment variable in the bin/hugegraph script:

#!/bin/bash

# Set environment here if needed
#export HUGEGRAPH_URL=
#export HUGEGRAPH_GRAPH=
#export HUGEGRAPH_USERNAME=
#export HUGEGRAPH_PASSWORD=
#export HUGEGRAPH_TIMEOUT=
#export HUGEGRAPH_TRUST_STORE_FILE=
#export HUGEGRAPH_TRUST_STORE_PASSWORD=

bin/hugegraph also reads JAVA_HOME (a warning is printed when it is not set, and it is needed for https) and JAVA_OPTIONS (JVM options; when it is empty the script uses -Xms512m plus an -Xmx computed from the free memory of the machine).

3.3 Graph Management Type, graph-mode-set, graph-mode-get, graph-list, graph-get, graph-clear, graph-create, graph-clone and graph-drop
  • graph-mode-set, set graph restore mode
    • –graph-mode or -m, required, specifies the mode to be set, legal values include [NONE, RESTORING, MERGING, LOADING]
  • graph-mode-get, get graph restore mode
  • graph-list, list all graphs in a HugeGraph-Server
  • graph-get, get a graph and its storage backend type
  • graph-clear, clear all schema and data of a graph
    • –confirm-message or -c, required, delete confirmation information, manual input is required, double confirmation to prevent accidental deletion, “I’m sure to delete all data”, including double quotes
  • graph-create, create a new graph with configuration file
    • –name or -n, optional, the name of the new graph, default is g
    • –file or -f, the path to the graph configuration file, the content of the file is sent to HugeGraph-Server as the config of the new graph
  • graph-clone, clone an existing graph
    • –name or -n, optional, the name of the cloned graph, default is g
    • –clone-graph-name, optional, the name of the source graph to clone from, default is hugegraph
  • graph-drop, drop a graph (different from graph-clear, this completely removes the graph)
    • –confirm-message or -c, required, confirmation message “I’m sure to drop the graph”, including double quotes

graph-create, graph-clone, graph-clear and graph-drop raise –timeout to at least 300 seconds.

When you need to restore the backup graph to a new graph, you need to set the graph mode to RESTORING mode; when you need to merge the backup graph into an existing graph, you need to first set the graph mode to MERGING model.

3.4 Asynchronous task management Type,task-list、task-get、task-delete、task-cancel and task-clear
  • task-list,List the asynchronous tasks in a graph, which can be filtered according to the status of the tasks
    • –status,Optional, specify the status of the task to view, i.e. filter tasks by status, legal values include [UNKNOWN, NEW, QUEUED, RESTORING, RUNNING, SUCCESS, CANCELLED, FAILED] (case insensitive)
    • –limit,Optional, specify the number of tasks to be obtained, the default is -1, which means to obtain all eligible tasks, a value passed explicitly must be positive
  • task-get,Get detailed information about an asynchronous task
    • –task-id,Required, specifies the ID of the asynchronous task
  • task-delete,Delete information about an asynchronous task
    • –task-id,Required, specifies the ID of the asynchronous task
  • task-cancel,Cancel the execution of an asynchronous task
    • –task-id,Required, the ID of the asynchronous task to cancel
  • task-clear,Clean up completed asynchronous tasks
    • –force,Optional. When set, it means to clean up all asynchronous tasks. Unfinished ones are canceled first, and then all asynchronous tasks are cleared. By default, only completed asynchronous tasks are cleaned up
3.5 Gremlin Type,gremlin-execute and gremlin-schedule

⚠️ SEC Reminder: The execution of Gremlin depends on the actual logic of the statements, which may involve scenarios such as large-scale data modification and high-risk system calls with potential implicit hazards. Please use this tool only in secure and trusted network environments. It is imperative to configure and secure HugeGraph-Server with the Authentication System (Auth) and an IP Whitelist to restrict execution requests on the server side. Never hand over the tool or expose the execution entry to unauthorized personnel.

  • gremlin-execute, send Gremlin statements to HugeGraph-Server to execute query or modification operations, execute synchronously, and return results after completion
    • –file or -f, specify the script file to execute, UTF-8 encoding, mutually exclusive with –script
    • –script or -s, specifies the script string to execute, mutually exclusive with –file
    • –aliases or -a, Gremlin alias settings, the format is: key1=value1,key2=value2,…
    • –bindings or -b, Gremlin binding settings, the format is: key1=value1,key2=value2,…
    • –language or -l, the language of the Gremlin script, the default is gremlin-groovy

    –file and –script are mutually exclusive, one of them must be set

  • gremlin-schedule, send Gremlin statements to HugeGraph-Server to perform query or modification operations, asynchronous execution, and return the asynchronous task id immediately after the task is submitted
    • –file or -f, specify the script file to execute, UTF-8 encoding, mutually exclusive with –script
    • –script or -s, specifies the script string to execute, mutually exclusive with –file
    • –bindings or -b, Gremlin binding settings, the format is: key1=value1,key2=value2,…
    • –language or -l, the language of the Gremlin script, the default is gremlin-groovy

    –file and –script are mutually exclusive, one of them must be set

3.6 Backup/Restore Type
  • backup, back up the schema or data in a certain graph out of the HugeGraph system, and store it on the local disk or HDFS in the form of JSON
    • –format, the backup format, optional values include [json, text], the default is json
    • –all-properties, whether to back up all properties of vertices/edges, only valid when –format is text, default false
    • –label, the vertex label or edge label to be backed up, only applied when –format is text; when it is set, –huge-types must name exactly one type and that type must be vertex or edge, otherwise the command fails
    • –properties, properties of vertices/edges to be backed up, separated by commas, only valid when –format is text, valid only when backing up vertices or edges
    • –compress, whether to compress data during backup, the default is true
    • –directory or -d, the directory to store schema or data, the default is ‘./{graphName}’ for local directory, and ‘{fs.default.name}/{graphName}’ for HDFS
    • –huge-types or -t, the data types to be backed up, separated by commas, the optional value is ‘all’ or a combination of one or more [vertex, edge, vertex_label, edge_label, property_key, index_label], ‘all’ Represents all 6 types, namely vertices, edges and all schemas, ‘schema’ represents the 4 schema types [vertex_label, edge_label, property_key, index_label]
    • –log or -l, specify the log directory, the default is ./logs
    • –retry, specify the number of failed retries, the default is 3
    • –thread-num or -T, the number of threads to use, default is Math.min(10, Math.max(4, CPUs / 2))
    • –split-size or -s, specifies the size of splitting vertices or edges when backing up, the default is 1048576, and it must be at least 1048576 (1M)
    • -D, use the mode of -Dkey=value to specify dynamic parameters, and specify HDFS configuration items when backing up data to HDFS, for example: -Dfs.default.name=hdfs://localhost:9000

    If –timeout is less than 120 seconds, backup (and the backup step of migrate) uses 120 seconds

  • restore, restore schema or data stored in JSON format to a new graph (RESTORING mode) or merge into an existing graph (MERGING mode)
    • –directory or -d, the directory to store schema or data, the default is ‘./{graphName}’ for local directory, and ‘{fs.default.name}/{graphName}’ for HDFS
    • –clean, whether to delete the directory specified by –directory after the recovery map is completed, the default is false
    • –huge-types or -t, data types to restore, separated by commas, optional value is ‘all’ or a combination of one or more [vertex, edge, vertex_label, edge_label, property_key, index_label], ‘all’ Represents all 6 types, namely vertices, edges and all schemas, ‘schema’ represents the 4 schema types [vertex_label, edge_label, property_key, index_label]
    • –log or -l, specify the log directory, the default is ./logs
    • –retry, specify the number of failed retries, the default is 3
    • –thread-num or -T, the number of threads to use, default is Math.min(10, Math.max(4, CPUs / 2))
    • -D, use the mode of -Dkey=value to specify dynamic parameters, which are used to specify HDFS configuration items when restoring graphs from HDFS, for example: -Dfs.default.name=hdfs://localhost:9000

    restore command can be used only if –format is executed as backup for json restore requires the graph to be in RESTORING or MERGING mode (set it with graph-mode-set first), otherwise the command fails

  • migrate, migrate the currently connected graph to another HugeGraphServer
    • –target-graph, the name of the target graph, the default is hugegraph
    • –target-url, the HugeGraphServer where the target graph is located, the default is http://127.0.0.1:8081
    • –target-user, the username used to access the target graph
    • –target-password, the password to access the target map
    • –target-timeout, the timeout for accessing the target map
    • –target-trust-store-file, access the truststore file used by the target graph
    • –target-trust-store-password, the password to access the truststore used by the target map
    • –directory or -d, during the migration process, the directory where the schema or data of the source graph is stored. For a local directory, the default is ‘./{graphName}’; for HDFS, the default is ‘{fs.default.name}/ {graphName}’
    • –huge-types or -t, the data types to be migrated, separated by commas, the optional value is ‘all’ or a combination of one or more [vertex, edge, vertex_label, edge_label, property_key, index_label], ‘all’ Represents all 6 types, namely vertices, edges and all schemas, ‘schema’ represents the 4 schema types [vertex_label, edge_label, property_key, index_label]
    • –log or -l, specify the log directory, the default is ./logs
    • –retry, specify the number of failed retries, the default is 3
    • –thread-num or -T, the number of threads to use, default is Math.min(10, Math.max(4, CPUs / 2))
    • –split-size or -s, specify the size of the vertex or edge block when backing up the source graph during the migration process, the default is 1048576, and it must be at least 1048576 (1M)
    • -D, use the mode of -Dkey=value to specify dynamic parameters, which are used to specify HDFS configuration items when the data needs to be backed up to HDFS during the migration process, for example: -Dfs.default.name=hdfs://localhost: 9000
    • –graph-mode or -m, the mode to set the target graph when restoring the source graph to the target graph, legal values include [RESTORING, MERGING], the default is RESTORING. The target graph is switched to this mode during the migration and switched back to its original mode afterwards
    • –keep-local-data, whether to keep the backup of the source map generated in the process of migrating the map, the default is false, that is, the backup of the source map is not kept after the default migration map ends
  • schedule-backup, periodically back up the graph and keep a certain number of the latest backups (currently only supports local file systems)
    • –directory or -d, required, specifies the directory of the backup data
    • –backup-num, optional, specifies the number of latest backups to save, defaults to 3
    • –interval, an optional item, specifies the backup cycle, the format is the same as the Linux crontab format, the default is “0 0 * * *” (every day at 00:00)

    schedule-backup adds a crontab entry that runs backup -t all into {directory}/{graph}/hugegraph-backup-{yyMMddHHmm}/ and keeps only the latest –backup-num backups. A relative –directory is resolved against the hugegraph-tools home directory, and {directory}/{graph} must not exist yet

  • dump, export all vertices and edges in the graph, using the vertex vertex-edge1 vertex-edge2... JSON format by default. To customize the format, implement a Formatter subclass such as CustomFormatter under hugegraph-tools/src/main/java/org/apache/hugegraph/formatter, then select it when running the command: bin/hugegraph dump -f CustomFormatter
    • –formatter or -f, specify the formatter to use, the default is JsonFormatter
    • –directory or -d, the directory where schema or data is stored, the default is ‘./{graphName}’ for local directory, and ‘{fs.default.name}/{graphName}’ for HDFS
    • –log or -l, specify the log directory, the default is ./logs
    • –retry, specify the number of failed retries, the default is 3
    • –thread-num or -T, the number of threads to use, default is Math.min(10, Math.max(4, CPUs / 2))
    • –split-size or -s, specifies the size of splitting vertices or edges when backing up, the default is 1048576, and it must be at least 1048576 (1M)
    • -D, use the mode of -Dkey=value to specify dynamic parameters, and specify HDFS configuration items when backing up data to HDFS, for example: -Dfs.default.name=hdfs://localhost:9000
3.7 Authentication data backup/restore type
  • auth-backup, backup authentication data to a specified directory
    • –types or -t, types of authentication data to back up, separated by commas, optional value is ‘all’ or a combination of one or more [user, group, target, belong, access], ‘all’ represents all 5 types; ‘belong’ requires ‘user’ and ‘group’ to be included, ‘access’ requires ‘group’ and ’target’ to be included
    • –directory, directory to store backup data, the default is ‘./auth-backup-restore’ for local directory, and ‘{fs.default.name}/auth-backup-restore’ for HDFS (this option has no -d short form)
    • –retry, specify the number of failed retries, the default is 3
    • -D, use the mode of -Dkey=value to specify dynamic parameters, and specify HDFS configuration items when backing up data to HDFS, for example: -Dfs.default.name=hdfs://localhost:9000
  • auth-restore, restore authentication data from a specified directory
    • –types or -t, types of authentication data to restore, separated by commas, optional value is ‘all’ or a combination of one or more [user, group, target, belong, access], ‘all’ represents all 5 types; ‘belong’ requires ‘user’ and ‘group’ to be included, ‘access’ requires ‘group’ and ’target’ to be included
    • –directory, directory where backup data is stored, the default is ‘./auth-backup-restore’ for local directory, and ‘{fs.default.name}/auth-backup-restore’ for HDFS (this option has no -d short form)
    • –retry, specify the number of failed retries, the default is 3
    • –strategy, conflict handling strategy, optional values are [stop, ignore], default is stop. stop means stop restoring when encountering conflicts, ignore means ignore conflicts and continue restoring
    • –init-password, initial password to set when restoring users, required when –types includes user
    • -D, use the mode of -Dkey=value to specify dynamic parameters, which are used to specify HDFS configuration items when restoring data from HDFS, for example: -Dfs.default.name=hdfs://localhost:9000
3.8 Install the deployment type
  • deploy, one-click download, install and start HugeGraph-Server and HugeGraph-Studio
    • -v, required, specifies the HugeGraph-Server and HugeGraph-Studio version to install, must be one of the versions listed in bin/version-map.yaml (0.6, 0.7, 0.8, 0.9, 0.10), which maps it to the matching server and studio release versions
    • -p, required, specifies the installed HugeGraph-Server and HugeGraph-Studio directories
    • -u, optional, specifies the link to download the HugeGraph-Server and HugeGraph-Studio compressed packages
  • clear, clean up HugeGraph-Server and HugeGraph-Studio directories and tarballs (refuses to run while a matching server or studio process is still alive, and prompts before each removal)
    • -p, required, specifies the directory of HugeGraph-Server and HugeGraph-Studio to be cleaned
  • start-all, start HugeGraph-Server and HugeGraph-Studio with one click
    • -v, required, specifies the installed HugeGraph-Server and HugeGraph-Studio version to start, same values as deploy
    • -p, required, specifies the directory where HugeGraph-Server and HugeGraph-Studio are installed
  • stop-all, close HugeGraph-Server and HugeGraph-Studio with one click

deploy, start-all, clear and stop-all are handed by bin/hugegraph straight to the shell scripts bin/deploy.sh, bin/start-all.sh, bin/clear.sh and bin/stop-all.sh, so the global options and environment variables in 3.2 do not apply to them.

There is an optional parameter -u in the deploy command. When provided, the specified download address will be used instead of the default download address to download the tar package, and the address will be written into the ~/hugegraph-download-url-prefix file; if no address is specified later When -u and ~/hugegraph-download-url-prefix are not specified, the tar package will be downloaded from the address specified by ~/hugegraph-download-url-prefix; if there is neither -u nor ~/hugegraph-download-url-prefix, it will be downloaded from the default download address https://github.com/hugegraph

3.9 Specific command parameters

The specific parameters of each subcommand are as follows:

Usage: hugegraph [options] [command] [command options]
  Options:
    --graph
      Name of graph
      Default: hugegraph
    --password
      Password of user
    --throw-mode
      Whether the hugegraph-tools work to throw an exception
      Default: false
    --timeout
      Connection timeout
      Default: 30
    --trust-store-file
      The path of client truststore file used when https protocol is enabled
    --trust-store-password
      The password of the client truststore file used when the https protocol 
      is enabled
    --url
      The URL of HugeGraph-Server
      Default: http://127.0.0.1:8080
    --user
      Name of user
  Commands:
    graph-create      Create graph with config
      Usage: graph-create [options]
        Options:
          --file, -f
            Creating graph config file
          --name, -n
            The name of new created graph, default is g
            Default: g

    graph-clone      Clone graph
      Usage: graph-clone [options]
        Options:
          --clone-graph-name
            The name of cloned graph, default is hugegraph
            Default: hugegraph
          --name, -n
            The name of new created graph, default is g
            Default: g

    graph-list      List all graphs
      Usage: graph-list

    graph-get      Get graph info
      Usage: graph-get

    graph-clear      Clear graph schema and data
      Usage: graph-clear [options]
        Options:
        * --confirm-message, -c
            Confirm message of graph clear is "I'm sure to delete all data". 
            (Note: include "")

    graph-drop      Drop graph
      Usage: graph-drop [options]
        Options:
        * --confirm-message, -c
            Confirm message of graph clear is "I'm sure to drop the graph". 
            (Note: include "")

    graph-mode-set      Set graph mode
      Usage: graph-mode-set [options]
        Options:
        * --graph-mode, -m
            Graph mode, include: [NONE, RESTORING, MERGING]
            Possible Values: [NONE, RESTORING, MERGING, LOADING]

    graph-mode-get      Get graph mode
      Usage: graph-mode-get

    task-list      List tasks
      Usage: task-list [options]
        Options:
          --limit
            Limit number, no limit if not provided
            Default: -1
          --status
            Status of task

    task-get      Get task info
      Usage: task-get [options]
        Options:
        * --task-id
            Task id
            Default: 0

    task-delete      Delete task
      Usage: task-delete [options]
        Options:
        * --task-id
            Task id
            Default: 0

    task-cancel      Cancel task
      Usage: task-cancel [options]
        Options:
        * --task-id
            Task id
            Default: 0

    task-clear      Clear completed tasks
      Usage: task-clear [options]
        Options:
          --force
            Force to clear all tasks, cancel all uncompleted tasks firstly, 
            and delete all completed tasks
            Default: false

    gremlin-execute      Execute Gremlin statements
      Usage: gremlin-execute [options]
        Options:
          --aliases, -a
            Gremlin aliases, valid format is: 'key1=value1,key2=value2...'
            Default: {}
          --bindings, -b
            Gremlin bindings, valid format is: 'key1=value1,key2=value2...'
            Default: {}
          --file, -f
            Gremlin Script file to be executed, UTF-8 encoded, exclusive to 
            --script 
          --language, -l
            Gremlin script language
            Default: gremlin-groovy
          --script, -s
            Gremlin script to be executed, exclusive to --file

    gremlin-schedule      Execute Gremlin statements as asynchronous job
      Usage: gremlin-schedule [options]
        Options:
          --bindings, -b
            Gremlin bindings, valid format is: 'key1=value1,key2=value2...'
            Default: {}
          --file, -f
            Gremlin Script file to be executed, UTF-8 encoded, exclusive to 
            --script 
          --language, -l
            Gremlin script language
            Default: gremlin-groovy
          --script, -s
            Gremlin script to be executed, exclusive to --file

    backup      Backup graph schema/data. If directory is on HDFS, use -D to 
            set HDFS params. For example: 
            -Dfs.default.name=hdfs://localhost:9000 
      Usage: backup [options]
        Options:
          --all-properties
            All properties to be backup flag
            Default: false
          --compress
            compress flag
            Default: true
          --directory, -d
            Directory of graph schema/data, default is './{graphname}' in 
            local file system or '{fs.default.name}/{graphname}' in HDFS
          --format
            File format, valid is [json, text]
            Default: json
          --huge-types, -t
            Type of schema/data. Concat with ',' if more than one. Other types 
            include 'all' and 'schema'. 'all' means all vertices, edges and 
            schema. In other words, 'all' equals with 'vertex, edge, 
            vertex_label, edge_label, property_key, index_label'. 'schema' 
            equals with 'vertex_label, edge_label, property_key, index_label'.
            Default: [PROPERTY_KEY, VERTEX_LABEL, EDGE_LABEL, INDEX_LABEL, VERTEX, EDGE]
          --label
            Vertex label or edge label, only valid when type is vertex or edge
          --log, -l
            Directory of log
            Default: ./logs
          --properties
            Vertex or edge properties to backup, only valid when type is 
            vertex or edge
            Default: []
          --retry
            Retry times, default is 3
            Default: 3
          --split-size, -s
            Split size of shard
            Default: 1048576
          --thread-num, -T
            Threads number to use, default is Math.min(10, Math.max(4, CPUs / 
            2)) 
            Default: 0
          -D
            HDFS config parameters
            Syntax: -Dkey=value
            Default: {}

    schedule-backup      Schedule backup task
      Usage: schedule-backup [options]
        Options:
          --backup-num
            The number of latest backups to keep
            Default: 3
        * --directory, -d
            The directory of backups stored
          --interval
            The interval of backup, format is: "a b c d e". 'a' means minute 
            (0 - 59), 'b' means hour (0 - 23), 'c' means day of month (1 - 
            31), 'd' means month (1 - 12), 'e' means day of week (0 - 6) 
            (Sunday=0), "*" means all
            Default: "0 0 * * *"

    dump      Dump graph to files
      Usage: dump [options]
        Options:
          --directory, -d
            Directory of graph schema/data, default is './{graphname}' in 
            local file system or '{fs.default.name}/{graphname}' in HDFS
          --formatter, -f
            Formatter to customize format of vertex/edge
            Default: JsonFormatter
          --log, -l
            Directory of log
            Default: ./logs
          --retry
            Retry times, default is 3
            Default: 3
          --split-size, -s
            Split size of shard
            Default: 1048576
          --thread-num, -T
            Threads number to use, default is Math.min(10, Math.max(4, CPUs / 
            2)) 
            Default: 0
          -D
            HDFS config parameters
            Syntax: -Dkey=value
            Default: {}

    restore      Restore graph schema/data. If directory is on HDFS, use -D to 
            set HDFS params if needed. For 
            example:-Dfs.default.name=hdfs://localhost:9000 
      Usage: restore [options]
        Options:
          --clean
            Whether to remove the directory of graph data after restored
            Default: false
          --directory, -d
            Directory of graph schema/data, default is './{graphname}' in 
            local file system or '{fs.default.name}/{graphname}' in HDFS
          --huge-types, -t
            Type of schema/data. Concat with ',' if more than one. Other types 
            include 'all' and 'schema'. 'all' means all vertices, edges and 
            schema. In other words, 'all' equals with 'vertex, edge, 
            vertex_label, edge_label, property_key, index_label'. 'schema' 
            equals with 'vertex_label, edge_label, property_key, index_label'.
            Default: [PROPERTY_KEY, VERTEX_LABEL, EDGE_LABEL, INDEX_LABEL, VERTEX, EDGE]
          --log, -l
            Directory of log
            Default: ./logs
          --retry
            Retry times, default is 3
            Default: 3
          --thread-num, -T
            Threads number to use, default is Math.min(10, Math.max(4, CPUs / 
            2)) 
            Default: 0
          -D
            HDFS config parameters
            Syntax: -Dkey=value
            Default: {}

    migrate      Migrate graph
      Usage: migrate [options]
        Options:
          --directory, -d
            Directory of graph schema/data, default is './{graphname}' in 
            local file system or '{fs.default.name}/{graphname}' in HDFS
          --graph-mode, -m
            Mode used when migrating to target graph, include: [RESTORING, 
            MERGING] 
            Default: RESTORING
            Possible Values: [NONE, RESTORING, MERGING, LOADING]
          --huge-types, -t
            Type of schema/data. Concat with ',' if more than one. Other types 
            include 'all' and 'schema'. 'all' means all vertices, edges and 
            schema. In other words, 'all' equals with 'vertex, edge, 
            vertex_label, edge_label, property_key, index_label'. 'schema' 
            equals with 'vertex_label, edge_label, property_key, index_label'.
            Default: [PROPERTY_KEY, VERTEX_LABEL, EDGE_LABEL, INDEX_LABEL, VERTEX, EDGE]
          --keep-local-data
            Whether to keep the local directory of graph data after restored
            Default: false
          --log, -l
            Directory of log
            Default: ./logs
          --retry
            Retry times, default is 3
            Default: 3
          --split-size, -s
            Split size of shard
            Default: 1048576
          --target-graph
            The name of target graph to migrate
            Default: hugegraph
          --target-password
            The password of target graph to migrate
          --target-timeout
            The timeout to connect target graph to migrate
            Default: 0
          --target-trust-store-file
            The trust store file of target graph to migrate
          --target-trust-store-password
            The trust store password of target graph to migrate
          --target-url
            The url of target graph to migrate
            Default: http://127.0.0.1:8081
          --target-user
            The username of target graph to migrate
          --thread-num, -T
            Threads number to use, default is Math.min(10, Math.max(4, CPUs / 
            2)) 
            Default: 0
          -D
            HDFS config parameters
            Syntax: -Dkey=value
            Default: {}

    deploy      Install HugeGraph-Server and HugeGraph-Studio
      Usage: deploy [options]
        Options:
        * -p
            Install path of HugeGraph-Server and HugeGraph-Studio
          -u
            Download url prefix path of HugeGraph-Server and HugeGraph-Studio
        * -v
            Version of HugeGraph-Server and HugeGraph-Studio

    start-all      Start HugeGraph-Server and HugeGraph-Studio
      Usage: start-all [options]
        Options:
        * -p
            Install path of HugeGraph-Server and HugeGraph-Studio
        * -v
            Version of HugeGraph-Server and HugeGraph-Studio

    clear      Clear HugeGraph-Server and HugeGraph-Studio
      Usage: clear [options]
        Options:
        * -p
            Install path of HugeGraph-Server and HugeGraph-Studio

    stop-all      Stop HugeGraph-Server and HugeGraph-Studio
      Usage: stop-all

    auth-backup      null
      Usage: auth-backup [options]
        Options:
          --directory
            Directory of auth information, default is 
            './{auth-backup-restore}' in local file system or 
            '{fs.default.name}/{auth-backup-restore}' in HDFS
          --retry
            Retry times, default is 3
            Default: 3
          --types, -t
            Type of auth data to restore and backup, concat with ',' if more 
            than one. 'all' means all auth information. In other words, 'all' 
            equals with 'user, group, target, belong, access'. In addition, 
            'belong' or 'access' can not backup or restore alone, if type 
            contains 'belong' then should contains 'user' and 'group'. If type 
            contains 'access' then should contains 'group' and 'target'.
            Default: [TARGET, GROUP, USER, ACCESS, BELONG]
          -D
            HDFS config parameters
            Syntax: -Dkey=value
            Default: {}

    auth-restore      null
      Usage: auth-restore [options]
        Options:
          --directory
            Directory of auth information, default is 
            './{auth-backup-restore}' in local file system or 
            '{fs.default.name}/{auth-backup-restore}' in HDFS
          --init-password
            Init user password, if restore type include 'user', please specify 
            the init-password of users.
            Default: <empty string>
          --retry
            Retry times, default is 3
            Default: 3
          --strategy
            The strategy needs to be chosen in the event of a conflict when 
            restoring. Valid strategies include 'stop' and 'ignore', default 
            is 'stop'. 'stop' means if there a conflict, stop restore. 
            'ignore' means if there a conflict, ignore and continue to 
            restore. 
            Default: STOP
            Possible Values: [STOP, IGNORE]
          --types, -t
            Type of auth data to restore and backup, concat with ',' if more 
            than one. 'all' means all auth information. In other words, 'all' 
            equals with 'user, group, target, belong, access'. In addition, 
            'belong' or 'access' can not backup or restore alone, if type 
            contains 'belong' then should contains 'user' and 'group'. If type 
            contains 'access' then should contains 'group' and 'target'.
            Default: [TARGET, GROUP, USER, ACCESS, BELONG]
          -D
            HDFS config parameters
            Syntax: -Dkey=value
            Default: {}

    help      Print usage
      Usage: help
3.10 Specific command example
1. gremlin statement
# Execute gremlin synchronously
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph gremlin-execute --script 'g.V().count()'

# Execute gremlin asynchronously
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph gremlin-schedule --script 'g.V().count()'
2. Show task status
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph task-list

./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph task-list --limit 5

./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph task-list --status success
3. Set and show graph mode
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph graph-mode-set -m RESTORING

./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph graph-mode-get

./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph graph-list
4. Cleanup Graph
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph graph-clear -c "I'm sure to delete all data"
5. Backup Graph
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph backup -t all --directory ./backup-test
6. Periodic Backup Graph
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph schedule-backup -d ./backup --interval "*/2 * * * *"
7. Recovery Graph
# set graph mode
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph graph-mode-set -m RESTORING

# recovery graph
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph restore -t all --directory ./backup-test

# restore graph mode
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph graph-mode-set -m NONE
8. Graph Migration
./bin/hugegraph --url http://127.0.0.1:8080 --graph hugegraph migrate --target-url http://127.0.0.1:8090 --target-graph hugegraph

7 - HugeGraph-Spark-Connector Quick Start

1 HugeGraph-Spark-Connector Overview

HugeGraph-Spark-Connector uses the Spark DataFrame API to write bulk data to HugeGraph. The current implementation provides vertex and edge writers.

Reading from HugeGraph is not implemented yet: the table only implements SupportsWrite, so spark.read.format(...) is not supported. The connector supports the CUSTOMIZE and PRIMARY_KEY vertex id strategies; the AUTOMATIC strategy is rejected.

2 Environment Requirements

  • Java 8+
  • Maven 3.6+
  • Spark 3.2.x (the module is built against Spark 3.2.2 with provided scope, so your Spark runtime must supply the Spark jars)
  • Scala 2.12 (built with Scala 2.12.11)

3 Building

3.1 Build without executing tests

git clone https://github.com/apache/hugegraph-toolchain.git
cd hugegraph-toolchain
mvn clean package -pl hugegraph-spark-connector -am -DskipTests -ntp

3.2 Build with default tests

mvn clean package -pl hugegraph-spark-connector -am -ntp

Both commands produce a fat jar at hugegraph-spark-connector/target/hugegraph-spark-connector-${revision}-jar-with-dependencies.jar (Spark itself is not bundled). Pass it to spark-submit --jars when you do not manage the dependency through Maven.

4 Usage

Add the dependency to pom.xml, replacing ${revision} with the release version you use:

<dependency>
    <groupId>org.apache.hugegraph</groupId>
    <artifactId>hugegraph-spark-connector</artifactId>
    <version>${revision}</version>
</dependency>

The format string must be the full class name org.apache.hugegraph.spark.connector.DataSource; the connector does not register a short name with Spark’s DataSourceRegister service loader. When HugeGraphServer has authentication enabled, add .option("username", ...) and .option("token", ...) to the examples below.

4.1 Schema Definition Example

If we have a graph, the schema is defined as follows:

schema.propertyKey("name").asText().ifNotExist().create()
schema.propertyKey("age").asInt().ifNotExist().create()
schema.propertyKey("city").asText().ifNotExist().create()
schema.propertyKey("weight").asDouble().ifNotExist().create()
schema.propertyKey("lang").asText().ifNotExist().create()
schema.propertyKey("date").asText().ifNotExist().create()
schema.propertyKey("price").asDouble().ifNotExist().create()

schema.vertexLabel("person")
        .properties("name", "age", "city")
        .useCustomizeStringId()
        .nullableKeys("age", "city")
        .ifNotExist()
        .create()

schema.vertexLabel("software")
        .properties("name", "lang", "price")
        .primaryKeys("name")
        .ifNotExist()
        .create()

schema.edgeLabel("knows")
        .sourceLabel("person")
        .targetLabel("person")
        .properties("date", "weight")
        .ifNotExist()
        .create()

schema.edgeLabel("created")
        .sourceLabel("person")
        .targetLabel("software")
        .properties("date", "weight")
        .ifNotExist()
        .create()

4.2 Vertex Sink (Scala)

val df = sparkSession.createDataFrame(Seq(
  Tuple3("marko", 29, "Beijing"),
  Tuple3("vadas", 27, "HongKong"),
  Tuple3("Josh", 32, "Beijing"),
  Tuple3("peter", 35, "ShangHai"),
  Tuple3("li,nary", 26, "Wu,han"),
  Tuple3("Bob", 18, "HangZhou"),
)) toDF("name", "age", "city")

df.show()

df.write
  .format("org.apache.hugegraph.spark.connector.DataSource")
  .option("host", "127.0.0.1")
  .option("port", "8080")
  .option("graph", "hugegraph")
  .option("data-type", "vertex")
  .option("label", "person")
  .option("id", "name")
  .option("batch-size", 2)
  .mode(SaveMode.Overwrite)
  .save()

4.3 Edge Sink (Scala)

val df = sparkSession.createDataFrame(Seq(
  Tuple4("marko", "vadas", "20160110", 0.5),
  Tuple4("peter", "Josh", "20230801", 1.0),
  Tuple4("peter", "li,nary", "20130220", 2.0)
)).toDF("source", "target", "date", "weight")

df.show()

df.write
  .format("org.apache.hugegraph.spark.connector.DataSource")
  .option("host", "127.0.0.1")
  .option("port", "8080")
  .option("graph", "hugegraph")
  .option("data-type", "edge")
  .option("label", "knows")
  .option("source-name", "source")
  .option("target-name", "target")
  .option("batch-size", 2)
  .mode(SaveMode.Overwrite)
  .save()

4.4 Vertex Sink with PRIMARY_KEY id strategy (Scala)

For a vertex label that uses primaryKeys(...), do not set the id option: the id is spliced from the primary key columns. Columns that are not part of the schema can be dropped with ignored-fields.

val df = sparkSession.createDataFrame(Seq(
  Tuple4("lop", "java", 328L, "ISBN978-7-107-18618-5"),
  Tuple4("ripple", "python", 199L, "ISBN978-7-100-13678-5"),
)).toDF("name", "lang", "price", "ISBN")

df.write
  .format("org.apache.hugegraph.spark.connector.DataSource")
  .option("host", "127.0.0.1")
  .option("port", "8080")
  .option("graph", "hugegraph")
  .option("data-type", "vertex")
  .option("label", "software")
  .option("ignored-fields", "ISBN")
  .option("batch-size", 2)
  .mode(SaveMode.Overwrite)
  .save()

4.5 Edge Sink with mixed id strategies (Scala)

source-name and target-name follow the id strategy of their own vertex label. Below, person uses a customized string id (one column) while software uses a primary key (its name column):

val df = sparkSession.createDataFrame(Seq(
  Tuple4("marko", "lop", "20171210", 0.5),
  Tuple4("Josh", "lop", "20091111", 0.4),
  Tuple4("peter", "ripple", "20171210", 1.0),
  Tuple4("vadas", "lop", "20171210", 0.2)
)).toDF("source", "name", "date", "weight")

df.write
  .format("org.apache.hugegraph.spark.connector.DataSource")
  .option("host", "127.0.0.1")
  .option("port", "8080")
  .option("graph", "hugegraph")
  .option("data-type", "edge")
  .option("label", "created")
  .option("source-name", "source") // customize id
  .option("target-name", "name")   // primary key
  .option("batch-size", 2)
  .mode(SaveMode.Overwrite)
  .save()

Note on save modes: SaveMode.Overwrite and SaveMode.Append both insert the rows. The overwrite path does not delete existing data from the graph first.

5 Configuration Parameters

Option keys are matched case-insensitively and trimmed. data-type and label are always required; source-name and target-name are required when data-type is edge; all other options have defaults.

5.1 Client Configs

Client Configs are used to configure hugegraph-client.

ParameterDefault ValueDescription
hostlocalhostAddress of HugeGraphServer. A bare host name or IP, or a full http:// / https:// prefix
port8080Port of HugeGraphServer
graphhugegraphGraph name
protocolhttpProtocol for sending requests to the server, optional http or https
usernamenullUsername of the current graph when HugeGraphServer enables permission authentication. When unset, the graph name is used as the username
tokennullToken of the current graph when HugeGraphServer has enabled authorization authentication
timeout60Timeout (seconds) for inserting results to return
max-connCPUS * 4The maximum number of HTTP connections between HugeClient and HugeGraphServer
max-conn-per-routeCPUS * 2The maximum number of HTTP connections for each route between HugeClient and HugeGraphServer
trust-store-filenullThe client’s certificate file path when the request protocol is https. When unset under https, the connector reads conf/hugegraph.truststore under the directory given by the JVM system property connector.home.path, which must then be set
trust-store-tokennullThe client’s certificate password when the request protocol is https. When unset under https, hugegraph is used

5.2 Graph Data Configs

Graph Data Configs describe how DataFrame columns map to vertices or edges.

ParameterDefault ValueDescription
data-typeRequired. Graph data type, must be vertex or edge
labelRequired. Label to which the vertex/edge data to be imported belongs
idSpecify a column as the id column of the vertex. When the vertex id policy is CUSTOMIZE, it is required; when the id policy is PRIMARY_KEY, it must be empty. The AUTOMATIC id policy is not supported
source-nameRequired when data-type is edge. Select certain columns of the input source as the id column of source vertex. When the id policy of the source vertex is CUSTOMIZE, a certain column must be specified as the id column of the vertex; when the id policy of the source vertex is PRIMARY_KEY, one or more columns must be specified for splicing the id of the generated vertex, that is, no matter which id strategy is used, this item is required. Multiple columns are separated by , (the delimiter option does not apply here)
target-nameRequired when data-type is edge. Specify certain columns as the id columns of target vertex, similar to source-name
selected-fieldsSelect some columns to insert, other unselected ones are not inserted, cannot exist at the same time as ignored-fields
ignored-fieldsIgnore some columns so that they do not participate in insertion, cannot exist at the same time as selected-fields
batch-size500The number of data items in each batch when importing data. Applied per Spark task: each partition writer flushes its buffer to the server once it holds this many vertices/edges, and again at commit for the remainder

5.3 Common Configs

Common Configs contains some common configurations.

ParameterDefault ValueDescription
delimiter,Separator of selected-fields and ignored-fields. source-name and target-name are always split on ,

6 Notes and Limitations

  • Each Spark write task opens its own HugeClient, switches the graph to LOADING mode before writing and sets it back to NONE at commit or abort.
  • Vertex ids are limited to 128 bytes (UTF-8). This applies to customized string ids and to ids spliced from primary keys.
  • The AUTOMATIC vertex id strategy is not supported; the write fails with an IllegalArgumentException when the writer is created.
  • Properties with SET or LIST cardinality are not supported yet; only SINGLE cardinality values are converted.
  • Date properties: string values must use the format yyyy-MM-dd HH:mm:ss and are parsed in the GMT+8 time zone; numeric values are treated as epoch milliseconds.
  • Boolean properties given as strings accept true, 1, yes, y and false, 0, no, n (case-insensitive).
  • Rows whose customized string id, or any primary key value, is an empty string are skipped. A null id or primary key value raises an error instead.

7 License

The same as HugeGraph, hugegraph-spark-connector is also licensed under Apache 2.0 License.