HugeGraph-Vermeer Quick Start
1. Overview of Vermeer
1.1 Architecture
Vermeer is a high-performance, memory-first graph computing framework written in Go: start once and execute repeatedly. It supports fast execution of 15+ OLAP algorithms, often in seconds to minutes; actual time depends on graph size, algorithm parameters, and resources. A single master currently schedules multiple workers.
The master handles communication, forwarding, and aggregation, with modest computation and resource usage. Workers store graph data and execute tasks, consuming most memory and CPU. gRPC handles internal communication; REST provides external APIs.
At startup, built-in defaults are loaded first, followed by the [default] section of config/<env>.ini under the working directory, then explicit command-line overrides. For example, --env=master reads config/master.ini. The Docker image copies repository configuration to /go/bin/config/ and works from /go/bin/. A host bind mount hides the bundled files, so it must contain every ini file used. The configuration reader does not read environment variables.
Default ports:
| Role | Configuration key | Default address | Purpose |
|---|---|---|---|
| master | http_peer | 0.0.0.0:6688 | REST API for host-side clients |
| master | grpc_peer | 0.0.0.0:6689 | Workers connect to master |
| worker | http_peer | 0.0.0.0:6788 | Worker HTTP service |
| worker | grpc_peer | 0.0.0.0:6789 | Worker gRPC listener, also advertised to master for node communication |
The Docker examples publish only master HTTP on host loopback 127.0.0.1:6688:6688; master and worker gRPC ports stay within the container network.
1.2 Running Vermeer
In production, enable Server authentication and authorization, an IP allowlist, and minimum permissions; retain audit-*.log. Server Auth does not protect independent Vermeer, PD, or Store APIs. Restrict these HTTP/gRPC ports to trusted networks and callers, and configure access controls at the Vermeer external entry point.
master.ini defaults to auth=none, disabling authentication for ordinary and administrative APIs. Local examples publish loopback only. Before remote access, enable auth=token or restrict callers through a protected network or gateway.
Both Docker options need a host configuration directory. From the Vermeer repository root, copy the supplied master.ini and worker.ini templates. The mount hides image configuration under /go/bin/config, so never mount an empty directory or your entire home directory there:
Keep master HTTP/gRPC listeners at 0.0.0.0:6688 and 0.0.0.0:6689. The Compose example assigns 172.20.0.10 to master and 172.20.0.11 to worker; set these options in the copied worker.ini:
This single-worker example uses literal worker_group=$, the general group used when no named group is bound. The repository template defaults to worker_group=default; if retaining that named group, bind it to the task space or graph before submitting tasks. For default space $DEFAULT, use:
An errcode=0 response confirms binding; with token authentication, also supply the authorization header. master_peer must reach master gRPC from the worker network. grpc_peer is both the listener and advertised worker address, so do not advertise 0.0.0.0, which other containers cannot reach. Update these IPs when changing the subnet or addresses; each additional worker needs a unique, mutually reachable grpc_peer.
- Option 1: Docker Compose (Recommended)
Run from the Vermeer repository root. Modify existing docker-compose.yaml or use the following example. The repository file currently mounts all of ~/ as configuration and does not publish master HTTP; replace the mount and add the port mapping before starting:
Replace /home/user/vermeer-config with your actual absolute configuration directory. Changing the subnet or static IPs of vermeer_network also requires updating grpc_peer and master_peer in worker.ini. Do not use the default grpc_peer=0.0.0.0:6789 as an advertised container address.
Build and start from the project directory:
View logs / stop / remove:
- Option 2: Separate docker run Commands (Manual Network and Static IPs)
Set CONFIG_DIR to the prepared configuration directory, with grpc_peer and master_peer matching the static container addresses. Replace the example CONFIG_DIR with your actual absolute path.
Build the image:
Create a custom bridge network (once):
Run master (use your absolute CONFIG_DIR and adjust IPs as needed):
Run worker:
View logs / stop / remove:
- Option 3: Build from Source
Build following the Vermeer README.
Start from the Vermeer root directory with ./vermeer --env=master and ./vermeer --env=worker01. In worker01.ini, set grpc_peer to an address bindable locally and reachable by master and other workers; point master_peer to master gRPC.
After starting master, check its HTTP port from the host:
Expect HTTP 200 and JSON errcode=0.
2. Task Creation REST API
2.1 Introduction
Submit a load task, wait for loading to finish, then submit a compute task. A loaded graph can support repeated computations and is not deleted after completion. Asynchronous APIs return creation results and task information, including ID, before completion; synchronous APIs wait for success or failure, so client and proxy HTTP timeouts must be sufficiently long. Query task states: loaded for successful loading, complete for successful computation, error for failure, and canceled for cancellation; other states remain waiting or running. Graphs loading or in an error state cannot be computed. Deletion requires a deletable graph state and no current usage.
Available URLs:
- Asynchronous:
POST http://master_ip:port/tasks/create; read the ID from responsetask.id. - Synchronous:
POST http://master_ip:port/tasks/create/sync; returns after the task ends. - Query task:
GET http://master_ip:port/task/{task_id}.errcode=0indicates a successful query; inspecttask.statefor task status. Set a client polling deadline for asynchronous jobs; reaching it stops client waiting without canceling the server task.
2.2 Loading Graph Data
The examples list common load parameters. Each loader reads its own keys; unused keys do not change behavior.
Vermeer provides three loading methods:
- Local files
load.vertex_files and load.edge_files map worker address host portions to file paths read by those workers. With containers, mount data into the worker and use container paths. In the Compose example, mount a host data directory at worker /data (such as - /host/data:/data:ro) and use 172.20.0.11 from grpc_peer as the mapping key. Other deployments use the host portion of their advertised worker addresses.
Obtain a dataset such as Twitter-2010; the first twitter-2010.txt.gz file is sufficient.
Request example:
- HugeGraph
Request example:
Replace request addresses, graph name, and credentials with your actual connection settings.
Vermeer master connects to PD through load.hg_pd_peers to look up partitions; workers read data from the returned Store addresses. These services must be reachable from the corresponding Vermeer hosts or containers. Inside Docker, 127.0.0.1 identifies the container itself, not the host or another container.
- HDFS
Request example:
2.3 Outputting Computation Results
Current result writers support local, hdfs, and hugegraph through output.type; none disables output. Source retains an afs constant, but current master registers no AFS loader or writer. Set output.need_statistics=1 to put statistics into task information; supported statistics operators depend on each algorithm implementation.
The examples list common computation and output parameters; supported algorithm parameters depend on current Vermeer implementations.
Request example:
output.type=local writes to the executing worker’s local filesystem. For containers, mount the output directory if the host needs to read results.
3. Supported Algorithms
3.1 PageRank
The PageRank algorithm, also known as the web ranking algorithm, is a technique used by search engines to calculate the relevance and importance of web pages (nodes) based on their mutual hyperlinks.
- If a web page is linked to by many other web pages, it indicates that the web page is relatively important, and its PageRank value will be relatively high.
- If a web page with a high PageRank value links to other web pages, the PageRank value of the linked web pages will also increase accordingly.
The PageRank algorithm is suitable for scenarios such as web page ranking and identifying key figures in social networks.
Request example:
3.2 WCC (Weakly Connected Components)
The weakly connected components algorithm calculates all connected subgraphs in an undirected graph and outputs the weakly connected subgraph ID to which each vertex belongs, indicating the connectivity between points and distinguishing different connected communities.
Request example:
3.3 LPA (Label Propagation Algorithm)
The label propagation algorithm is a graph clustering algorithm commonly used in social networks to discover potential communities.
Request example:
3.4 Degree Centrality
The degree centrality algorithm calculates the degree centrality value of each node in the graph, supporting both undirected and directed graphs. Degree centrality is an important indicator of node importance; the more edges a node has with other nodes, the higher its degree centrality value, and the more important the node is in the graph. In an undirected graph, degree centrality is calculated based on edge information to count the number of times a node appears, resulting in the degree centrality value of the node. In a directed graph, it is based on the direction of the edges, filtering based on input or output-edge information to count the number of times a node appears, resulting in the in-degree or out-degree value of the node. It indicates the importance of each point, with more important points having higher degrees.
Request example:
3.5 Closeness Centrality
Closeness centrality is used to calculate the inverse of the shortest distance from a node to all other reachable nodes, accumulating and normalizing the value. Closeness centrality can be used to measure the time it takes for information to be transmitted from the node to other nodes. The larger the closeness centrality of a node, the closer its position in the graph is to the center, suitable for scenarios such as identifying key nodes in social networks.
Request example:
3.6 Betweenness Centrality
The betweenness centrality algorithm determines the value of a node as a “bridge” node; the larger the value, the more likely it is to be a necessary path between two points in the graph. Typical examples include mutual followers in social networks. It is suitable for measuring the degree of aggregation around a node in a community.
Request example:
3.7 Triangle Count
The triangle count algorithm calculates the number of triangles passing through each vertex, suitable for calculating the relationships between users and whether the associations form triangles. The more triangles, the higher the degree of association between nodes in the graph, and the tighter the organizational relationship. In social networks, triangles indicate cohesive communities, and identifying triangles helps understand clustering and interconnections among individuals or groups in the network. In financial or transaction networks, the presence of triangles may indicate suspicious or fraudulent activities, and triangle counting can help identify transaction patterns that may require further investigation.
The output result is the Triangle Count corresponding to each vertex, i.e., the number of triangles the vertex is part of.
Note: This algorithm is for undirected graphs and ignores edge directions.
Request example:
3.8 K-Core
The K-Core algorithm marks all vertices with a degree of K, suitable for graph pruning and finding the core part of the graph.
Request example:
3.9 SSSP (Single Source Shortest Path)
The single source the shortest path algorithm calculates the shortest distance from one point to all other points.
Request example:
3.10 KOUT
Starting from a point, get the k-layer nodes of this point.
Request example:
3.11 Louvain
The Louvain algorithm is a community detection algorithm based on modularity. The basic idea is that nodes in the network try to traverse all neighbor community labels and choose the community label that maximizes the modularity increment. After maximizing modularity, each community is regarded as a new node, and the process is repeated until the modularity no longer increases.
The distributed Louvain algorithm implemented on Vermeer is affected by factors such as node order and parallel computation. Due to the random traversal order of the Louvain algorithm, community compression also has a certain randomness, leading to different results in multiple executions. However, the overall trend will not change significantly.
Request example:
3.12 Jaccard Similarity Coefficient
The Jaccard index, also known as the Jaccard similarity coefficient, is used to compare the similarity and diversity between finite sample sets. The larger the Jaccard coefficient value, the higher the similarity of the samples. It is used to calculate the Jaccard similarity coefficient between a given source point and all other points in the graph.
Request example:
3.13 Personalized PageRank
The goal of personalized PageRank is to calculate the relevance of all nodes relative to user u. Starting from the node corresponding to user u, at each node, there is a probability of 1-d to stop walking and start again from u, or a probability of d to continue walking, randomly selecting a node from the nodes pointed to by the current node to walk down. It is used to calculate the personalized PageRank score starting from a given starting point, suitable for scenarios such as social recommendations.
Since the calculation requires using out-degree, load.use_out_degree needs to be set to 1 when reading the graph.
Request example:
3.14 Global Kout
Calculate the k-degree neighbors of all nodes in the graph (excluding themselves and 1~k-1 degree neighbors). Due to the severe memory expansion of the global kout algorithm, k is currently limited to 1 and 2. Additionally, the global kout algorithm supports filtering functions (parameters such as “compute.filter”:“risk_level==1”), and the filtering condition is judged when calculating the k-degree. The final result set includes those that meet the filtering condition. The algorithm’s final output is the number of neighbors that meet the condition.
Request example:
3.15 Clustering Coefficient
The clustering coefficient represents the coefficient of the clustering degree of nodes in a graph. In real networks, especially in specific networks, nodes tend to establish a tightly organized relationship due to relatively high-density connection points. The clustering coefficient algorithm (Cluster Coefficient) is used to calculate the clustering degree of nodes in the graph. This algorithm is for local clustering coefficients. The local clustering coefficient can measure the clustering degree around each node in the graph.
Request example:
3.16 SCC (Strongly Connected Components)
In the mathematical theory of directed graphs, if every vertex of a graph can be reached from any other point in the graph, the graph is said to be strongly connected. The parts of any directed graph that can achieve strong connectivity are called strongly connected components. It indicates the connectivity between points and distinguishes different connected communities.
Request example:
🚧, further updates and improvements will be made at any time. Suggestions and feedback are welcome.