Using HBase Through API

Last updated: 2023-12-25 16:20:03
HBase is a highly reliable, high-performance, column-oriented, scalable distributed storage system, serving as an open-source implementation of Google's BigTable. HBase employs Hadoop HDFS as its file storage system, utilizes Hadoop MapReduce to process the vast amount of data within HBase, and uses Zookeeper for coordination services.
HBase primarily consists of Zookeeper, HMaster, and HRegionServer. Specifically:
ZooKeeper mitigates the single point of failure of HMaster, its master election mechanism ensures the provision of service by a single Master.
HMaster manages user operations for adding, deleting, modifying, and querying tables, as well as managing the load balancing of HRegionServers. It can also adjust the distribution of Regions, migrating the HRegions within an exiting HRegionServer to other HRegionServers.
HRegionServer is the most crucial module within HBase, primarily responsible for responding to user I/O requests and reading and writing data to the HDFS file system. Internally, HRegionServer manages a series of HRegion objects, each corresponding to a Region, and composed of multiple Stores within HRegion. Each Store corresponds to the storage of a Column Family.
This development guide, from a technical perspective, will assist users in developing with the EMR cluster. Considering user data security, currently, EMR only supports VPC network access.

1. Development Preparation

Ensure that you have activated Tencent Cloud and have created an EMR cluster. When creating the EMR cluster, you need to select the HBase and ZooKeeper components in the software configuration interface.

2. Utilizing HBase Shell

Before using HBase Shell, please log in to the Master node of the EMR cluster. The method of logging into EMR can be referred to in Log in to the Linux instance. Here, you can choose to log in using WebShell. Click on the login on the right side of the corresponding cloud server to enter the login interface. The username is set to root by default, and the password is the one entered by the user when creating EMR. After entering correctly, you can access the EMR command line interface.
In the EMR command line, first use the following command to switch to the Hadoop user, and enter the directory /usr/local/service/hbase:
[root@172 ~]# su hadoop
[hadoop@10root]$ cd /usr/local/service/hbase
You can enter HBase Shell using the following command:
[hadoop@10hbase]$ bin/hbase shell
In the HBase shell, you can view basic usage information and example commands by typing 'help'. Next, we will use the following command to create a new table:
hbase(main):001:0> create 'test', 'cf'
After the table is established, you can use the list command to check whether the table you created already exists.
hbase(main):002:0> list 'test'
TABLE
test
1 row(s) in 0.0030 seconds

=> ["test"]
Utilize the put command to incorporate elements into the table you have created:
hbase(main):003:0> put 'test', 'row1', 'cf:a', 'value1'
0 row(s) in 0.0850 seconds

hbase(main):004:0> put 'test', 'row2', 'cf:b', 'value2'
0 row(s) in 0.0110 seconds

hbase(main):005:0> put 'test', 'row3', 'cf:c', 'value3'
0 row(s) in 0.0100 seconds
We have incorporated three values into the table we created. Initially, a value "value1" was inserted into the "cf:a" column of the "row1" row, and so forth.
Employ the scan command to traverse the entire table:
hbase(main):006:0> scan 'test'
ROW COLUMN+CELL
row1 column=cf:a, timestamp=1530276759697, value=value1
row2 column=cf:b, timestamp=1530276777806, value=value2
row3 column=cf:c, timestamp=1530276792839, value=value3
3 row(s) in 0.2110 seconds
Use the get command to obtain the value of a specified row in the table:
hbase(main):007:0> get 'test', 'row1'
COLUMN CELL
cf:a timestamp=1530276759697, value=value
1 row(s) in 0.0790 seconds
Use the drop command to delete a table. Prior to deleting a table, it is necessary to first employ the disable command to deactivate a table:
hbase(main):010:0> disable 'test'
hbase(main):011:0> drop 'test'
Finally, you can employ the quit command to close the hbase shell.
For additional Hbase shell commands, please refer to the official documentation.

3. Utilizing Hbase through API

Initially, download and install Maven, configure Maven's environmental variables. If you are using an IDE, please set up the relevant Maven configurations within the IDE.

Create a new Maven project

Navigate to the directory where you wish to create a new project from the command line, for instance, within D://mavenWorkplace, and input the following command to create a new Maven project:
mvn archetype:generate -DgroupId=$yourgroupID -DartifactId=$yourartifactID
-DarchetypeArtifactId=maven-archetype-quickstart
Where $yourgroupID is your package name. $yourartifactID is your project name, and maven-archetype-quickstart indicates the creation of a Maven Java project. Some files need to be downloaded during the project creation process, so please ensure a stable internet connection.
Upon successful creation, a project folder named $yourartifactID will be generated in the D://mavenWorkplace directory. The structure of the files within is as follows:
simple
   ---pom.xml    Core configuration, located at the root of the project
   ---src
     ---main      
       ---java    Directory for Java source code
    ---resources  Directory for Java configuration files
    ---test
      ---java    Directory for test source code
      ---resources  Directory for test configurations
Our primary focus lies on the pom.xml file and the Java folder under main. The pom.xml file is primarily used for dependency and packaging configurations, while the Java folder houses your source code.

Incorporate Hadoop dependencies and sample code

Firstly, incorporate Maven dependencies into the pom.xml file:
<dependencies>
<dependency>
<groupId>org.apache.hbase</groupId>
<artifactId>hbase-client</artifactId>
<version>1.2.4</version>
</dependency>
</dependencies>
Subsequently, incorporate packaging and compilation plugins into the pom.xml file:
<build>
<plugins>
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-compiler-plugin</artifactId>
<configuration>
<source>1.8</source>
<target>1.8</target>
<encoding>utf-8</encoding>
</configuration>
</plugin>
<plugin>
<artifactId>maven-assembly-plugin</artifactId>
<configuration>
<descriptorRefs>
<descriptorRef>jar-with-dependencies</descriptorRef>
</descriptorRefs>
</configuration>
<executions>
<execution>
<id>make-assembly</id>
<phase>package</phase>
<goals>
<goal>single</goal>
</goals>
</execution>
</executions>
</plugin>
</plugins>
</build>
Before adding the sample code, users need to obtain the zookeeper address of the Hbase cluster. Log into any Master or Core node of EMR, navigate to the /usr/local/service/hbase/conf directory, and view the hbase.zookeeper.quorum configuration in hbase-site.xml to obtain the IP address $quorum of zookeeper, the hbase.zookeeper.property.clientPort configuration to obtain the port number $clientPort of zookeeper, and the zookeeper.znode.parent configuration to obtain the znode path $znodePath used by hbase.
Next, add the sample code. Create a new Java Class named PutExample.java in the main>java folder, and incorporate the following code into it:
import org.apache.hadoop.conf.Configuration;
import org.apache.hadoop.hbase.*;
import org.apache.hadoop.hbase.client.*;
import org.apache.hadoop.hbase.util.Bytes;
import org.apache.hadoop.hbase.io.compress.Compression.Algorithm;

import java.io.IOException;

/**
* Created by tencent on 2018/6/30.
*/
public class PutExample {
public static void main(String[] args) throws IOException {
Configuration conf = HBaseConfiguration.create();
conf.set("hbase.zookeeper.quorum","$quorum");
conf.set("hbase.zookeeper.property.clientPort","$clientPort");
conf.set("zookeeper.znode.parent", "$znodePath");

Connection connection = ConnectionFactory.createConnection(conf);
Admin admin = connection.getAdmin();

HTableDescriptor table = new HTableDescriptor(TableName.valueOf("test1"));
table.addFamily(new HColumnDescriptor("cf").setCompressionType(Algorithm.NONE));

System.out.print("Creating table. ");
if (admin.tableExists(table.getTableName())) {
admin.disableTable(table.getTableName());
admin.deleteTable(table.getTableName());
}
admin.createTable(table);

Table table1 = connection.getTable(TableName.valueOf("test1"));
Put put1 = new Put(Bytes.toBytes("row1"));
put1.addColumn(Bytes.toBytes("cf"), Bytes.toBytes("a"),
Bytes.toBytes("value1"));
table1.put(put1);
Put put2 = new Put(Bytes.toBytes("row2"));
put2.addColumn(Bytes.toBytes("cf"), Bytes.toBytes("b"),
Bytes.toBytes("value2"));
table1.put(put2);
Put put3 = new Put(Bytes.toBytes("row3"));
put3.addColumn(Bytes.toBytes("cf"), Bytes.toBytes("c"),
Bytes.toBytes("value3"));
table1.put(put3);

System.out.println(" Done.");
}
}

Compile the code and package it for upload.

Navigate to the project directory using the local command line and execute the following command to compile and package the project:
mvn package
A display of 'build success' indicates a successful operation. The packaged file can be found in the target folder within the project directory.
Use scp or sftp tools to upload the packaged file to the EMR cluster. It is essential to upload the jar package that has been packaged together with the dependencies. Run the following in the local command line mode:
scp $localfile root@PublicIPAddress:$remotefolder
In this context, $localfile refers to the path and name of your local file, 'root' is the username of the CVM server, and the public IP can be viewed in the node information of the EMR console or in the cloud server console. $remotefolder is the path on the CVM server where you wish to store the file. Once the upload is complete, you can check in the EMR cluster command line to see if the corresponding file is in the respective folder.

4. Run the Sample

Log into the Master node of the EMR cluster and switch to the hadoop user. Execute the sample using the following command:
[hadoop@10 hadoop]$ java –jar $package.jar
Upon the console outputting "Done", it indicates that all operations have been completed. You can switch to the hbase shell and use the list command to check whether the Hbase table created using the API was successful. If successful, you can use the scan command to view the specific contents of the table.
[hadoop@10hbase]$ bin/hbase shell
hbase(main):002:0> list 'test1'
TABLE
Test1
1 row(s) in 0.0030 seconds

=> ["test1"]
hbase(main):006:0> scan 'test1'
ROW COLUMN+CELL
row1 column=cf:a, timestamp=1530276759697, value=value1
row2 column=cf:b, timestamp=1530276777806, value=value2
row3 column=cf:c, timestamp=1530276792839, value=value3
3 row(s) in 0.2110 seconds
For more detailed instructions on API usage, please refer to the official documentation.