HBase is a highly reliable, high-performance, column-oriented, scalable distributed storage system, serving as an open-source implementation of Google's BigTable. HBase employs Hadoop HDFS as its file storage system, utilizes Hadoop MapReduce to process the vast amount of data within HBase, and uses Zookeeper for coordination services.
HBase primarily consists of Zookeeper, HMaster, and HRegionServer. Specifically:
ZooKeeper mitigates the single point of failure of HMaster, its master election mechanism ensures the provision of service by a single Master.
HMaster manages user operations for adding, deleting, modifying, and querying tables, as well as managing the load balancing of HRegionServers. It can also adjust the distribution of Regions, migrating the HRegions within an exiting HRegionServer to other HRegionServers.
HRegionServer is the most crucial module within HBase, primarily responsible for responding to user I/O requests and reading and writing data to the HDFS file system. Internally, HRegionServer manages a series of HRegion objects, each corresponding to a Region, and composed of multiple Stores within HRegion. Each Store corresponds to the storage of a Column Family.
This development guide, from a technical perspective, will assist users in developing with the EMR cluster. Considering user data security, currently, EMR only supports VPC network access.
1. Development Preparation
Ensure that you have activated Tencent Cloud and have created an EMR cluster. When creating the EMR cluster, you need to select the HBase and ZooKeeper components in the software configuration interface.
2. Utilizing HBase Shell
Before using HBase Shell, please log in to the Master node of the EMR cluster. The method of logging into EMR can be referred to in Log in to the Linux instance. Here, you can choose to log in using WebShell. Click on the login on the right side of the corresponding cloud server to enter the login interface. The username is set to root by default, and the password is the one entered by the user when creating EMR. After entering correctly, you can access the EMR command line interface.
In the EMR command line, first use the following command to switch to the Hadoop user, and enter the directory
/usr/local/service/hbase:[root@172 ~]# su hadoop[hadoop@10root]$ cd /usr/local/service/hbase
You can enter HBase Shell using the following command:
[hadoop@10hbase]$ bin/hbase shell
In the HBase shell, you can view basic usage information and example commands by typing 'help'. Next, we will use the following command to create a new table:
hbase(main):001:0> create 'test', 'cf'
After the table is established, you can use the
list command to check whether the table you created already exists.hbase(main):002:0> list 'test'TABLEtest1 row(s) in 0.0030 seconds=> ["test"]
Utilize the
put command to incorporate elements into the table you have created:hbase(main):003:0> put 'test', 'row1', 'cf:a', 'value1'0 row(s) in 0.0850 secondshbase(main):004:0> put 'test', 'row2', 'cf:b', 'value2'0 row(s) in 0.0110 secondshbase(main):005:0> put 'test', 'row3', 'cf:c', 'value3'0 row(s) in 0.0100 seconds
We have incorporated three values into the table we created. Initially, a value "value1" was inserted into the "cf:a" column of the "row1" row, and so forth.
Employ the
scan command to traverse the entire table:hbase(main):006:0> scan 'test'ROW COLUMN+CELLrow1 column=cf:a, timestamp=1530276759697, value=value1row2 column=cf:b, timestamp=1530276777806, value=value2row3 column=cf:c, timestamp=1530276792839, value=value33 row(s) in 0.2110 seconds
Use the
get command to obtain the value of a specified row in the table:hbase(main):007:0> get 'test', 'row1'COLUMN CELLcf:a timestamp=1530276759697, value=value1 row(s) in 0.0790 seconds
Use the
drop command to delete a table. Prior to deleting a table, it is necessary to first employ the disable command to deactivate a table:hbase(main):010:0> disable 'test'hbase(main):011:0> drop 'test'
Finally, you can employ the
quit command to close the hbase shell.3. Utilizing Hbase through API
Initially, download and install Maven, configure Maven's environmental variables. If you are using an IDE, please set up the relevant Maven configurations within the IDE.
Create a new Maven project
Navigate to the directory where you wish to create a new project from the command line, for instance, within
D://mavenWorkplace, and input the following command to create a new Maven project:mvn archetype:generate -DgroupId=$yourgroupID -DartifactId=$yourartifactID-DarchetypeArtifactId=maven-archetype-quickstart
Where $yourgroupID is your package name. $yourartifactID is your project name, and maven-archetype-quickstart indicates the creation of a Maven Java project. Some files need to be downloaded during the project creation process, so please ensure a stable internet connection.
Upon successful creation, a project folder named $yourartifactID will be generated in the
D://mavenWorkplace directory. The structure of the files within is as follows:simple---pom.xml Core configuration, located at the root of the project---src---main---java Directory for Java source code---resources Directory for Java configuration files---test---java Directory for test source code---resources Directory for test configurations
Our primary focus lies on the pom.xml file and the Java folder under main. The pom.xml file is primarily used for dependency and packaging configurations, while the Java folder houses your source code.
Incorporate Hadoop dependencies and sample code
Firstly, incorporate Maven dependencies into the pom.xml file:
<dependencies><dependency><groupId>org.apache.hbase</groupId><artifactId>hbase-client</artifactId><version>1.2.4</version></dependency></dependencies>
Subsequently, incorporate packaging and compilation plugins into the pom.xml file:
<build><plugins><plugin><groupId>org.apache.maven.plugins</groupId><artifactId>maven-compiler-plugin</artifactId><configuration><source>1.8</source><target>1.8</target><encoding>utf-8</encoding></configuration></plugin><plugin><artifactId>maven-assembly-plugin</artifactId><configuration><descriptorRefs><descriptorRef>jar-with-dependencies</descriptorRef></descriptorRefs></configuration><executions><execution><id>make-assembly</id><phase>package</phase><goals><goal>single</goal></goals></execution></executions></plugin></plugins></build>
Before adding the sample code, users need to obtain the zookeeper address of the Hbase cluster. Log into any Master or Core node of EMR, navigate to the
/usr/local/service/hbase/conf directory, and view the hbase.zookeeper.quorum configuration in hbase-site.xml to obtain the IP address $quorum of zookeeper, the hbase.zookeeper.property.clientPort configuration to obtain the port number $clientPort of zookeeper, and the zookeeper.znode.parent configuration to obtain the znode path $znodePath used by hbase.Next, add the sample code. Create a new Java Class named PutExample.java in the main>java folder, and incorporate the following code into it:
import org.apache.hadoop.conf.Configuration;import org.apache.hadoop.hbase.*;import org.apache.hadoop.hbase.client.*;import org.apache.hadoop.hbase.util.Bytes;import org.apache.hadoop.hbase.io.compress.Compression.Algorithm;import java.io.IOException;/*** Created by tencent on 2018/6/30.*/public class PutExample {public static void main(String[] args) throws IOException {Configuration conf = HBaseConfiguration.create();conf.set("hbase.zookeeper.quorum","$quorum");conf.set("hbase.zookeeper.property.clientPort","$clientPort");conf.set("zookeeper.znode.parent", "$znodePath");Connection connection = ConnectionFactory.createConnection(conf);Admin admin = connection.getAdmin();HTableDescriptor table = new HTableDescriptor(TableName.valueOf("test1"));table.addFamily(new HColumnDescriptor("cf").setCompressionType(Algorithm.NONE));System.out.print("Creating table. ");if (admin.tableExists(table.getTableName())) {admin.disableTable(table.getTableName());admin.deleteTable(table.getTableName());}admin.createTable(table);Table table1 = connection.getTable(TableName.valueOf("test1"));Put put1 = new Put(Bytes.toBytes("row1"));put1.addColumn(Bytes.toBytes("cf"), Bytes.toBytes("a"),Bytes.toBytes("value1"));table1.put(put1);Put put2 = new Put(Bytes.toBytes("row2"));put2.addColumn(Bytes.toBytes("cf"), Bytes.toBytes("b"),Bytes.toBytes("value2"));table1.put(put2);Put put3 = new Put(Bytes.toBytes("row3"));put3.addColumn(Bytes.toBytes("cf"), Bytes.toBytes("c"),Bytes.toBytes("value3"));table1.put(put3);System.out.println(" Done.");}}
Compile the code and package it for upload.
Navigate to the project directory using the local command line and execute the following command to compile and package the project:
mvn package
A display of 'build success' indicates a successful operation. The packaged file can be found in the target folder within the project directory.
Use scp or sftp tools to upload the packaged file to the EMR cluster. It is essential to upload the jar package that has been packaged together with the dependencies. Run the following in the local command line mode:
scp $localfile root@PublicIPAddress:$remotefolder
In this context, $localfile refers to the path and name of your local file, 'root' is the username of the CVM server, and the public IP can be viewed in the node information of the EMR console or in the cloud server console. $remotefolder is the path on the CVM server where you wish to store the file. Once the upload is complete, you can check in the EMR cluster command line to see if the corresponding file is in the respective folder.
4. Run the Sample
Log into the Master node of the EMR cluster and switch to the hadoop user. Execute the sample using the following command:
[hadoop@10 hadoop]$ java –jar $package.jar
Upon the console outputting "Done", it indicates that all operations have been completed. You can switch to the hbase shell and use the
list command to check whether the Hbase table created using the API was successful. If successful, you can use the scan command to view the specific contents of the table.[hadoop@10hbase]$ bin/hbase shellhbase(main):002:0> list 'test1'TABLETest11 row(s) in 0.0030 seconds=> ["test1"]hbase(main):006:0> scan 'test1'ROW COLUMN+CELLrow1 column=cf:a, timestamp=1530276759697, value=value1row2 column=cf:b, timestamp=1530276777806, value=value2row3 column=cf:c, timestamp=1530276792839, value=value33 row(s) in 0.2110 seconds