Compare commits

...

533 Commits

Author SHA1 Message Date
i-robot e755263a3a
!1531 Fixed 0 row count issue while recovering
Merge pull request !1531 from Kishore Kumar M V/0rows_after_recovery
2022-06-23 06:57:51 +00:00
Kishore 2876fc31f0 Fixed 0 row count issue after recovering 2022-06-23 11:23:37 +05:30
i-robot 8289812e65
!1530 translate docs for new feature
Merge pull request !1530 from tushengxia/master
2022-06-22 11:06:07 +00:00
tushengxia 3c01a18784 translate docs for new feature 2022-06-22 17:57:59 +08:00
i-robot fac10963d5
!1525 add more introcution for openlookeng ranger plugin
Merge pull request !1525 from Maxiaoqi/documents
2022-06-22 09:32:07 +00:00
i-robot 0c6ad6bf69
!1527 [I59M5L] Suspend/Resume Hang Fix and UI changes
Merge pull request !1527 from Surya Sumanth/suspend_resume_hang_and_UI_fix
2022-06-22 09:20:07 +00:00
i-robot a1d0e90407
!1453 【轻量级 PR】:update hetu-docs/zh/admin/properties.md:Hive
Merge pull request !1453 from 暮暮七/N/A
2022-06-22 06:09:46 +00:00
i-robot e911b79869
!1528 [I4Y3TQ] Handling Recovery of Worker - To -Worker Interaction Timeout
Merge pull request !1528 from Surya Sumanth/worker_to_worker_interaction_timeout
2022-06-22 00:37:16 +00:00
Surya Sumanth N 97a3394d97 [I4Y3TQ] Handling Worker to Worker Interaction Timeout 2022-06-21 23:51:41 +05:30
Surya Sumanth N bd2e8edf79 [I59M5L] Suspend/Resume Hang Fix and UI Changes 2022-06-21 23:32:00 +05:30
maxiaoqi2020 155d6c52da update ranger documents 2022-06-21 14:28:08 +08:00
i-robot 6d7c691b7a
!1522 修复openLooKeng社区交流问题
Merge pull request !1522 from DOU/master
2022-06-21 01:26:49 +00:00
DOU 9b49a61f93 Fixed community communication issues 2022-06-20 16:06:14 +08:00
i-robot 01e63fea84
!1514 fixed bug of I5BJ3R and I5BG5H
Merge pull request !1514 from 孙锐/master
2022-06-17 08:03:31 +00:00
i-robot 22b7e688ee
!1517 [I5CM66] Fix Not Serializable Exception for SpilledBlooms
Merge pull request !1517 from i-robot/pull360
2022-06-17 04:51:27 +00:00
i-robot 660d5d6584
!1516 hetu core clean code improment
Merge pull request !1516 from chenpingzeng/clean_code_modify
2022-06-17 01:43:26 +00:00
Surya Sumanth N 2d9b099728 [I5CM66] Fix Not Serializable Exception for SpilledBlooms 2022-06-16 18:34:04 +05:30
chenpingzeng aeecaa8390 hetu core clean code
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-06-16 20:08:17 +08:00
i-robot 509bda89f0
!1507 hetu core upgrade opensource version to solve CVEs
Merge pull request !1507 from chenpingzeng/software_upgrade
2022-06-15 23:01:46 +00:00
sunrui 636bbe7779 fixed bug of I5BJ3R and I5BG5H 2022-06-15 17:34:02 +08:00
chenpingzeng ac5d8d3dca upgrade software dependency to solve CVEs
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-06-15 16:37:25 +08:00
i-robot 4235a0560f
!1511 Configuring Default Spiller Path for spill to hdfs
Merge pull request !1511 from Surya Sumanth/default_spill_hdfs_path
2022-06-11 04:16:40 +00:00
Surya Sumanth N c4f8da5daa Configuring Default Spiller Path for spill to hdfs 2022-06-10 19:38:08 +05:30
i-robot cc14517482
!1509 Fixed infinite loop while refreshing node states before reschedling.
Merge pull request !1509 from i-robot/pull359
2022-06-10 11:37:03 +00:00
i-robot 40303c65a9
!1472 [I58HBD] Support Spill To Hdfs Extension for Snapshot Feature
Merge pull request !1472 from Surya Sumanth/spill_to_hdfs_snapshot
2022-06-10 05:10:35 +00:00
Surya Sumanth N 378fcdd589 Spill To Hdfs Support for Snapshot 2022-06-09 23:56:09 +05:30
Kishore a14c23916e Fixed infinite loop while refreshing node states before reschedling. 2022-06-09 17:57:11 +05:30
i-robot 72b573ef62
!1508 use encodeURIComponent to fix bug of issue I5B9I7
Merge pull request !1508 from 孙锐/master
2022-06-09 11:18:15 +00:00
sunrui f158251296 use encodeURIComponent to fix bug of issue #I5B9I7 2022-06-08 15:51:31 +08:00
i-robot 1547e6b6b1
!1485 fixed bug issues--I59QGY by change queryHistory UI
Merge pull request !1485 from 孙锐/master
2022-06-06 09:35:36 +00:00
sunrui aa8f1e8ba2 fixed bug issues--I59QGY by change queryHistory UI 2022-06-02 09:18:07 +08:00
i-robot 9992c4bf5a
!1481 [I59M5L] Suspend-Resume hang fix
Merge pull request !1481 from i-robot/pull357
2022-06-01 10:23:03 +00:00
i-robot 5dbe49614b
!1484 Fixing issue while recovering from snapshot in HA mode
Merge pull request !1484 from i-robot/pull358
2022-06-01 09:27:03 +00:00
i-robot e3dd0d7847
!1479 To disable unsupported optimizations when snapshot is enabled for query
Merge pull request !1479 from i-robot/pull356
2022-06-01 07:25:43 +00:00
i-robot 57aef4b7d1
!1483 Fixed openlookeng UI dynamically add mongodb catalog
Merge pull request !1483 from lslzz/lslzz001
2022-06-01 02:53:54 +00:00
Kishore 7279350d7a Fixing issue while recovering from snapshot in HA mode 2022-05-31 20:04:12 +05:30
lslzz af5277656d Fixed openlookeng UI dynamically add mongodb catalog 2022-05-31 10:31:57 +08:00
Nitin Kashyap 47f1696ab6
[I59M5L] Resume hang fix 2022-05-31 00:07:57 +05:30
i-robot b3f1c1810f
!1480 fixed bug issue I597T9 by change version of JQuery
Merge pull request !1480 from 孙锐/master
2022-05-28 09:03:49 +00:00
sunrui 95ea5d4b29 fixed bug issues--I597T9 by change version of jQuery 2022-05-28 15:18:58 +08:00
i-robot cfd547b432
!1477 fixed bug of blank ui and hetu.collectionsql.max-count error
Merge pull request !1477 from 孙锐/master
2022-05-27 01:54:44 +00:00
sunrui b6095ab2dd fixed bug of I5995H and I597T9 2022-05-26 14:50:19 +08:00
Kishore 444c2c2819 To disable unsupported optimizations when snapshot is enabled for query 2022-05-26 09:23:34 +05:30
i-robot f9c2ed92f9
!1467 Add read only for file based system access control at catalog level
Merge pull request !1467 from zhaoyaqian/master
2022-05-24 03:18:18 +00:00
i-robot c328d17b1f
!1476 Feature-Suspend, Resume query when low resources
Merge pull request !1476 from Nitin-Kashyap/feature-SuspendResumeQuery
2022-05-23 19:54:32 +00:00
Nitin Kashyap 9a6ad76160
Feature - suspend, and resume memory intensive queries in low memory situation. 2022-05-24 00:31:48 +05:30
i-robot 5803556080
!1473 Refactoring of snapshot config to use recovery framework and to support recovery without snapshot capture
Merge pull request !1473 from Surya Sumanth/recovery_framework
2022-05-23 18:17:29 +00:00
Kishore 3190b64cf8 Refactoring of snapshot config to use recovery framework and to support recovery without snapshot capture 2022-05-23 22:43:35 +05:30
i-robot 3be5c9c199
!1474 Gossip Protocol Review comment fixes
Merge pull request !1474 from ahanapradhan/gossip_fix
2022-05-23 15:17:05 +00:00
Ahana 1697da2f3b review comment fixes
review comment fixes
2022-05-23 17:24:10 +05:30
i-robot d313b8f951
!1471 Branch-KunPengZhongZhi WebUI Enhancement code
Merge pull request !1471 from 孙锐/master
2022-05-23 07:23:06 +00:00
i-robot 7cb87ba6db
!1449 add redis connector
Merge pull request !1449 from shiming/master
2022-05-23 03:48:30 +00:00
i-robot cf7e84b742
!1469 Gossip protocol implementation for failure detection
Merge pull request !1469 from i-robot/pull347
2022-05-20 10:52:54 +00:00
sunrui 1b4c9bc842 Merge branch 'Branch-KunPengZhongzhi'
# Conflicts:
#	hetu-carbondata/pom.xml
#	hetu-clickhouse/pom.xml
#	hetu-common/pom.xml
#	hetu-cube/pom.xml
#	hetu-datacenter/pom.xml
#	hetu-docs/en/admin/web-interface.md
#	hetu-docs/zh/admin/web-interface.md
#	hetu-filesystem-client/pom.xml
#	hetu-function-namespace-managers/pom.xml
#	hetu-greenplum/pom.xml
#	hetu-hana/pom.xml
#	hetu-hazelcast/pom.xml
#	hetu-hbase/pom.xml
#	hetu-heuristic-index/pom.xml
#	hetu-hive-functions/pom.xml
#	hetu-kylin/pom.xml
#	hetu-listener/pom.xml
#	hetu-metastore/pom.xml
#	hetu-mongodb/pom.xml
#	hetu-opengauss/pom.xml
#	hetu-oracle/pom.xml
#	hetu-seed-store/pom.xml
#	hetu-server-rpm/pom.xml
#	hetu-server/pom.xml
#	hetu-sql-migration-tool/pom.xml
#	hetu-startree/pom.xml
#	hetu-state-store/pom.xml
#	hetu-transport/pom.xml
#	hetu-vdm/pom.xml
#	pom.xml
#	presto-array/pom.xml
#	presto-atop/pom.xml
#	presto-base-jdbc/pom.xml
#	presto-benchmark-driver/pom.xml
#	presto-benchmark/pom.xml
#	presto-benchto-benchmarks/pom.xml
#	presto-cli/pom.xml
#	presto-client/pom.xml
#	presto-elasticsearch/pom.xml
#	presto-example-http/pom.xml
#	presto-expressions/pom.xml
#	presto-geospatial-toolkit/pom.xml
#	presto-geospatial/pom.xml
#	presto-hive-hadoop2/pom.xml
#	presto-hive/pom.xml
#	presto-jdbc/pom.xml
#	presto-jmx/pom.xml
#	presto-kafka/pom.xml
#	presto-local-file/pom.xml
#	presto-main/pom.xml
#	presto-main/src/main/java/io/prestosql/catalog/AbstractCatalogStore.java
#	presto-main/src/main/java/io/prestosql/memory/ClusterMemoryManager.java
#	presto-main/src/main/java/io/prestosql/queryeditorui/resources/LoginResource.java
#	presto-main/src/main/resources/webapp/dist/index.js
#	presto-matching/pom.xml
#	presto-memory-context/pom.xml
#	presto-memory/pom.xml
#	presto-ml/pom.xml
#	presto-mysql/pom.xml
#	presto-orc/pom.xml
#	presto-parquet/pom.xml
#	presto-parser/pom.xml
#	presto-password-authenticators/pom.xml
#	presto-plugin-toolkit/pom.xml
#	presto-postgresql/pom.xml
#	presto-product-tests/pom.xml
#	presto-proxy/pom.xml
#	presto-rcfile/pom.xml
#	presto-record-decoder/pom.xml
#	presto-resource-group-managers/pom.xml
#	presto-session-property-managers/pom.xml
#	presto-spi/pom.xml
#	presto-sqlserver/pom.xml
#	presto-teradata-functions/pom.xml
#	presto-testing-docker/pom.xml
#	presto-testing-server-launcher/pom.xml
#	presto-tests/pom.xml
#	presto-thrift-api/pom.xml
#	presto-thrift-testing-server/pom.xml
#	presto-thrift/pom.xml
#	presto-tpcds/pom.xml
#	presto-tpch/pom.xml
#	presto-verifier/pom.xml
2022-05-20 15:44:32 +08:00
sunrui a8939f88f1 Modify Hetu-listener to monitor openLooKeng cluster startup and shutdown,WebUi user login and exit;
show catalog by reading properties file;
make queryInfo persistence by hetu-metadata, collect user sql by hetu-metadata;
Code completion of sql-editor;
add pagenation for AuditLog WebUi;
use airlift.log to enhance hetu-log rather than log4j;
change collect button;
add config for queryHistory-max-count and collectSql-max-count;
2022-05-20 15:39:14 +08:00
Ahana a3446685c2 codecheck
Gossip Protocol for Failure Detection
2022-05-20 12:19:44 +05:30
i-robot 8363281e01
!1460 [I4MHGW] Fix for Query Restore Failing From Successfully Captured Snapshot Issue
Merge pull request !1460 from i-robot/pull338
2022-05-19 17:09:54 +00:00
i-robot 44cbeb8eb6
!1470 BugFix: Fix for immediate query fail due to unhandled exception in httpRequest for resumable failure
Merge pull request !1470 from i-robot/pull351
2022-05-19 06:44:19 +00:00
Ahana ff420218b8 Fix for immediate query fail due to unhandled exception in httpRequest for resumable failure 2022-05-18 17:23:52 +05:30
zhaoyaqian 11cc99b4fe Add read only for file based system access control at catalog level 2022-05-18 09:30:55 +08:00
i-robot 56305dcfc2
!1463 Introducing Failure Retry Policies
Merge pull request !1463 from ahanapradhan/newpr
2022-05-17 10:36:48 +00:00
Ahana a36635f676 Introducing failure retry profiles 2022-05-17 11:36:22 +05:30
i-robot e21b036efb
!1465 Revert the pr:1430
Merge pull request !1465 from zengchen1024/revert-merge-1430-master
2022-05-16 07:52:18 +00:00
tushengxia c17eaf8491
回退 'Pull Request !1430 : Add read only for file based system access control at catalog level' 2022-05-16 06:16:15 +00:00
qzweng 7cb0b9c017
!1430 Add read only for file based system access control at catalog level
Merge pull request !1430 from zhaoyaqian/master
2022-05-16 03:06:07 +00:00
chenshiming2 590d62637a add redis connector 2022-05-12 22:47:02 +08:00
zhousipei f4a150c07f amend docs about extension execution planner 2022-05-11 10:27:43 +08:00
i-robot 6b005c02f0
!1462 fix featuredQueries json deserialization bug
Merge pull request !1462 from tianyi.tu/master
2022-05-10 12:07:50 +00:00
Surya Sumanth N 40f7d57104 [I4MHGW] Fix for Query Restore Failing From Successfully Captured Snapshot 2022-05-09 11:23:38 +05:30
tianyitu 2d16d74cbe 【bugfix】 Fixed featuredQueries json deserialization bug. 2022-05-05 15:16:55 +08:00
i-robot bb76006d46
!1458 add release note for 1.6.1
Merge pull request !1458 from tushengxia/releasenotes1.6.1
2022-04-27 10:26:29 +00:00
tushengxia 4d9d15e28e add release note for 1.6.1 2022-04-27 15:26:08 +08:00
i-robot 842339fc8e
!1457 [343] Extend Snapshot Support for Spilling in LookUpJoinOperator
Merge pull request !1457 from i-robot/pull344
2022-04-27 05:55:08 +00:00
i-robot 8a4b071a13
!1456 adapt to module presto hive function namespace
Merge pull request !1456 from wyy566/branch-1.7
2022-04-26 11:29:32 +00:00
wyy566 ce641847a4 adapt to module presto hive function namespace 2022-04-26 16:49:24 +08:00
i-robot d946b22c19
!1450 add API to get statistics from pageSource
Merge pull request !1450 from guojunfei399/master
2022-04-25 06:07:35 +00:00
暮暮七 95802f2fe1
update hetu-docs/zh/admin/properties.md:Hive
在大数据语境中Hive一般不译为“蜂巢”,并且对应的英文本句含有2个“supports”导致本句翻译有些异常,建议修改。
2022-04-24 05:58:10 +00:00
i-robot 6a79256341
!1447 [I4Y3TQ] handle exception in httpRequest for resumable failure
Merge pull request !1447 from i-robot/pull336
2022-04-22 08:32:53 +00:00
i-robot 12d82cf262
!1446 Recovery state displayed in CLI during snapshot restore
Merge pull request !1446 from i-robot/pull327
2022-04-21 10:23:30 +00:00
Qiaa-n d0f22c8f0b modify field accessMode 2022-04-21 14:32:59 +08:00
guojunfei 28091e6e3d add API to get statistics from pageSource 2022-04-20 19:55:28 +08:00
i-robot d69ea6e00f
!1448 Added null check to avoid crash in unusual flows
Merge pull request !1448 from i-robot/pull337
2022-04-20 10:51:51 +00:00
Kishore 4976377fc7 Snapshot capture/restore information displayed in CLI while running in debug mode 2022-04-20 16:08:32 +05:30
Nitin Kashyap e864bb6aed
[I4Y3TQ] defect fix - handle exception in httpRequest for resumable failure. 2022-04-20 15:50:42 +05:30
i-robot a76d1b6246
!1444 Add docs about extension execution planner
Merge pull request !1444 from zhousipei/add_docs
2022-04-20 07:17:51 +00:00
zhousipei fe8ceb70c0 add docs about extension execution planner 2022-04-19 17:08:48 +08:00
i-robot db1c66c8e5
!1436 support omniruntime
Merge pull request !1436 from zhousipei/support_omniruntime
2022-04-16 07:15:09 +00:00
zhousipei 4b0edc0804 support omniruntime 2022-04-15 11:37:50 +08:00
i-robot 6ff4d8264c
!1439 support querying hudi tables configured with kerberos
Merge pull request !1439 from 建康/master
2022-04-14 12:07:32 +00:00
Surya Sumanth N b433fda9c3 Extend Snapshot Support for Spilling in LookUpJoinOperator 2022-04-14 11:19:14 +05:30
lijiankang 96d30d62e9 support querying hudi tables configured with kerberos 2022-04-12 16:19:41 +08:00
i-robot 7ad71942a1
!1392 [I4V5KY] mysql and postgresql connector do not support create or drop schema
Merge pull request !1392 from futureltl/master
2022-04-01 03:39:05 +00:00
i-robot 69424f8064
!1431 [master][1.6.0RC5]修正docs文档
Merge pull request !1431 from DOU/master
2022-03-30 09:57:51 +00:00
DOU 4714c7702b 修正docs文档 2022-03-30 17:13:37 +08:00
Kishore 8271999e4a Added null check to avoid crash in unusual flows 2022-03-30 09:30:41 +05:30
Raghunandan 630bf19d05 [maven-release-plugin] prepare for next development iteration 2022-03-30 09:14:32 +05:30
Raghunandan f3a95a750f [maven-release-plugin] prepare branch branch-1.6 2022-03-30 09:14:31 +05:30
i-robot f306ab666f
!1429 update release note for 1.6.0
Merge pull request !1429 from tushengxia/releasenotes1.6.0
2022-03-30 01:45:52 +00:00
Qiaa-n 1bea8956d5 Add read only for file based system access control at catalog level 2022-03-29 21:52:59 +08:00
tushengxia fab527d82e add release notes for 1.6.0 2022-03-29 14:31:33 +08:00
i-robot 1f178dac3a
!1428 hetu core clean code
Merge pull request !1428 from chenpingzeng/clean_code_modify
2022-03-27 01:52:23 +00:00
chenpingzeng 94a944cb7b clean code optimize
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-03-26 14:41:08 +08:00
i-robot 2a834c4797
!1427 [I4XU9C] Empty Spilled partition identification for outer fixed
Merge pull request !1427 from i-robot/pull334
2022-03-25 13:03:07 +00:00
Nitin Kashyap 019561dffb
[I4XU9C] defect fix - partition identification of Spilled pages outer tracking corrected. 2022-03-25 17:52:35 +05:30
i-robot adc323fb7b
!1425 Too low value of exchange.max-retry-count doesn't take …
Merge pull request !1425 from i-robot/pull333
2022-03-24 16:43:07 +00:00
Ahana 86411fa5a7 Too low value of exchange.max-retry-count doesn't take effect 2022-03-24 20:47:21 +05:30
i-robot 5475f9b7d5
!1424 [I4PIKD] Document update stating spilling not supported for cross join
Merge pull request !1424 from i-robot/pull332
2022-03-24 14:37:08 +00:00
i-robot 829dfa4bd2
!1423 [I4WGE1] Fix Resetting Retry Count on Restore Success
Merge pull request !1423 from i-robot/pull331
2022-03-24 13:45:09 +00:00
Surya Sumanth N dd905b18a1 [I4PIKD] Document update stating spilling not supported for cross join 2022-03-24 18:41:39 +05:30
i-robot d1d87d04f9
!1422 [I44QYL] Fix for Hive Split Source is already closed Issue
Merge pull request !1422 from i-robot/pull328
2022-03-24 12:01:08 +00:00
Surya Sumanth N ccfbee3d7e [I4ZB03] Fix Resetting Retry Count on Restore Success 2022-03-24 13:58:10 +05:30
Surya Sumanth N 5a66ab4154 [I44QYL] Fix for Hive Split Source is already closed Issue 2022-03-23 16:03:19 +05:30
i-robot fc320c73fa
!1420 upgrade hazelcast from v4.0.3 to v5.1 to solve CVE
Merge pull request !1420 from chenpingzeng/software_upgrade
2022-03-22 02:47:55 +00:00
chenpingzeng 97d0355587 upgrade hazelcast from 4.0.3 to 5.1
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-03-22 10:19:51 +08:00
i-robot 698c4570df
!1421 [I4Y3TQ] Fix Polling Workers Sync Issue For Resume Flow
Merge pull request !1421 from i-robot/pull324
2022-03-22 01:53:56 +00:00
Surya Sumanth N b3973fcb82 Fix Polling Workers Sync Issue For Resume Flow 2022-03-22 01:30:14 +05:30
i-robot be22256fb6
!1417 Reload Cube support without any predicate
Merge pull request !1417 from mahtabahmed/issueFix
2022-03-19 15:59:15 +00:00
mahtabahmed 7892307723 reload cube fixed without predicate 2022-03-18 13:42:35 -04:00
i-robot 4ea79e57ca
!1419 Synchronize Cube metadata updates to handle concurrent inserts into Cube.
Merge pull request !1419 from sundarannamalai/cube-0322
2022-03-16 20:57:51 +00:00
Sundar Annamalai 7f4991d50b Synchronize Cube metadata updates to streamline concurrent inserts into Cube. 2022-03-16 15:42:47 -04:00
i-robot cbbd6eb82a
!1414 [I4XGEK] Considering query restart also as restore
Merge pull request !1414 from i-robot/pull318
2022-03-15 16:06:31 +00:00
i-robot 9ef5bd8547
!1416 [I4XU9C] fixed partition identification of Spilled outer tracker
Merge pull request !1416 from Nitin-Kashyap/defectFix-OuterMismatchOnSpill
2022-03-15 13:30:37 +00:00
i-robot 9e946f8d23
!1413 Fixed updating the restore CPU time properly
Merge pull request !1413 from i-robot/pull315
2022-03-15 11:56:34 +00:00
i-robot ed9a1160ac
!1415 [I4VX53] Skipping deletion of the mergefiles incase of cancel-to-resume
Merge pull request !1415 from i-robot/pull321
2022-03-15 11:02:33 +00:00
i-robot 69b713cd7b
!1410 issue# 307 fix: Failure detection by maximum retry fail doesnt take the exact value set in exchange.max-retry-count config parameter
Merge pull request !1410 from i-robot/pull308
2022-03-15 09:08:32 +00:00
Nitin Kashyap 99f83ad087
[I4XU9C] fixed partition identification of Spilled pages outer tracking corrected. 2022-03-15 14:12:31 +05:30
Surya Sumanth N 2df33c2010 [I4VX53] Skipping deletion of the mergefiles incase of race-condition between coordinator cancel-to-resume flow 2022-03-15 10:26:11 +05:30
i-robot 54154ba4fe
!1412 Defect fix spill bloom finish after probe
Merge pull request !1412 from i-robot/pull317
2022-03-15 02:52:32 +00:00
Kishore 1ec4eb22d4 Considering query restart also as restore 2022-03-14 17:59:02 +05:30
Nitin Kashyap bc4c1d9b4f
DefectFix - Blooms for spilled Hash partition might not have finished on finish spilling as more data may be expected on build side. 2022-03-14 17:18:31 +05:30
i-robot fdd3acfdad
!1409 Fixed the execution of Create Cube command with Where clause twice
Merge pull request !1409 from mahtabahmed/issueFix
2022-03-14 11:45:16 +00:00
i-robot bb91e034f1
!1408 upgrade spring-framework to solve CVEs
Merge pull request !1408 from chenpingzeng/software_upgrade
2022-03-14 10:49:16 +00:00
i-robot 7e43ba1ae2
!1411 fix:The specified HDFS client options are not loaded when using federated HDFS or NameNode high availability;
Merge pull request !1411 from 建康/master
2022-03-14 03:37:16 +00:00
lijiankang 4b6e9c34e9 fix:https://gitee.com/openlookeng/hetu-core/issues/I4XMKD?from=project-issue#note_9188748_link 2022-03-14 09:37:39 +08:00
Kishore e387e4f8a2 Fixed updating the restore CPU time properly 2022-03-12 17:45:27 +05:30
mahtabahmed 3ca8144eb3 fixed create cube with where twice 2022-03-10 10:48:52 -05:00
Ahana 5df46ebd06 [307] Failure detection by maximum retry fail doesnt take the exact value set in exchange.max-retry-count config parameter 2022-03-10 20:33:27 +05:30
chenpingzeng bda86acbd0 upgrade spring-framework to solve CVE problems
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-03-10 19:31:24 +08:00
i-robot 441b414a83
!1407 Fix Kryo Serialization Issue for empty VariableWidthBlock
Merge pull request !1407 from i-robot/pull306
2022-03-09 13:28:47 +00:00
i-robot b21aca43a9
!1405 translate new feature docs
Merge pull request !1405 from tushengxia/translate-docs
2022-03-09 12:14:47 +00:00
i-robot 3b2506101d
!1406 Fix For Kryo Serialization Issue - fieldBlockOffsets length less than positionCount
Merge pull request !1406 from SubhraJyotiBaroi/kryo-spill-issue
2022-03-09 06:22:51 +00:00
i-robot ca94f1ce20
!1399 Update version of log4j 2 to latest to resolve security vulnerabilities
Merge pull request !1399 from chenpingzeng/log4j_upgrade
2022-03-09 04:20:48 +00:00
i-robot 2468b57afc
!1398 Update version of jquery to 3.5.1 above to resolve security vulnerabilities
Merge pull request !1398 from chenpingzeng/jquery_upgrade
2022-03-09 02:31:29 +00:00
i-robot 220278a774
!1404 Fixed Startree issues and Partition-group overlap
Merge pull request !1404 from mahtabahmed/issueFix
2022-03-08 21:19:24 +00:00
mahtabahmed f1c2ab096d fixed partition-group overlap with unit test and Startree issues 2022-03-08 12:10:58 -05:00
SJBaroi 0dedd94e31 Fix For Kryo Serialization Issue - fieldBlockOffsets. 2022-03-07 19:46:54 +05:30
tushengxia 5011b0dfd4 translate hetu-docs 2022-03-07 19:59:51 +08:00
i-robot 07d57b6872
!1402 Removing nodeId for creating HDFS spill subdirectories.
Merge pull request !1402 from SubhraJyotiBaroi/hdfs-spill
2022-03-05 01:17:25 +00:00
SJBaroi 9fe785ba53 Removing nodeId for creating HDFS spill subdirectories. 2022-03-05 02:30:55 +05:30
i-robot bef93ec616
!1403 Updated 'varchar' predicate limitation in Cube documentation
Merge pull request !1403 from sundarannamalai/cube-0322
2022-03-04 20:15:23 +00:00
Sundar Annamalai 75ecc511cf Add varchar predicate limitation in cube documentation 2022-03-04 12:15:58 -05:00
i-robot bffc6d375f
!1401 Corrected spelling mistakes in snapshot doc
Merge pull request !1401 from i-robot/pull304
2022-03-04 15:59:24 +00:00
Surya Sumanth N 3cce55738c Fix Kryo Serialization Issue for empty VariableWidthBlock 2022-03-04 17:44:54 +05:30
Kishore 3aed01e48b Corrected spelling mistakes in snapshot doc 2022-03-04 13:04:20 +05:30
i-robot d739daa154
!1396 Adding Configuration To Support Spilling In HDFS.
Merge pull request !1396 from SubhraJyotiBaroi/hdfs-spill
2022-03-03 16:44:30 +00:00
SJBaroi 3617b76600 Adding Configuration To Support Spilling In HDFS. 2022-03-03 20:36:15 +05:30
i-robot 0e8c11e532
!1400 Updating snapshot documentation with statastics details
Merge pull request !1400 from i-robot/pull303
2022-03-03 11:58:28 +00:00
i-robot 56b4eec973
!1395 Spiller for join and using Spilled buildSide for RightOuter queries
Merge pull request !1395 from i-robot/pull301
2022-03-03 11:12:29 +00:00
Kishore 1ca96b8ab6 Updating snapshot documentation with statastics details 2022-03-03 15:31:49 +05:30
Nitin Kashyap 0c18f876bd
SpilledJoinOptimizations, spiller blooms for eliminating spill probe
Added support for right outer scan when Build side spills.
2022-03-03 14:37:36 +05:30
chenpingzeng 9fc9a1c516 log4j upgrade to latest version
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-03-03 10:32:17 +08:00
chenpingzeng 509b8850a4 update jquery and Bootstrap to 0 CVE version
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-03-03 10:31:26 +08:00
i-robot 7c498e0c43
!1397 updated ui execution timeout to 100 days, also updated document
Merge pull request !1397 from i-robot/pull298
2022-03-02 13:22:29 +00:00
i-robot 94e52ebd1f
!1393 Changes to record snapshot capture metrics
Merge pull request !1393 from i-robot/pull291
2022-03-02 08:24:31 +00:00
Kishore 2ff0fb2f30 Changes to record snapshot capture metrics 2022-03-02 10:47:21 +05:30
i-robot 2e513c05c2
!1394 Document update to show how to enable asynchronous spill mechanism for order by.
Merge pull request !1394 from i-robot/pull296
2022-03-02 02:51:12 +00:00
i-robot 8dcf0bbf02
!1390 [I4V0HU] handle exclusion of incomplete spill files for snapshots
Merge pull request !1390 from Nitin-Kashyap/snapshot-for-asyncOrderBySpill
2022-03-02 02:09:08 +00:00
Nitin Kashyap cfe310724c
handle exclusion of incomplete spill files for snapshots 2022-03-01 15:33:55 +05:30
aloknath 396fce3419 updated ui execution timeout to 100 days, also updated document 2022-02-24 19:02:43 +05:30
liangtl 6b84dcace0 presto-base-jdbc支持create&drop schema 2022-02-24 17:41:48 +08:00
SJBaroi 6fd046f332 Document update for enabling asynchronous spill mechanism for order by. 2022-02-24 11:39:48 +05:30
i-robot 9f1c40ca2b
!1387 #292 #294 updated the ui document to show execution timeout property with its d…
Merge pull request !1387 from i-robot/pull295
2022-02-23 21:58:53 +00:00
i-robot 76b90b6e06
!1391 Doc changes for Reload cube and Show create cube command
Merge pull request !1391 from mahtabahmed/docUpdate
2022-02-23 21:24:53 +00:00
mahtabahmed 1cc1fc46fe doc changes for RELOAD CUBE and SHOW CREATE CUBE 2022-02-23 15:11:10 -05:00
i-robot 9999c198f0
!1389 Adding Secondary Spilling For OrderByOperator.
Merge pull request !1389 from i-robot/pull289
2022-02-22 17:16:52 +00:00
SJBaroi 7e4713a4db Adding Secondary Spilling For OrderByOperator. 2022-02-22 21:49:45 +05:30
i-robot ffcb38b3ad
!1388 Fault detection
Merge pull request !1388 from i-robot/pull287
2022-02-22 15:56:54 +00:00
aloknath 87bdb97308 updated the ui document to show execution timeout property with its default value 2022-02-22 20:50:05 +05:30
Ahana 9bd193238f failure-detection changes
PR review comment fixes
2022-02-22 18:21:01 +05:30
i-robot dd7f7fa8b1
!1386 [I4UMZH] Kryo Serialization Integration for Snapshot
Merge pull request !1386 from Surya Sumanth/kryo_serialization_snapshot
2022-02-22 06:06:53 +00:00
Surya Sumanth N e87f2a08d3 [I4UMZH] Kryo Serialization Integration For Snapshot 2022-02-22 10:51:14 +05:30
i-robot b62a9fe636
!1385 restore code logic of getRandomString for external function register
Merge pull request !1385 from chenpingzeng/clean_code_modify
2022-02-22 02:20:57 +00:00
chenpingzeng 3464f93d22 restore logic of getRandomString for external function register
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-02-21 20:00:48 +08:00
i-robot a9a27d895b
!1369 added Kryo Integration for Spiller Serialization
Merge pull request !1369 from i-robot/pull278
2022-02-21 11:57:34 +00:00
i-robot 8bc16aef96
!1383 Copyright Header check added for 2022
Merge pull request !1383 from Nitin-Kashyap/copyright-2022
2022-02-21 09:53:34 +00:00
Nitin Kashyap f0bbae9b92
[I4UJJB] Copyright header check added for 2022 2022-02-21 12:25:00 +05:30
Nitin Kashyap a50d60b2ed
[277] added Kryo Integration for Spiller Serialization 2022-02-21 09:57:54 +05:30
i-robot 247d7850e6
!1374 Adding RELOAD CUBE [CUBENAME] command
Merge pull request !1374 from mahtabahmed/test
2022-02-19 02:29:55 +00:00
mahtabahmed 479e1a36a1 support for reload cube 2022-02-18 16:48:32 -05:00
i-robot 68d84d95b4
!1382 hetu-core clean code modify
Merge pull request !1382 from chenpingzeng/clean_code_modify
2022-02-18 07:27:14 +00:00
chenpingzeng a4bdcadb3e hetu-core clean code
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-02-18 14:57:36 +08:00
i-robot 1512e1adaf
!1380 Support PostgreSQL and openGauss Update/Delete
Merge pull request !1380 from Anllick/openGaussPgUpdate
2022-02-17 11:18:37 +00:00
i-robot a031b83a1e
!1381 hetu-core clean code modify
Merge pull request !1381 from chenpingzeng/clean_code_modify
2022-02-17 09:36:38 +00:00
chenpingzeng fd85d47d53 hetu clean code
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-02-17 16:45:52 +08:00
Anllick 64e21895a0 Support PostgreSQL and openGauss Update/Delete 2022-02-17 14:16:56 +08:00
i-robot 40730650cd
!1375 Cube Support for 'varchar' range predicate
Merge pull request !1375 from sundarannamalai/cube-0322
2022-02-16 22:39:50 +00:00
i-robot 26944890e3
!1366 Adjust some codes according to community rules
Merge pull request !1366 from cmwenxin/master
2022-02-15 08:26:34 +00:00
cmwenxin db75e89b96 Make code readable based on openLooKeng community 2022-02-15 15:13:09 +08:00
i-robot 0103406c19
!1378 fix remaining cleancode problems after we use new rules
Merge pull request !1378 from tushengxia/codecheck-planner-other
2022-02-11 06:40:44 +00:00
tushengxia 53a1819776 fix remaining codecheck problems 2022-02-11 13:17:16 +08:00
i-robot 2af9b39351
!1376 Cleancode refactor based on the community rules
Merge pull request !1376 from lizheng(lifengzi)/fix-dailychecks
2022-02-11 01:39:27 +00:00
i-robot d601e4313f
!1367 fix cleancode problems after we use new rules
Merge pull request !1367 from tushengxia/codecheck-planner-other
2022-02-10 16:11:26 +00:00
i-robot 3a3c2eda67
!1377 hetu-core clean code modify
Merge pull request !1377 from chenpingzeng/clean_code_modify
2022-02-10 15:19:26 +00:00
tushengxia 970c63927a fix docs problem and log problem 2022-02-10 20:09:36 +08:00
lizheng920625 77790e5639 Make code more clear in several modules based on openLooKeng community rules 2022-02-10 20:00:52 +08:00
chenyidao1 eaaf065220 fix hetu-core clean code daily check result
Signed-off-by: chenyidao1 <979136761@qq.com>
2022-02-10 19:53:27 +08:00
i-robot b20c9ce2c9
!1368 clean_code_new_1
Merge pull request !1368 from chen/clean_code_new_1
2022-02-10 11:52:54 +00:00
tushengxia 9b2f67c22c fix codecheck problems of presto-main module 2022-02-09 20:08:45 +08:00
i-robot 8152686fa1
!1360 Adjust some codes according to community rules
Merge pull request !1360 from lizheng(lifengzi)/fix-dailychecks
2022-02-09 10:33:55 +00:00
chenyidao1 24113c7444 openlookeng clean_code 2022-02-09 17:33:14 +08:00
i-robot e7d41141a8
!1363 hetu-core clean code modify
Merge pull request !1363 from chenpingzeng/clean_code_modify
2022-02-09 09:13:56 +00:00
lizheng920625 117a7fed93 Make code more clear in several modules based on openLooKeng community rules 2022-02-09 17:00:21 +08:00
i-robot 849e94b856
!1357 Clean code according to the community rule
Merge pull request !1357 from zhousipei/localmaster
2022-02-09 08:26:01 +00:00
Sundar Annamalai edb66cb49a Cube 'varchar' range predicate support. 2022-02-07 10:46:58 -05:00
zhousipei e2de0da06f Clean code according to the community rule 2022-01-29 11:39:39 +08:00
chenpingzeng 0f9f5ddc40 hetu-core clean code modify
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2022-01-27 17:03:16 +08:00
i-robot 7c785c9e6f
!1358 Hindex: fix bloom index size too large
Merge pull request !1358 from peiwangdb/bloom-size-analyze
2022-01-24 20:53:17 +00:00
i-robot 99bc09758d
!1356 Fix I4QU7A and two errors in memory connector
Merge pull request !1356 from peiwangdb/fix-doc
2022-01-24 19:41:16 +00:00
peiwangdb 9d0614378d fix bloom index size too large issue 2022-01-24 10:16:04 -05:00
peiwangdb eb7d166a23 fix memory connector doc description error 2022-01-20 09:13:48 -05:00
peiwangdb 985e61a357 fix-I4QU7A 2022-01-20 08:43:05 -05:00
i-robot 6bcbbaefd8
!1355 Fix for I4M2LW
Merge pull request !1355 from jessica-surya/I4M2LW-potential-fix
2022-01-05 19:22:39 +00:00
i-robot 1198e8c365 !1353 Disable StarTree Cube test because of decimal comparison issue
Merge pull request !1353 from sundarannamalai/cube-0322
2022-01-01 02:09:11 +00:00
Sundar Annamalai e4bab6c5c6 Disable StarTree Cube test due to decimal comparison issue. 2021-12-31 11:04:29 -05:00
i-robot 335831ba98 !1349 fix error in doc release note 1.5.0
Merge pull request !1349 from xudezhi/docErrorFix
2021-12-31 03:19:17 +00:00
xudezhi 2ab1d44af0
release note1.5.0 error fix 2021-12-31 03:01:49 +00:00
i-robot 6fdabd8035 !1346 fix docs problem of star tree
Merge pull request !1346 from tushengxia/master
2021-12-30 09:17:40 +00:00
tushengxia 8e583bf775 fix docs problem of star tree 2021-12-30 16:19:56 +08:00
i-robot a8f72db0db !1344 add index for release notes
Merge pull request !1344 from tushengxia/master
2021-12-30 01:07:29 +00:00
tushengxia bc975fb840 add index for release notes 2021-12-29 22:08:00 +08:00
Raghunandan e1043f02a2 [maven-release-plugin] prepare for next development iteration 2021-12-29 13:20:26 +05:30
Raghunandan 58c2bbed54 [maven-release-plugin] prepare branch branch-1.5 2021-12-29 13:20:26 +05:30
i-robot 663e698814 !1343 update release notes for 1.5.0 and fix docs problem of star tree
Merge pull request !1343 from tushengxia/master
2021-12-29 07:13:38 +00:00
tushengxia 8df2e6fb15 1.fix docs issue of star tree 2.add release note for 1.5.0 2021-12-29 11:51:58 +08:00
Kevin Wan 9a7b0f9210 Potential fix for #I4M2LW 2021-12-24 15:34:27 -05:00
i-robot c5c75043b8 !1342 [264]: Spill to disk log
Merge pull request !1342 from i-robot/pull265
2021-12-24 07:48:25 +00:00
rajeevrastogi d5e3339614 [264]: Spill to disk log 2021-12-24 11:59:26 +05:30
i-robot 625b878006 !1341 upgrade thirdparty software for hetu core
Merge pull request !1341 from chenpingzeng/1230_rc1_release_safety_issue
2021-12-24 02:12:17 +00:00
i-robot 5aee70d1b2 !1340 StarTree Cube documentation updated
Merge pull request !1340 from sundarannamalai/cube-123021
2021-12-24 00:04:16 +00:00
Sundar Annamalai 701d1a27c4 StarTree Cube documentation - Index updated. 2021-12-23 18:47:20 -05:00
i-robot b896411631 !1338 Memory Connector: fix I4NVW3, capture null bloomfilter case
Merge pull request !1338 from peiwangdb/fix-I4NVW3
2021-12-23 22:26:20 +00:00
i-robot bd7d2339eb !1339 StarTree Cube documentation updated
Merge pull request !1339 from sundarannamalai/cube-123021
2021-12-23 21:24:20 +00:00
Sundar Annamalai 81233f6492 StarTree Cube documentation updated for Join query support. 2021-12-23 15:47:05 -05:00
peiwangdb 111aea473b capture null bloomfilter 2021-12-23 14:53:03 -05:00
chenpingzeng df60d4d6de upgrade software for hetu to enhense robust
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2021-12-23 20:57:33 +08:00
i-robot d5dbb7a6b1 !1336 INSERT OVERWRITE not supported on partitioned cubes
Merge pull request !1336 from sundarannamalai/cube-123021
2021-12-22 21:24:18 +00:00
Sundar Annamalai 158b6fc355 Fail INSERT OVERWRITE on partitioned cubes. 2021-12-22 13:38:09 -05:00
i-robot 330d229c95 !1327 Spiller microbenchmark enhancements additional params
Merge pull request !1327 from i-robot/pull253
2021-12-22 08:16:18 +00:00
i-robot 7d36bfda72 !1335 [I4NBP3] [260] Single Column Read or Update Issue Fix
Merge pull request !1335 from i-robot/pull261
2021-12-22 07:42:16 +00:00
Surya Sumanth N c84c9d2fc2 [I4NBP3] [260] Single Column Read or Update Issue Fix 2021-12-22 12:30:10 +05:30
i-robot 79dc5040ad !1332 add the Chinese configurations of the HTTP Client
Merge pull request !1332 from wyy566/master
2021-12-22 03:48:17 +00:00
i-robot b68a4f4fa0 !1333 Memory Connector: add more debug information to the data spill process
Merge pull request !1333 from peiwangdb/debug-spill
2021-12-22 00:20:16 +00:00
peiwangdb 503825fe3e Debug LP processing 2021-12-21 16:55:52 -05:00
i-robot 8fac6f337d !1334 Extract quoted value properly from CREATE CUBE where clause predicate
Merge pull request !1334 from sundarannamalai/cube-123021
2021-12-21 21:52:15 +00:00
Sundar Annamalai 41502fc40d Parse CREATE CUBE where clause properly and extract quoted value. 2021-12-21 16:02:55 -05:00
wyy566 f9910f25c1 add the Chinese configurations of the HTTP Client 2021-12-21 22:05:29 +08:00
i-robot a98d4393c4 !1331 upgrade io.netty to latest version
Merge pull request !1331 from chenpingzeng/1230_rc1_software_upgrade
2021-12-21 09:04:51 +00:00
chenpingzeng 97b0133ec9 upgrade airlift netty to latest 4.1.72.final
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2021-12-21 15:20:39 +08:00
i-robot b0fd71fbcb !1329 fix resource leak and path manipulation problem
Merge pull request !1329 from chenpingzeng/1230_rc1_release_safety_issue
2021-12-21 06:24:53 +00:00
i-robot 534df0117a !1330 Fix incorrect AVG result when StarTree Cube is enabled.
Merge pull request !1330 from sundarannamalai/cube-123021
2021-12-20 23:58:52 +00:00
Sundar Annamalai cc4c4377e4 Cast sum, count values to avg result type to get accurate decimal values. 2021-12-20 17:31:03 -05:00
chenpingzeng 8f391b32d5 fix safety issues for hetu core to make it more robust
Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2021-12-20 21:03:16 +08:00
i-robot db202e451a !1324 Fix I4M0YA, partitioned pages with null as partitionkey can't be processed
Merge pull request !1324 from peiwangdb/fix-I4M0YA-new
2021-12-20 08:31:00 +00:00
i-robot a9e2ad6b39 !1328 [258] Enable httpClient timeout configs for Exchange,Scheduler,MemMgr
Merge pull request !1328 from i-robot/pull259
2021-12-20 07:56:53 +00:00
Nitin Kashyap ee8aa48c27
[258] Enable httpClient timeout configs for Exchange,Scheduler,MemoryManager client 2021-12-20 12:55:26 +05:30
i-robot 0317c531ee !1325 use stringbuilder concatenation in TestMemorySelection
Merge pull request !1325 from i-robot/pull255
2021-12-17 15:02:51 +00:00
mudit khasgiwale 831a81fd6f use stringbuilder concatenation 2021-12-17 17:58:51 +05:30
i-robot e0779440a3 !1304 OpenStream and return value ignored fixed
Merge pull request !1304 from i-robot/pull242
2021-12-17 06:02:49 +00:00
i-robot ce0f9daa8a !1323 Fixes incorrect result returned by partitioned_cube
Merge pull request !1323 from sundarannamalai/partitioned_cube
2021-12-17 00:58:57 +00:00
peiwangdb d0a103e2ed fix-I4M0YA 2021-12-16 16:58:26 -05:00
Sundar Annamalai 3c50840931 Rewrite Cube TableScan output symbol with partition columns at the end of the list.
This reverts commit cb74f083ff.
2021-12-16 16:05:25 -05:00
i-robot 68468ba34b !1320 Fix I4KUOB build failure
Merge pull request !1320 from Kevin Wan/fix-thread
2021-12-16 16:44:50 +00:00
rishabhmurarka7 e89f948193 OpenStream and return value ignored 2021-12-16 21:28:44 +05:30
i-robot 123fa7ace0 !1322 [246] Local of Known Non Null
Merge pull request !1322 from i-robot/pull247
2021-12-16 13:04:49 +00:00
i-robot 490c9110a7 !1317 spot bug fix for BitmapIndex File
Merge pull request !1317 from i-robot/pull209
2021-12-15 09:56:56 +00:00
i-robot 20dee80bf2 !1316 fixed spotbugs in file DefaultConnectorConfigFunctionRewriter
Merge pull request !1316 from i-robot/pull227
2021-12-15 07:26:57 +00:00
Kevin Wan 583878c3c4 Fix issue I4KUOB 2021-12-14 17:13:50 -05:00
i-robot b776254af3 !1309 Optimize Aggregation over Join using StarTree Cube
Merge pull request !1309 from sundarannamalai/cube-123021
2021-12-14 20:20:53 +00:00
Sundar Annamalai f6cd23a099 Optimize aggregation queries over join with StarTree Cube 2021-12-14 14:32:03 -05:00
i-robot aec057c3e2 !1303 [243] Dead Local Store Issue
Merge pull request !1303 from i-robot/pull244
2021-12-14 12:30:28 +00:00
i-robot fa9c0b606d !1315 [I4L3DN] [I4M172] [248] Fix Read Issue for Struct Data Column
Merge pull request !1315 from i-robot/pull249
2021-12-14 11:58:34 +00:00
Nitin Kashyap 95bc57f7cd
Spiller microbenchmark enhancements additional params(direct,compression,prefetch,encrypt) 2021-12-14 11:52:47 +05:30
Surya Sumanth N 24c23be12d [I4L3DN] [I4M172] [248] Fix Read Issue for Struct Data Column 2021-12-14 10:30:36 +05:30
i-robot 623504e178 !1308 StarTree Cube: Cube Range Visitor Comparison Expression Operation Adjustment
Merge pull request !1308 from Daniel Zhang/cube-canonicalizer-expression-fix
2021-12-13 20:44:58 +00:00
i-robot 917dbbc63c !1314 Fix I4KUOB build failure
Merge pull request !1314 from peiwangdb/fix-I4KUOB
2021-12-13 19:18:53 +00:00
i-robot 781d350eee !1313 Memory Connector: fix a bug for handling double type
Merge pull request !1313 from peiwangdb/fix-I4M0YA
2021-12-13 18:48:56 +00:00
peiwangdb 462945f2c1 fix build failure I4KUOB 2021-12-13 11:52:44 -05:00
i-robot 234ce472a2 !1311 CubeConsole - CreateCube fix incorrect catalog selection
Merge pull request !1311 from sundarannamalai/ch-cube-issue
2021-12-13 16:18:52 +00:00
peiwangdb 291d375963 memory connector: fix a bug for restoring tables with double type 2021-12-13 10:23:51 -05:00
i-robot c19bb24c24 !1305 remove dead variables in MemoryTableProperties
Merge pull request !1305 from i-robot/pull240
2021-12-13 15:18:57 +00:00
i-robot c01bc66010 !1301 [237] return value ignored
Merge pull request !1301 from i-robot/pull238
2021-12-13 14:48:58 +00:00
i-robot e890b8e958 !1312 Remove log4j-core in ES connector
Merge pull request !1312 from lizheng(lifengzi)/log4j_es
2021-12-13 14:16:55 +00:00
lizheng920625 dce05ac35b update log4j2 version 2021-12-13 19:26:35 +08:00
i-robot b1e49eb42a !1300 Fix write to static variable from instance method warning
Merge pull request !1300 from i-robot/pull235
2021-12-11 11:06:32 +00:00
i-robot ffe1016db0 !1297 return value ignored no side effect fix
Merge pull request !1297 from i-robot/pull229
2021-12-11 09:18:33 +00:00
Daniel Zhang cb74f083ff StarTree Cube: Cube Range Visitor Comparison Expression Operation Adjustment 2021-12-10 18:35:57 -05:00
Sundar Annamalai 9f13593c0e Use correct catalog to retrieve datatype of the predicate column while creating cube. 2021-12-10 16:53:50 -05:00
i-robot de9bc95cea !1302 StarTree Cube: Cube Creation Processing Type Identifier Case Sensitive Adjustment for Between Operations
Merge pull request !1302 from Daniel Zhang/cube-canonicalizer-expression-fix
2021-12-10 15:53:02 +00:00
i-robot 3976f667d5 !1292 Spotbug fix for TestHbase file
Merge pull request !1292 from i-robot/pull222
2021-12-10 13:11:04 +00:00
i-robot 3c314d3388 !1295 SpotBug fix for FunctionWriterManager
Merge pull request !1295 from i-robot/pull225
2021-12-10 12:41:04 +00:00
i-robot dae8ed7c45 !1293 [177] Spotbugs fix for TestResult file
Merge pull request !1293 from i-robot/pull223
2021-12-10 12:11:00 +00:00
i-robot 1b092e7c8b !1291 [183] handled resource leakage for ResultSet object.
Merge pull request !1291 from i-robot/pull218
2021-12-10 11:41:04 +00:00
i-robot 56bc94a529 !1290 Non null return violation fix
Merge pull request !1290 from i-robot/pull216
2021-12-10 10:51:00 +00:00
i-robot e0a6774f14 !1280 [#177] Fix: spotbug issue TestFunctionMetadata
Merge pull request !1280 from i-robot/pull185
2021-12-10 10:21:01 +00:00
i-robot 2b80321f6e !1289 return value ignored no side effect fix
Merge pull request !1289 from i-robot/pull213
2021-12-10 09:53:00 +00:00
i-robot da749d3827 !1307 [171] Fix for to_unixtime() issue
Merge pull request !1307 from i-robot/pull170
2021-12-10 07:58:04 +00:00
i-robot 0e2980d569 !1288 [177]: Fix for spot bug OrcPageSource
Merge pull request !1288 from i-robot/pull207
2021-12-10 07:26:08 +00:00
Daniel Zhang 0f87b88d50 StarTree Cube: Cube Creation Processing Type Identifier Case Sensitive Adjustment for Between Operations 2021-12-09 10:38:33 -05:00
Vandana B T 976c696e55 Local of Known Non Null 2021-12-09 15:56:32 +05:30
Sharanya Desai 794f9c5a04 [#177] Fix: spotbug issue 2021-12-09 14:56:56 +05:30
C S Likhith 271e584d34 [208]: spot bug fix for BitmapIndex File 2021-12-09 14:54:13 +05:30
Aman Omer d8c1371086 to_unixtime Timezone issue 2021-12-09 13:59:30 +05:30
i-robot a0b663227d !1286 [177]: Fix spot bug issue in OrcSelectivePageSourceFactory
Merge pull request !1286 from i-robot/pull205
2021-12-09 03:54:07 +00:00
i-robot e6c2e4f6a8 !1173 [fix] Some update and delete SQL statements are pushed down incorrectly
Merge pull request !1173 from chenpingzeng/delete_update_pushdown_fix
2021-12-09 02:38:11 +00:00
i-robot ef3f78f497 !1285 [177]: Fix for spot bug in TestArbitraryOutputBuffer
Merge pull request !1285 from i-robot/pull202
2021-12-08 15:50:06 +00:00
Abhishek Gupta 95e181994e Dead Local Store Issue 2021-12-08 20:33:30 +05:30
i-robot 3723f05362 !1287 [177]: fix spot bug issue in TestHiveFileFormats
Merge pull request !1287 from i-robot/pull206
2021-12-08 14:56:07 +00:00
Akanksha Kedia 350cd812bd remove dead local variables 2021-12-08 19:54:49 +05:30
i-robot ee454137fd !1283 spotbug fix for TestBTreeIndex
Merge pull request !1283 from i-robot/pull192
2021-12-08 13:58:06 +00:00
i-robot e90d0c8d1f !1298 added assert to ensure catalog is not null
Merge pull request !1298 from i-robot/pull230
2021-12-08 13:28:06 +00:00
i-robot f55f8d8b3b !1282 [175]: Fix for spot bug
Merge pull request !1282 from i-robot/pull176
2021-12-08 11:40:05 +00:00
Nancy Saini 483f3cca67 return value ignored 2021-12-08 15:29:57 +05:30
KushalSankanna 839ecf1c3b write to static variable from instance method 2021-12-08 15:24:36 +05:30
i-robot 0e6cee04be !1269 Fixed spotbug issues for IndexServiceUtils
Merge pull request !1269 from i-robot/pull182
2021-12-08 07:32:15 +00:00
akanksha 3ff1eefa7e added assert to ensure catalog is not null 2021-12-08 12:57:13 +05:30
rajesh322 6c7b7d1c17 [175]: Fix for spot bug 2021-12-08 12:56:56 +05:30
manjunath 12809c812a fixed spotbugs in file DefaultConnectorConfigFunctionRewriter 2021-12-08 12:53:43 +05:30
apeksha g raj bf9857cfea return value ignored no side effect fix 2021-12-08 12:35:06 +05:30
Deepthi V S 9b72833821 SpotBug fix for FunctionWriterManager 2021-12-08 12:29:18 +05:30
Geetha 06ff037625 [177] Spotbugs fix for TestResult file 2021-12-08 12:18:13 +05:30
Vaishnavi-A27 33bea350a3 [177]: Fix for spot bug in TestArbitraryOutputBuffer 2021-12-08 12:02:32 +05:30
Ganga Bhavani 123dd66a46 Spotbug fix for TestHbase file 2021-12-08 11:27:50 +05:30
M C Shravan 00d3512e95 [177] handled resource leakage for ResultSet object. 2021-12-08 11:01:53 +05:30
Ramyashree S 3dd489164e Non null return violation fix 2021-12-08 10:52:24 +05:30
Neha Harish c38757e7e4 return value ignored no side effect fix 2021-12-08 10:48:06 +05:30
jyotsna bellary eba1e05aa9 [177]: fix spot bug issue in TestHiveFileFormats 2021-12-08 10:14:12 +05:30
srima 028d4bc099 [177]: Fix for spot bug OrcPageSource 2021-12-08 10:11:26 +05:30
POOJA GUTTAL 106b5f05d8 [177]: Fix spot bug issue in OrcSelectivePageSourceFactory 2021-12-08 09:51:50 +05:30
Rakshith Reddy A da6bb031ac spotbug fix for TestBTreeIndex 2021-12-08 09:50:10 +05:30
Manoj S 4ed49deaa4 Fixed spotbug issues for IndexServiceUtils 2021-12-08 09:36:54 +05:30
i-robot 7c37e6e9d6 !1277 removing dead variable
Merge pull request !1277 from i-robot/pull184
2021-12-08 03:30:08 +00:00
i-robot a3ca110485 !1278 dead local store fix
Merge pull request !1278 from i-robot/pull189
2021-12-08 02:50:11 +00:00
i-robot c3e23e9980 !1275 [#172] Fix: gramatical error in comment.
Merge pull request !1275 from i-robot/pull173
2021-12-08 01:08:12 +00:00
i-robot 22efb5dd75 !1274 dead local store fix
Merge pull request !1274 from i-robot/pull201
2021-12-08 00:34:04 +00:00
i-robot fbaee6d06e !1276 use stringbuilder concatenation
Merge pull request !1276 from i-robot/pull199
2021-12-07 18:20:05 +00:00
i-robot 01e86e5ece !1273 static field correction
Merge pull request !1273 from i-robot/pull179
2021-12-07 17:52:04 +00:00
i-robot c76f022e54 !1271 [177]:"fix for spotbug in bloomindex"
Merge pull request !1271 from i-robot/pull187
2021-12-07 17:22:05 +00:00
i-robot c8e2574c8d !1268 remove the unused variable
Merge pull request !1268 from i-robot/pull181
2021-12-07 16:52:04 +00:00
i-robot e7d544b09b !1272 spot bug fixed for test participation index
Merge pull request !1272 from i-robot/pull194
2021-12-07 16:22:06 +00:00
i-robot a3635eaaef !1267 Fixed Spotbug reported issue DLS-DEAD_LOCAL_STORE
Merge pull request !1267 from i-robot/pull174
2021-12-07 15:50:05 +00:00
i-robot 08e7ae2ce9 !1270 spotbug fixes for TestIndexResources.java
Merge pull request !1270 from i-robot/pull186
2021-12-07 12:04:05 +00:00
shobhitha m c827ae04d8 dead local store fix 2021-12-07 16:08:02 +05:30
sai chandana dfdf3e33ff use stringbuilder concatenation 2021-12-07 15:58:46 +05:30
KAVYA B e5a2bdb87b dead local store fix 2021-12-07 15:12:28 +05:30
tejasvinu 71da88b633 [177]:"fix for spotbug in bloomindex" 2021-12-07 15:08:26 +05:30
Prathik S Shah 06b382dd04 spotbug fixes for TestIndexResources.java 2021-12-07 15:03:25 +05:30
venkatadri072 4f1593870a spot bug fixed for test participation index 2021-12-07 14:54:02 +05:30
kavya T 82708e1294 removing dead variable 2021-12-07 14:44:28 +05:30
dras227 4de66a1d73 remove the unused variable 2021-12-07 14:32:42 +05:30
Nachiketh N 8a9333880e static field correction 2021-12-07 14:23:38 +05:30
chenpingzeng 151d11848a [fix] delete and update with filter condition does not pushdown to
datasource

Signed-off-by: chenpingzeng <chenpingzeng@huawei.com>
2021-12-07 15:17:08 +08:00
Ragavendra Kumar S 0e56f27815 Fixed Spotbug reported issue DLS-DEAD_LOCAL_STORE 2021-12-07 12:31:00 +05:30
Vihaashetty fc1c66a347 [#172] Fix: gramatical error in comment. 2021-12-07 09:39:18 +05:30
i-robot 8d5f2a8592 !1265 fix I4KUP5 build failure
Merge pull request !1265 from peiwangdb/fix-I4KUP5
2021-12-06 22:02:07 +00:00
peiwangdb 2a0ae9e0a5 fix-I4KUP5 2021-12-06 16:15:40 -05:00
i-robot cd2b7ccefd !1261 Optimize partitioning strategy
Merge pull request !1261 from peiwangdb/sort-part
2021-12-02 16:38:08 +00:00
i-robot ee0853f716 !1243 Make Memory Connector support statistic features
Merge pull request !1243 from peiwangdb/mem-stats
2021-12-02 14:24:11 +00:00
peiwangdb dee92cf70d memory connector: adding sorting to prevent generating many small partitioned pages 2021-12-01 21:09:05 -05:00
i-robot b27aa53e27 !1259 StarTree Cube: Cube Creation Processing Type Identifier Case Sensitive Adjustment
Merge pull request !1259 from Daniel Zhang/cube-type-adjust
2021-12-01 20:40:06 +00:00
i-robot 1235ff63ef !1254 Remove seeds from on-yarn seedstore file
Merge pull request !1254 from lilianyuan_c78e/master
2021-12-01 19:44:06 +00:00
i-robot 5cde9e95a9 !1258 StarTree Cube: Cube Canonicalizer Comparison Expression SymbolReference Cast Type Fix
Merge pull request !1258 from Daniel Zhang/cube-canonicalizer-expression-fix
2021-12-01 19:00:07 +00:00
peiwangdb ecaa6b4075 Memory Connector: inject stats features 2021-12-01 13:31:29 -05:00
i-robot cca8adb4d7 !1253 Track memory usage in SingleInputSnapshotState
Merge pull request !1253 from Kevin Wan/single-state-tracking
2021-12-01 18:26:06 +00:00
Kevin Wan 0576c49961 Track memory usage in SingleInputSnapshotState 2021-12-01 12:29:29 -05:00
i-robot 0edd69ecc1 !1257 Spill Prefetch Read and Prioritizing Larger splits in Memory Revoke
Merge pull request !1257 from i-robot/pull158
2021-12-01 10:42:06 +00:00
Daniel Zhang db01dcbf88 StarTree Cube: Cube Creation Processing Type Identifier Case Sensitive Adjustment 2021-11-30 16:05:34 -05:00
Daniel Zhang bdd3131bd7 StarTree Cube: Cube Canonicalizer Comparison Expression SymbolReference Cast Type Fix 2021-11-30 14:10:43 -05:00
SURYA SUMANTH N b26b7b1d9f Spill Prefetch and Prioritize Larger Splits 2021-11-30 19:50:59 +05:30
i-robot c96aa28292 !1239 [124] [I3BXGG] Output is not same as hive when data contains '\0'
Merge pull request !1239 from i-robot/pull123
2021-11-30 05:14:05 +00:00
Lilian Yuan c1262c18e8 Remove seed from yarn seedstore if a coordinator goes offline 2021-11-29 08:45:06 -08:00
i-robot aeda1421e7 !1251 support time column and testcase for all data types
Merge pull request !1251 from i-robot/pull133
2021-11-29 10:46:13 +00:00
i-robot 3605efa14a !1256 [I44SI7] [I3PSDU] Fix Alter Column DataType Issue
Merge pull request !1256 from i-robot/pull165
2021-11-29 09:08:23 +00:00
aloknath bea799f5f2 support time column and testcase for all data types 2021-11-29 12:29:09 +05:30
i-robot 9b95e57de1 !1252 [157] Spiller serialization optimizations
Merge pull request !1252 from i-robot/pull166
2021-11-27 12:40:09 +00:00
i-robot 5c43b17d5c !1255 [#167] Query Execution Metric Correction
Merge pull request !1255 from i-robot/pull168
2021-11-27 11:58:09 +00:00
Nitin Kashyap f16d1e49e2
[157] Spiller serialization optimizations 2021-11-27 16:44:43 +05:30
rajeevrastogi 86713234df [#167] Query Execution Metric Correction 2021-11-27 16:02:34 +05:30
Aman Omer b3db79791f [I3BXGG] Output is not same as hive when data contains '\0' 2021-11-26 16:00:51 +05:30
i-robot efa2bc3ef5 !1250 Fix snapshot scheduling bug causing queries to hang
Merge pull request !1250 from Kevin Wan/hanging-fix
2021-11-25 16:12:10 +00:00
Kevin Wan 1afc642499 Fix snapshot hanging queries that are not related to resource usage
Add tests for more complicated Join/Exchange trees
2021-11-24 13:05:03 -05:00
SURYA SUMANTH N e56166b183 Fix Alter Column DataType Issue 2021-11-24 22:46:38 +05:30
i-robot b311335c7e !1249 Add null checks in MemoryColumnHandle
Merge pull request !1249 from farhan3/I4IVOB
2021-11-24 16:32:12 +00:00
i-robot cadaf36c32 !1218 [129] [I4BK7C] [I3WK54] Fix Alter Table Rename Column Bug
Merge pull request !1218 from i-robot/pull130
2021-11-23 07:12:10 +00:00
farhan3 bfa5194373 Add null checks in MemoryColumnHandle 2021-11-22 15:06:10 -05:00
i-robot 85738c416c !1227 'ALTER TABLE' on StarTree Cube should fail and doc changes
Merge pull request !1227 from mahtabahmed/debug_alter_doc
2021-11-22 15:06:06 +00:00
mahtabahmed dbceb4c808 fixed the alter command bugs, added testAlterTableOnCube() and updated the website document 2021-11-19 15:38:52 -05:00
SURYA SUMANTH N fdae299b64 Fix Alter Table Rename Column Bug 2021-11-19 14:30:47 +05:30
i-robot 268250457b !1244 Memory connector: Handle null partitions correctly
Merge pull request !1244 from farhan3/I4HW8H
2021-11-19 02:13:50 +00:00
i-robot ff8bbfc692 !1248 Hindex: make loading index into cache a blocking operation
Merge pull request !1248 from farhan3/I4IPA2
2021-11-19 01:43:44 +00:00
farhan3 d8ec99c8c9 Hindex: make loading index into cache a blocking operation 2021-11-18 18:01:49 -05:00
i-robot 16f5bde8d3 !1247 [160][I4HBFC] hetu-carbondata UT fix
Merge pull request !1247 from i-robot/pull159
2021-11-18 12:25:43 +00:00
Aman Omer 302c8678ed [I4HBFC] hetu-carbondata UT (TestAllCarbonType.test_writer_count) fix 2021-11-18 11:17:54 +05:30
farhan3 46ba508d45 Memory connector: Handle null partitions correctly 2021-11-17 16:18:45 -05:00
i-robot 26ed51b07f !1245 Hindex: print index loaded message at correct time
Merge pull request !1245 from farhan3/I4IPA2
2021-11-17 21:03:22 +00:00
i-robot 2c9508ad45 !1246 Fix I4I83S
Merge pull request !1246 from peiwangdb/fix-I4I83S
2021-11-17 20:15:13 +00:00
peiwangdb ccdbb48add fix issue I4I83S 2021-11-17 14:48:50 -05:00
farhan3 034724224d Hindex: print index loaded message at correct time 2021-11-17 14:22:42 -05:00
i-robot 4266feeb37 !1241 spacing correction in deployment-ha.md, cache-table.md
Merge pull request !1241 from i-robot/pull152
2021-11-17 08:11:40 +00:00
Sailahari_Gorli 929b11b5e6 spacing correction in deployment-ha.md, cache-table.md 2021-11-16 21:00:41 +05:30
i-robot 98e7e914c6 !1187 [I3YM5D] The execution time printed from CLI is not correct
Merge pull request !1187 from i-robot/pull104
2021-11-16 07:35:19 +00:00
i-robot 84e135c90b !1230 Handle missing and corrupted Parquet statistics
Merge pull request !1230 from farhan3/fix-prq-stats
2021-11-16 06:29:19 +00:00
i-robot 24524b3436 !1235 add UT for SHOW VIEWS
Merge pull request !1235 from i-robot/pull150
2021-11-15 14:32:24 +00:00
i-robot c928647355 !1237 [147] document correction in select.md
Merge pull request !1237 from i-robot/pull148
2021-11-15 13:16:05 +00:00
i-robot a168ef96ae !1234 [153] grammar correction in create-view.md, insert-overwrite.md
Merge pull request !1234 from i-robot/pull154
2021-11-15 13:10:05 +00:00
Aman Omer 2f9e2b54f6 I3YM5D The execution time printed from CLI is not correct 2021-11-15 16:22:02 +05:30
Lokesh Pantam 73e8ce0170 grammar correction in create-view.md, insert-overwrite.md 2021-11-14 21:50:44 +05:30
Hari Chandana 3ecd80cb15 add UT for SHOW VIEWS 2021-11-13 23:27:00 +05:30
Sweta Kota 7099386e76 document correction in select.md 2021-11-13 12:51:47 +05:30
farhan3 d24e580147 Handle missing and corrupted Parquet statistics
Changes cherry-picked from:
- https://github.com/trinodb/trino/issues/1798
- https://github.com/trinodb/trino/issues/3517
Commits/PRs:
- d408b6c85d
- f8899cc723
- 7feccf946d
- https://github.com/trinodb/trino/pull/3541/files
2021-11-12 11:13:52 -05:00
i-robot f6976233f7 !1232 update release notes for 1.4.1
Merge pull request !1232 from tushengxia/check-java-version
2021-11-12 03:54:57 +00:00
tushengxia 38ea3c4cfd update release notes for 1.4.1 2021-11-12 11:18:50 +08:00
i-robot 82cd2a33f7 !1229 Spelling corrections in new-index.md and odbc.md
Merge pull request !1229 from i-robot/pull146
2021-11-11 11:42:56 +00:00
i-robot 22d9c329e4 !1228 Spelling & grammar corrections in deployment-auto.md
Merge pull request !1228 from i-robot/pull144
2021-11-11 11:40:55 +00:00
i-robot 5c07b2572f !1219 Add docs for OmniData connector.
Merge pull request !1219 from jiaotongZou/omnidata
2021-11-11 11:38:55 +00:00
jiaotongZou 48485a7d91 Add docs for omnidata connector 2021-11-11 19:31:38 +08:00
Sneha 47737bc467 Spelling corrections in new-index.md and odbc.md 2021-11-11 14:36:09 +05:30
Advith 8c5da51f58 Spelling & grammar corrections in deployment-auto.md 2021-11-11 14:16:32 +05:30
i-robot 18238c2aae !1224 window.md document file correction
Merge pull request !1224 from i-robot/pull140
2021-11-10 17:12:55 +00:00
i-robot 0e74f5a079 !1223 Spelling correction in lambda.md
Merge pull request !1223 from i-robot/pull138
2021-11-10 17:10:54 +00:00
i-robot 4e4aa884a0 !1225 Language correction geospatial.md
Merge pull request !1225 from i-robot/pull142
2021-11-10 17:06:55 +00:00
i-robot 0529643a7b !1222 alignment changes in conversion
Merge pull request !1222 from i-robot/pull136
2021-11-10 17:02:55 +00:00
i-robot 8c9842259e !1221 Language corrections in memory.md
Merge pull request !1221 from i-robot/pull48
2021-11-10 17:00:56 +00:00
VeenaCutie 0496a3fa92 Language correction geospatial.md 2021-11-10 22:20:20 +05:30
Pranav Bharadwaj 91e71b5f92 Correct language in memory.md 2021-11-10 22:10:03 +05:30
MurliSK 86b7a21195 window.md document file correction 2021-11-10 22:01:39 +05:30
RashmiLaxmeshwar 19f72fc6cd Spelling correction in lambda.md 2021-11-10 21:07:02 +05:30
Manoj Kulkarni 721a00fc02 alignment changes in conversion 2021-11-10 19:52:00 +05:30
i-robot 6ad6a01d5c !1217 Fix StarTree Cube test failures
Merge pull request !1217 from sundarannamalai/st_build_issues
2021-11-09 15:08:57 +00:00
Sundar Annamalai 8e640a46e9 Fix StarTree Cube test failures. 2021-11-08 22:18:15 -05:00
i-robot 922c65889f !1210 [118] [I455PL] Pom corrections for hive-contrib jar
Merge pull request !1210 from i-robot/pull131
2021-11-08 12:32:33 +00:00
i-robot 265509b23a !1216 Heuristic Index: UT Index Load Wait Time Adjustment
Merge pull request !1216 from Daniel Zhang/hindex-ut-changes
2021-11-05 13:58:35 +00:00
Daniel Zhang 9bcec81abd Heuristic Index: UT Index Load Wait Time Adjustment 2021-11-04 14:29:30 -04:00
i-robot 69c81b5f28 !1213 Memory Connector: capture the nullpointer case to fix I4GCFD
Merge pull request !1213 from peiwangdb/fix-I4GCFD
2021-11-03 17:19:11 +00:00
peiwangdb 44ec52126d Memory Connector: capture a null case to fix I4GCFD 2021-11-03 11:55:44 -04:00
i-robot 3d0beed3ab !1209 Memory Connector: Current Bytes Memory Usage Non-Negative Fix
Merge pull request !1209 from Daniel Zhang/jmx-memory-info-cur-bytes
2021-11-02 17:28:47 +00:00
Daniel Zhang 4b9fb3e75d Memory Connector: Usage Info Viewable Using JMX 2021-11-02 12:29:47 -04:00
i-robot d615426513 !1214 check java version for arm, unlimit java 11
Merge pull request !1214 from tushengxia/check-java-version
2021-11-02 06:14:38 +00:00
tushengxia 0d91be94f0 check java version for arm, unlimit java 11 2021-11-02 11:00:59 +08:00
i-robot 3aad0a10fd !1212 Heuristic Index: Index creation for multi-column partition tables document change
Merge pull request !1212 from Daniel Zhang/hindex-doc-changes
2021-11-01 21:04:36 +00:00
Daniel Zhang 2362faf364 Heuristic Index: Index creation for multi-column partition tables document change 2021-11-01 17:00:47 -04:00
i-robot 9105c7da73 !1205 Optimize Star Tree Performance
Merge pull request !1205 from Debasatwa/optimization-startree-performance-branch-v26
2021-11-01 14:24:40 +00:00
Debasatwa Dutta 81163123f6 commit for optimizing the star tree index
fix the isseue with grouping match

added average aggregation column

changes to support avg for group by columns match

added the changes for average

removed the avg changes

average table scan optimization

fix UTs

fix UTs

add UTs

added UTs

fix UTs

added UT

added UTs

changed the UTs

edited UTs

edited UTs

added the session startree false in UTs

changes for group by match prior to cube scan node

avg aggregator source symbol changes

segregated the unit tests

checkstyle fixes

changes for Unit Tests, lineitem table replace the nation table

UTs are divided into smaller methods

added documentation comments for test method

changed the UT tests

changes for group match

updated hetu-docs

edits the create cubes of the UTs

fixes for avg aggregation

fix uts
2021-10-29 21:45:15 -04:00
i-robot 4f7e37834f !1191 Memory connector: add partition functionality
Merge pull request !1191 from peiwangdb/lp-add-part
2021-10-29 22:11:21 +00:00
i-robot 5ed90383f9 !1211 Heuristic Index: BTree Index Properties Access Exception Fix
Merge pull request !1211 from Daniel Zhang/btree-access-fix
2021-10-29 20:39:19 +00:00
Daniel Zhang 82179f00cc Heuristic Index: BTree Index Properties Access Exception Fix 2021-10-29 15:38:29 -04:00
peiwangdb a29d6001cf memory-connector: add partition layer 2021-10-29 08:38:26 -07:00
SURYA SUMANTH N 69073adaeb Pom corrections for hive-contrib jar 2021-10-29 18:12:02 +05:30
i-robot c4d2858d33 !1204 Memory Connector: Memory and Disk Byte Usage Info Viewable Using JMX
Merge pull request !1204 from Daniel Zhang/jmx-memory-info
2021-10-26 18:36:56 +00:00
Daniel Zhang 764df6d9e5 Memory Connector: Memory and Disk Byte Usage Info Viewable Using JMX 2021-10-26 12:48:46 -04:00
i-robot 7b7a7f163e !1190 [102] [I29SBC]Support select ArrayType in Elastic Search connector
Merge pull request !1190 from i-robot/pull103
2021-10-26 09:34:55 +00:00
i-robot f0edc297f3 !1196 Support Hive tables with customized delimiters
Merge pull request !1196 from i-robot/pull119
2021-10-25 12:40:56 +00:00
i-robot a28a95b80d !1206 [125][I4DK84] fixed incorrect codegen for BETWEEN clause while splitting expression
Merge pull request !1206 from i-robot/pull126
2021-10-25 09:24:55 +00:00
Nitin Kashyap 53efdd0651
[I4DK84] fixed incorrect codegen for BETWEEN clause while splitting expression 2021-10-25 11:28:05 +05:30
i-robot dfed60a295 !1203 fix issue-I4EXK9 a few doc typos
Merge pull request !1203 from peiwangdb/fix-I4EXK9
2021-10-23 04:57:14 +00:00
i-robot 91d6566c92 !1179 Add large number format function
Merge pull request !1179 from i-robot/pull101
2021-10-23 04:55:01 +00:00
NALAM JYOTHIRMAYE 2aa8322a9a Support Hive tables with customized delimiters 2021-10-22 21:37:04 +05:30
i-robot a525dc3878 !1195 Added Support for Show Views
Merge pull request !1195 from i-robot/pull113
2021-10-22 08:43:00 +00:00
peiwangdb 43d817d07d hindex: fix typos in zh-doc 2021-10-21 18:48:23 -07:00
i-robot 98eb9f55f8 !1201 fix I4EHMI
Merge pull request !1201 from peiwangdb/fix-I4EHMI-new
2021-10-21 15:45:03 +00:00
i-robot 9ec096beaa !1198 The partial fix for case senstivity issue and average aggregation function issue of Star Tree
Merge pull request !1198 from Debasatwa/startree-fix-branch-v24
2021-10-21 15:41:00 +00:00
Yalamanchili Tirumala Sai Prasad 1427c8e94d support select ArrayType in elastic search connector 2021-10-21 09:14:28 +05:30
Debasatwa Dutta e987270b8a Added the fix for case senstivitiy and aggregation function issue
checkstyle fix

fix the average aggregation function issue

added unit tests

unit test fixex

fix for symbol assignments

added ut tests, removed the conditions for checking duplicate symbol in AggregationRewriteWithCube

Added UT

code clean up

changed the UT with table scan aserts

edited the UTs

fixed cube name

removed duplicates

fix UTs

edited UTs

edited UTs

edit UTs
2021-10-20 12:09:29 -04:00
Nikhil Sinha 7ed56e41cb Added Support for Show Views 2021-10-20 12:44:41 +05:30
i-robot 70155a5d7b !1200 Spelling correction in externalfunction-registration-pushdown.md
Merge pull request !1200 from i-robot/pull122
2021-10-20 04:28:48 +00:00
peiwangdb c714a79558 hindex: align zh and en doc, fix I4EHMI 2021-10-19 20:22:22 -07:00
i-robot 4091438651 !1193 Hindex: disable autoload check for update case
Merge pull request !1193 from peiwangdb/fix-I4E25Q
2021-10-19 17:48:50 +00:00
Sanjeet Nishad 2805f710de Spelling correction in externalfunction-registration-pushdown.md 2021-10-19 20:29:46 +04:00
i-robot dbb922751d !1197 [114]spell fix in types.md
Merge pull request !1197 from i-robot/pull115
2021-10-19 16:16:51 +00:00
Teja Vegi 69c3ac0805 Add large number format function 2021-10-19 20:17:50 +05:30
Tirumala Sai Prasad 3eb763557a spell fix in types.md 2021-10-19 14:05:01 +05:30
i-robot f7ebb98b76 !1166 [79] [I44VUV] Fix Dynamic Filters Issue when ORC Predicate Pushdown is enabled
Merge pull request !1166 from i-robot/pull80
2021-10-19 07:12:49 +00:00
i-robot 029974a27a !1175 [96][I4CVQX] join distribution improvement based on size
Merge pull request !1175 from i-robot/pull97
2021-10-19 05:28:50 +00:00
peiwangdb e1228462f2 hindex: remove autoload check for update case 2021-10-18 18:38:31 -07:00
i-robot ec1b838b69 !1188 Fix a decimal type conversion bug
Merge pull request !1188 from peiwangdb/fix-I4E27U
2021-10-18 21:00:49 +00:00
i-robot 17029ea49a !1192 Spelling Correction in doc elasticsearch.md
Merge pull request !1192 from i-robot/pull110
2021-10-18 15:26:50 +00:00
rakheetest ada54ac887 Spelling Correction in doc elasticsearch.md 2021-10-18 20:00:42 +05:30
peiwangdb dc62673337 fix issue I4E27U, a type conversion error 2021-10-18 06:08:18 -07:00
i-robot cf0e9c4f4c !1189 Add random function taking range min, max
Merge pull request !1189 from i-robot/pull107
2021-10-18 11:00:49 +00:00
soham 4234ceddcf Add random function taking range min, max 2021-10-18 14:58:16 +05:30
i-robot 00edc82a7b !1178 Add reverse function for varbinary
Merge pull request !1178 from i-robot/pull99
2021-10-18 08:46:50 +00:00
i-robot c3d2c1883a !1186 【轻量级 PR】:remove the redundant symbol in the url of StarTree
Merge pull request !1186 from Jiwen/N/A
2021-10-18 07:00:48 +00:00
Jiwen 90adb65bc2 remove the redundant symbol in the url of StarTree 2021-10-16 06:08:51 +00:00
i-robot 856893354e !1185 Fixed index.md issue
Merge pull request !1185 from yumei/index
2021-10-15 09:40:50 +00:00
senny456 754410f06a updata index.md 2021-10-15 17:31:04 +08:00
i-robot a04dfb9544 !1184 add index info of 1.4.0 in doc
Merge pull request !1184 from xudezhi/master
2021-10-15 02:28:48 +00:00
xudezhi 141a772845 add index info of 1.4.0 in doc 2021-10-15 02:21:23 +00:00
i-robot 14b3c25a9f !1182 update faq doc
Merge pull request !1182 from xudezhi/master
2021-10-15 01:44:47 +00:00
i-robot a215311e66 !1183 Hetu-Docs Changes: Preagg Chinese Translation Document Additions
Merge pull request !1183 from Daniel Zhang/preagg_doc_changes
2021-10-14 18:44:48 +00:00
Daniel Zhang 2b533be3ba Hetu-Docs Changes: Preagg Chinese Translation Document Additions 2021-10-14 13:30:55 -04:00
SURYA SUMANTH N 48878999ec Fix Creation of Dynamic Filters Issue when ORC Predicate Pushdown is enabled 2021-10-14 16:23:40 +05:30
xudezhi 5c0188882f update faq doc 2021-10-14 09:24:12 +00:00
Jiwen 37ec5ffdd6 !1181 fix the lack of "relref" which cause the sync failure of doc.
Merge pull request !1181 from Jiwen/N/A
2021-10-14 08:27:41 +00:00
Jiwen 147868e416 fix the lack of "relref" which cause the sync failure of doc. 2021-10-14 08:27:12 +00:00
Jiwen 45ac934a1a !1180 remove the redundant dash symbo in the index.md
Merge pull request !1180 from Jiwen/N/A
2021-10-14 07:41:07 +00:00
Jiwen a2682903e9 remove the redundant dash symbo in the index.md 2021-10-14 07:40:01 +00:00
i-robot b92c581993 !1177 update the Chinese doc of resource group
Merge pull request !1177 from xudezhi/master
2021-10-14 06:06:50 +00:00
Sade Asish Bhushan 4c0a0ab995 Add reverse function for varbinary 2021-10-13 20:27:47 +05:30
Nitin Kashyap 4931f1e4e3
[96][I4CVQX] join distribution improvement based on size 2021-10-13 15:24:02 +05:30
xudezhi 3493ee3b35 update the Chinese doc of resource group 2021-10-13 07:55:05 +00:00
i-robot 88d427fcb4 !1170 add kylin connector in dynamic catalog
Merge pull request !1170 from xudezhi/master
2021-10-13 07:36:54 +00:00
i-robot e16c83d946 !1172 grammar correction hetu-docs/en/develop/functions.md
Merge pull request !1172 from i-robot/pull88
2021-10-12 11:12:50 +00:00
i-robot aae14179e4 !1171 spelling correction in connectors.md
Merge pull request !1171 from i-robot/pull86
2021-10-12 07:59:00 +00:00
i-robot 3ac90935cb !1169 [#91] Integer overflow when parsing dates beyond representable range
Merge pull request !1169 from i-robot/pull92
2021-10-12 07:20:50 +00:00
i-robot ad4aad3e18 !1165 Modified Spelling in Kafka-Tutorial
Merge pull request !1165 from i-robot/pull78
2021-10-12 04:14:48 +00:00
xudezhi 6aaf1690b2 add kylin connector in dynamic catalog 2021-10-12 11:01:39 +08:00
i-robot 5df1e98a16 !1162 spelling correction in hetu-docs/en/admin/spill.md
Merge pull request !1162 from i-robot/pull72
2021-10-11 14:31:50 +00:00
vishalvikal 6c67bee150 Integer overflow when parsing dates beyond representable 2021-10-11 17:00:32 +05:30
madhumita91 06b9a3f77f grammar correction hetu-docs/en/develop/functions.md 2021-10-08 11:23:25 +05:30
07ARB af46cb0071 spelling correction in connectors.md 2021-10-07 07:36:57 +04:00
i-robot 88abb330b0 !1164 Fix incorrect cube selection when query matches mutliple cubes
Merge pull request !1164 from sundarannamalai/cube-123021
2021-10-06 18:41:51 +00:00
Sundar Annamalai 5e080a50e2 Fix incorrect cube selection when multiple cube matches 2021-10-04 17:10:39 -04:00
Ojaswini12 8f57df81d3 Kafka-Tutorial spell fix 2021-10-04 20:37:09 +05:30
Ojaswini12 19adf17e60 Modified Spelling in Kafka-Tutorial 2021-10-04 19:50:10 +05:30
Babu A ea75283000 spelling correction in hetu-docs/en/admin/spill.md 2021-10-04 16:06:18 +05:30
i-robot 3ef3d037d8 !1161 spelling correction in kafka-tutorial.md
Merge pull request !1161 from i-robot/pull70
2021-10-04 04:39:51 +00:00
i-robot 479ad1bc1c !1159 Presto Tests Test Queue Db Query Runner Setup Fix
Merge pull request !1159 from Daniel Zhang/test-queue-db-setup-fix
2021-10-01 21:01:49 +00:00
Daniel Zhang 07bbd37bb8 Presto Tests Test Queue Db Query Runner Setup Fix 2021-10-01 15:43:09 -04:00
i-robot e6eb49a5cb !1157 Heuristic Index Auto Loading Cache Initialization Fix
Merge pull request !1157 from Daniel Zhang/auto-load-fix
2021-10-01 17:41:49 +00:00
i-robot 2c49e982a7 !1146 Hindex: fix an indexcachekey insertion bug
Merge pull request !1146 from peiwangdb/fix-partition-name-bug
2021-10-01 17:09:49 +00:00
Rajesh c1c20df222 spelling correction in kafka-tutorial.md 2021-10-01 13:22:32 +05:30
Raghunandan 3793aa170a [maven-release-plugin] prepare for next development iteration 2021-10-01 13:13:30 +05:30
Daniel Zhang 9cd2531dea Heuristic Index Auto Loading Cache Initialization Fix 2021-09-30 15:28:25 -04:00
peiwangdb cf6288e3ea hindex: fix cachekey inserting bug for partitioned table 2021-09-29 09:41:35 -07:00
1633 changed files with 187458 additions and 16721 deletions

View File

@ -22,7 +22,7 @@
<parent>
<groupId>io.hetu.core</groupId>
<artifactId>presto-root</artifactId>
<version>1.4.0</version>
<version>1.7.0-SNAPSHOT</version>
</parent>
<artifactId>hetu-carbondata</artifactId>

View File

@ -108,8 +108,7 @@ public class CarbondataAutoVacuumThread
AutoVacuumScanTask(SemiTransactionalHiveMetastore metastore)
{
this.metastore = metastore;
this.schemaName = null;
this(metastore, null);
}
AutoVacuumScanTask(SemiTransactionalHiveMetastore metastore, String schemaName)
@ -233,7 +232,6 @@ public class CarbondataAutoVacuumThread
private void submitTaskScanning(CarbondataAutoVacuumThread instanceAutoVacuum, SemiTransactionalHiveMetastore metastore)
{
//trigger task to do scanning of tables
//instanceAutoVacuum.executorService.submit(new AutoVacuumScanTask(metastore));
if (enableTracingCleanupTask) {
queuedTasks.add(instanceAutoVacuum.executorService.submit(new AutoVacuumScanTask(metastore)));
}

View File

@ -73,16 +73,17 @@ public class CarbondataColumnVectorWrapper
@Override
public void putShorts(int rowId, int count, short value)
{
int inputRowId = rowId;
if (filteredRowsExist) {
for (int i = 0; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putShort(counter++, value);
}
rowId++;
inputRowId++;
}
}
else {
columnVector.putShorts(rowId, count, value);
columnVector.putShorts(inputRowId, count, value);
}
}
@ -97,16 +98,17 @@ public class CarbondataColumnVectorWrapper
@Override
public void putInts(int rowId, int count, int value)
{
int inputRowId = rowId;
if (filteredRowsExist) {
for (int i = 0; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putInt(counter++, value);
}
rowId++;
inputRowId++;
}
}
else {
columnVector.putInts(rowId, count, value);
columnVector.putInts(inputRowId, count, value);
}
}
@ -121,16 +123,17 @@ public class CarbondataColumnVectorWrapper
@Override
public void putLongs(int rowId, int count, long value)
{
int inputRowId = rowId;
if (filteredRowsExist) {
for (int i = 0; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putLong(counter++, value);
}
rowId++;
inputRowId++;
}
}
else {
columnVector.putLongs(rowId, count, value);
columnVector.putLongs(inputRowId, count, value);
}
}
@ -145,11 +148,12 @@ public class CarbondataColumnVectorWrapper
@Override
public void putDecimals(int rowId, int count, BigDecimal value, int precision)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putDecimal(counter++, value, precision);
}
rowId++;
inputRowId++;
}
}
@ -164,16 +168,17 @@ public class CarbondataColumnVectorWrapper
@Override
public void putDoubles(int rowId, int count, double value)
{
int inputRowId = rowId;
if (filteredRowsExist) {
for (int i = 0; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putDouble(counter++, value);
}
rowId++;
inputRowId++;
}
}
else {
columnVector.putDoubles(rowId, count, value);
columnVector.putDoubles(inputRowId, count, value);
}
}
@ -196,11 +201,12 @@ public class CarbondataColumnVectorWrapper
@Override
public void putByteArray(int rowId, int count, byte[] value)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putByteArray(counter++, value);
}
rowId++;
inputRowId++;
}
}
@ -223,16 +229,17 @@ public class CarbondataColumnVectorWrapper
@Override
public void putNulls(int rowId, int count)
{
int inputRowId = rowId;
if (filteredRowsExist) {
for (int i = 0; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putNull(counter++);
}
rowId++;
inputRowId++;
}
}
else {
columnVector.putNulls(rowId, count);
columnVector.putNulls(inputRowId, count);
}
}
@ -319,66 +326,72 @@ public class CarbondataColumnVectorWrapper
@Override
public void putFloats(int rowId, int count, float[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = srcIndex; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putFloat(counter++, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putShorts(int rowId, int count, short[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = srcIndex; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putShort(counter++, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putInts(int rowId, int count, int[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = srcIndex; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putInt(counter++, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putLongs(int rowId, int count, long[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = srcIndex; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putLong(counter++, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putDoubles(int rowId, int count, double[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = srcIndex; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putDouble(counter++, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putBytes(int rowId, int count, byte[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = srcIndex; i < count; i++) {
if (!filteredRows[rowId]) {
if (!filteredRows[inputRowId]) {
columnVector.putByte(counter++, src[i]);
}
rowId++;
inputRowId++;
}
}

View File

@ -127,15 +127,16 @@ public class CarbondataFileWriter
private boolean isInitDone;
private boolean isCommitDone;
public CarbondataFileWriter(Path outPutPath, List<String> inputColumnNames, Properties properties,
public CarbondataFileWriter(Path paramOutPutPath, List<String> inputColumnNames, Properties properties,
JobConf configuration, TypeManager typeManager, Optional<AcidOutputFormat.Options> acidOptions,
Optional<HiveACIDWriteType> acidWriteType, OptionalInt taskId) throws SerDeException
{
this.outPutPath = requireNonNull(outPutPath, "path is null");
Path localOutPutPath = paramOutPutPath;
this.outPutPath = requireNonNull(localOutPutPath, "path is null");
// in table creation this can be null
if (null != properties.getProperty("location")) {
this.outPutPath = new Path(properties.getProperty("location"));
outPutPath = new Path(properties.getProperty("location"));
localOutPutPath = new Path(properties.getProperty("location"));
}
this.configuration = requireNonNull(configuration, "conf is null");
this.properties = requireNonNull(properties, "Properties is null");
@ -211,7 +212,7 @@ public class CarbondataFileWriter
Object writer =
Class.forName(MapredCarbonOutputFormat.class.getName()).getConstructor().newInstance();
recordWriter = ((MapredCarbonOutputFormat<?>) writer)
.getHiveRecordWriter(this.configuration, outPutPath, Text.class, compress,
.getHiveRecordWriter(this.configuration, localOutPutPath, Text.class, compress,
properties, Reporter.NULL);
}
@ -226,25 +227,25 @@ public class CarbondataFileWriter
private FileSinkOperator.RecordWriter getHiveWriter(String segmentId, long taskNo) throws Exception
{
Path outPutPath = this.outPutPath;
Properties properties = this.properties;
JobConf configuration = this.configuration;
boolean compress = HiveConf.getBoolVar(configuration, COMPRESSRESULT);
Path finalOutPutPath = this.outPutPath;
Properties finalProperties = this.properties;
JobConf finalConfiguration = this.configuration;
boolean compress = HiveConf.getBoolVar(finalConfiguration, COMPRESSRESULT);
CarbonLoadModel carbonLoadModel = HiveCarbonUtil.getCarbonLoadModel(properties, configuration);
CarbonLoadModel carbonLoadModel = HiveCarbonUtil.getCarbonLoadModel(finalProperties, finalConfiguration);
carbonLoadModel.setSegmentId(segmentId);
carbonLoadModel.setTaskNo(String.valueOf(taskNo));
carbonLoadModel.setFactTimeStamp(Long.parseLong(txnTimeStamp));
carbonLoadModel.setBadRecordsAction(TableOptionConstant.BAD_RECORDS_ACTION.getName() + ",force");
CarbonTableOutputFormat.setLoadModel(configuration, carbonLoadModel);
CarbonTableOutputFormat.setLoadModel(finalConfiguration, carbonLoadModel);
this.configuration.set(CarbondataConstants.TaskId, getTaskAttemptId(String.valueOf(taskNo)));
Object writer =
Class.forName(MapredCarbonOutputFormat.class.getName()).getConstructor().newInstance();
return ((MapredCarbonOutputFormat<?>) writer)
.getHiveRecordWriter(configuration, outPutPath, Text.class, compress,
properties, Reporter.NULL);
.getHiveRecordWriter(finalConfiguration, finalOutPutPath, Text.class, compress,
finalProperties, Reporter.NULL);
}
@Override
@ -285,7 +286,7 @@ public class CarbondataFileWriter
public void appendRow(Page dataPage, int position)
{
FileSinkOperator.RecordWriter recordWriter = null;
FileSinkOperator.RecordWriter finalRecordWriter = null;
if (HiveACIDWriteType.isUpdateOrDelete(acidWriteType)) {
try {
DeleteDeltaBlockDetails deleteDeltaBlockDetails = null;
@ -334,7 +335,7 @@ public class CarbondataFileWriter
return;
}
recordWriter = segmentRecordWriterMap.computeIfAbsent(segmentId, v ->
finalRecordWriter = segmentRecordWriterMap.computeIfAbsent(segmentId, v ->
{
try {
return getHiveWriter(segmentId, CarbonUpdateUtil.getLatestTaskIdForSegment(new Segment(segmentId), tablePath) + 1);
@ -351,7 +352,7 @@ public class CarbondataFileWriter
}
}
else {
recordWriter = this.recordWriter;
finalRecordWriter = this.recordWriter;
}
for (int field = 0; field < fieldCount; field++) {
@ -365,8 +366,8 @@ public class CarbondataFileWriter
}
try {
if (recordWriter != null) {
recordWriter.write(serDe.serialize(row, tableInspector));
if (finalRecordWriter != null) {
finalRecordWriter.write(serDe.serialize(row, tableInspector));
}
}
catch (SerDeException | IOException e) {

View File

@ -48,6 +48,7 @@ public class CarbondataHandleResolver
return CarbonDeleteAsInsertTableHandle.class;
}
@Override
public Class<? extends ConnectorOutputTableHandle> getOutputTableHandleClass()
{
return CarbondataOutputTableHandle.class;

View File

@ -317,11 +317,11 @@ public class CarbondataHetuFilterUtil
if (rawData instanceof Slice) {
String value = ((Slice) rawData).toStringUtf8();
if (type.getTypeInfo() instanceof CharTypeInfo) {
String padding = "";
StringBuilder padding = new StringBuilder();
int paddedLength = ((CharTypeInfo) type.getTypeInfo()).getLength();
int truncatedLength = value.length();
for (int i = 0; i < paddedLength - truncatedLength; i++) {
padding += " ";
padding.append(" ");
}
return value + padding;
}

View File

@ -103,7 +103,6 @@ public class CarbondataHetuOutputFormat<T>
OutputCommitter carbonOutputCommitter = super.getOutputCommitter(context);
JobContextImpl jobContext = new JobContextImpl(jc, new JobID());
carbonOutputCommitter.setupJob(jobContext);
CarbonLoadModel updatedCarbonLoadModel = CarbonTableOutputFormat.getLoadModel(jc);
org.apache.hadoop.mapreduce.RecordWriter re = super.getRecordWriter(context);
return new FileSinkOperator.RecordWriter()
{

View File

@ -80,8 +80,6 @@ public class CarbondataLocationService
{
// TODO: check and make it compatible for cloud scenario
HdfsEnvironment.HdfsContext context =
new HdfsEnvironment.HdfsContext(session, table.getDatabaseName(), table.getTableName());
Path targetPath = new Path(table.getStorage().getLocation());
return new LocationHandle(targetPath, targetPath, true,

View File

@ -306,13 +306,13 @@ public class CarbondataMetadata
private void setupCommitWriter(Properties hiveSchema, Path outputPath, Configuration initialConfiguration, boolean isOverwrite) throws PrestoException
{
CarbonLoadModel carbonLoadModel;
CarbonLoadModel finalCarbonLoadModel;
TaskAttemptID taskAttemptID = TaskAttemptID.forName(initialConfiguration.get("mapred.task.id"));
try {
ThreadLocalSessionInfo.setConfigurationToCurrentThread(initialConfiguration);
carbonLoadModel = HiveCarbonUtil.getCarbonLoadModel(hiveSchema, initialConfiguration);
carbonLoadModel.setBadRecordsAction(TableOptionConstant.BAD_RECORDS_ACTION.getName() + ",force");
CarbonTableOutputFormat.setLoadModel(initialConfiguration, carbonLoadModel);
finalCarbonLoadModel = HiveCarbonUtil.getCarbonLoadModel(hiveSchema, initialConfiguration);
finalCarbonLoadModel.setBadRecordsAction(TableOptionConstant.BAD_RECORDS_ACTION.getName() + ",force");
CarbonTableOutputFormat.setLoadModel(initialConfiguration, finalCarbonLoadModel);
}
catch (IOException ex) {
LOG.error("Error while creating carbon load model", ex);
@ -360,13 +360,13 @@ public class CarbondataMetadata
this.user = session.getUser();
return hdfsEnvironment.doAs(user, () -> {
SchemaTableName tableName = parent.getSchemaTableName();
Optional<Table> table =
Optional<Table> finalTable =
metastore.getTable(new HiveIdentity(session), tableName.getSchemaName(), tableName.getTableName());
if (table.isPresent() && table.get().getPartitionColumns().size() > 0) {
if (finalTable.isPresent() && finalTable.get().getPartitionColumns().size() > 0) {
throw new PrestoException(NOT_SUPPORTED, "Operations on Partitioned CarbonTables is not supported");
}
this.table = table;
this.table = finalTable;
Path outputPath =
new Path(parent.getLocationHandle().getJsonSerializableTargetPath());
initialConfiguration = ConfigurationUtils.toJobConf(this.hdfsEnvironment
@ -382,7 +382,7 @@ public class CarbondataMetadata
}
/* Create committer object */
setupCommitWriter(table, outputPath, initialConfiguration, isOverwrite);
setupCommitWriter(finalTable, outputPath, initialConfiguration, isOverwrite);
return new CarbondataInsertTableHandle(parent.getSchemaName(),
parent.getTableName(),
@ -416,13 +416,13 @@ public class CarbondataMetadata
currentState = State.UPDATE;
HiveInsertTableHandle parent = super.beginInsert(session, tableHandle);
SchemaTableName tableName = parent.getSchemaTableName();
Optional<Table> table =
Optional<Table> finalTable =
this.metastore.getTable(new HiveIdentity(session), tableName.getSchemaName(), tableName.getTableName());
if (table.isPresent() && table.get().getPartitionColumns().size() > 0) {
if (finalTable.isPresent() && finalTable.get().getPartitionColumns().size() > 0) {
throw new PrestoException(NOT_SUPPORTED, "Operations on Partitioned CarbonTables is not supported");
}
this.table = table;
this.table = finalTable;
this.user = session.getUser();
hdfsEnvironment.doAs(user, () -> {
initialConfiguration = ConfigurationUtils.toJobConf(this.hdfsEnvironment
@ -430,8 +430,8 @@ public class CarbondataMetadata
new HdfsEnvironment.HdfsContext(session, parent.getSchemaName(),
parent.getTableName()),
new Path(parent.getLocationHandle().getJsonSerializableWritePath())));
Properties schema = MetastoreUtil.getHiveSchema(table.get());
schema.setProperty("tablePath", table.get().getStorage().getLocation());
Properties schema = MetastoreUtil.getHiveSchema(finalTable.get());
schema.setProperty("tablePath", finalTable.get().getStorage().getLocation());
carbonTable = getCarbonTable(parent.getSchemaName(),
parent.getTableName(),
schema,
@ -470,13 +470,13 @@ public class CarbondataMetadata
HiveInsertTableHandle parent = super.beginInsert(session, tableHandle);
List<HiveColumnHandle> inputColumns = parent.getInputColumns().stream().filter(HiveColumnHandle::isRequired).collect(toList());
SchemaTableName tableName = parent.getSchemaTableName();
Optional<Table> table =
Optional<Table> finalTable =
this.metastore.getTable(new HiveIdentity(session), tableName.getSchemaName(), tableName.getTableName());
if (table.isPresent() && table.get().getPartitionColumns().size() > 0) {
if (finalTable.isPresent() && finalTable.get().getPartitionColumns().size() > 0) {
throw new PrestoException(NOT_SUPPORTED, "Operations on Partitioned CarbonTables is not supported");
}
this.table = table;
this.table = finalTable;
this.user = session.getUser();
hdfsEnvironment.doAs(user, () -> {
initialConfiguration = ConfigurationUtils.toJobConf(this.hdfsEnvironment
@ -484,8 +484,8 @@ public class CarbondataMetadata
new HdfsEnvironment.HdfsContext(session, parent.getSchemaName(),
parent.getTableName()),
new Path(parent.getLocationHandle().getJsonSerializableWritePath())));
Properties schema = MetastoreUtil.getHiveSchema(table.get());
schema.setProperty("tablePath", table.get().getStorage().getLocation());
Properties schema = MetastoreUtil.getHiveSchema(finalTable.get());
schema.setProperty("tablePath", finalTable.get().getStorage().getLocation());
carbonTable = getCarbonTable(parent.getSchemaName(),
parent.getTableName(),
schema,
@ -643,7 +643,7 @@ public class CarbondataMetadata
return hdfsEnvironment.doAs(session.getUser(), () -> {
Properties hiveSchema = MetastoreUtil.getHiveSchema(this.table.get());
CarbonTable carbonTable = getCarbonTable(carbondataVacuumTableHandle.getSchemaName(),
CarbonTable finalCarbonTable = getCarbonTable(carbondataVacuumTableHandle.getSchemaName(),
carbondataVacuumTableHandle.getTableName(),
hiveSchema,
initialConfiguration);
@ -705,7 +705,7 @@ public class CarbondataMetadata
SegmentFileStore.mergeSegmentFiles(readPath, segmentFileName, CarbonTablePath.getSegmentFilesLocation(carbonLoadModel.getTablePath()));
String source;
for (String currPartitionName : partitionNames) {
source = carbonTable.getTablePath() + "/" + currPartitionName;
source = finalCarbonTable.getTablePath() + "/" + currPartitionName;
moveFromTempFolder(source + "/" + carbonLoadModel.getSegmentId() + "_" + timeStamp + ".tmp", source);
}
segmentFilesToBeUpdatedLatest.add(new Segment(carbonLoadModel.getSegmentId(), segmentFileName));
@ -719,7 +719,7 @@ public class CarbondataMetadata
for (CarbondataSegmentInfoUtil segmentInfo : newMergedSegmentInfoUtilList) {
String mergedLoadNumber = segmentInfo.getDestinationSegment();
try {
String segmentFileName = SegmentFileStore.writeSegmentFile(carbonTable, mergedLoadNumber, String.valueOf(carbonLoadModel.getFactTimeStamp()));
String segmentFileName = SegmentFileStore.writeSegmentFile(finalCarbonTable, mergedLoadNumber, String.valueOf(carbonLoadModel.getFactTimeStamp()));
}
catch (IOException e) {
throw new PrestoException(GENERIC_INTERNAL_ERROR, "Failed while merging segment files", e);
@ -900,9 +900,9 @@ public class CarbondataMetadata
private LocationHandle getCarbonDataTableCreationPath(ConnectorSession session, ConnectorTableMetadata tableMetadata, HiveWriteUtils.OpertionType opertionType) throws PrestoException
{
Path targetPath = null;
SchemaTableName schemaTableName = tableMetadata.getTable();
String schemaName = schemaTableName.getSchemaName();
String tableName = schemaTableName.getTableName();
SchemaTableName finalSchemaTableName = tableMetadata.getTable();
String finalSchemaName = finalSchemaTableName.getSchemaName();
String tableName = finalSchemaTableName.getTableName();
Optional<String> location = getCarbondataLocation(tableMetadata.getProperties());
LocationHandle locationHandle;
FileSystem fileSystem;
@ -914,32 +914,32 @@ public class CarbondataMetadata
throw new PrestoException(NOT_SUPPORTED, format("Setting %s property is not allowed", LOCATION_PROPERTY));
}
/* if path not having prefix with filesystem type, than we will take fileSystem type from core-site.xml using below methods */
fileSystem = hdfsEnvironment.getFileSystem(new HdfsEnvironment.HdfsContext(session, schemaName), new Path(location.get()));
fileSystem = hdfsEnvironment.getFileSystem(new HdfsEnvironment.HdfsContext(session, finalSchemaName), new Path(location.get()));
targetLocation = fileSystem.getFileStatus(new Path(location.get())).getPath().toString();
targetPath = getPath(new HdfsEnvironment.HdfsContext(session, schemaName, tableName), targetLocation, false);
targetPath = getPath(new HdfsEnvironment.HdfsContext(session, finalSchemaName, tableName), targetLocation, false);
}
else {
updateEmptyCarbondataTableStorePath(session, schemaName);
updateEmptyCarbondataTableStorePath(session, finalSchemaName);
targetLocation = carbondataTableStore;
targetLocation = targetLocation + File.separator + schemaName + File.separator + tableName;
targetLocation = targetLocation + File.separator + finalSchemaName + File.separator + tableName;
targetPath = new Path(targetLocation);
}
}
catch (IllegalArgumentException | IOException e) {
throw new PrestoException(NOT_SUPPORTED, format("Error %s store path %s ", e.getMessage(), targetLocation));
}
locationHandle = locationService.forNewTable(metastore, session, schemaName, tableName, Optional.empty(), Optional.of(targetPath), opertionType);
locationHandle = locationService.forNewTable(metastore, session, finalSchemaName, tableName, Optional.empty(), Optional.of(targetPath), opertionType);
return locationHandle;
}
@Override
public void createTable(ConnectorSession session, ConnectorTableMetadata tableMetadata, boolean ignoreExisting)
{
SchemaTableName schemaTableName = tableMetadata.getTable();
String schemaName = schemaTableName.getSchemaName();
String tableName = schemaTableName.getTableName();
SchemaTableName localSchemaTableName = tableMetadata.getTable();
String localSchemaName = localSchemaTableName.getSchemaName();
String tableName = localSchemaTableName.getTableName();
this.user = session.getUser();
this.schemaName = schemaName;
this.schemaName = localSchemaName;
currentState = State.CREATE_TABLE;
List<String> partitionedBy = new ArrayList<String>();
List<SortingColumn> sortBy = new ArrayList<SortingColumn>();
@ -947,29 +947,29 @@ public class CarbondataMetadata
Map<String, String> tableProperties = new HashMap<String, String>();
getParametersForCreateTable(session, tableMetadata, partitionedBy, sortBy, columnHandles, tableProperties);
metastore.getDatabase(schemaName).orElseThrow(() -> new SchemaNotFoundException(schemaName));
metastore.getDatabase(localSchemaName).orElseThrow(() -> new SchemaNotFoundException(localSchemaName));
BaseStorageFormat hiveStorageFormat = CarbondataTableProperties.getCarbondataStorageFormat(tableMetadata.getProperties());
// it will get final path to create carbon table
LocationHandle locationHandle = getCarbonDataTableCreationPath(session, tableMetadata, HiveWriteUtils.OpertionType.CREATE_TABLE);
Path targetPath = locationService.getQueryWriteInfo(locationHandle).getTargetPath();
AbsoluteTableIdentifier absoluteTableIdentifier = AbsoluteTableIdentifier.from(targetPath.toString(),
new CarbonTableIdentifier(schemaName, tableName, UUID.randomUUID().toString()));
AbsoluteTableIdentifier finalAbsoluteTableIdentifier = AbsoluteTableIdentifier.from(targetPath.toString(),
new CarbonTableIdentifier(localSchemaName, tableName, UUID.randomUUID().toString()));
hdfsEnvironment.doAs(session.getUser(), () -> {
initialConfiguration = ConfigurationUtils.toJobConf(this.hdfsEnvironment.getConfiguration(
new HdfsEnvironment.HdfsContext(session, schemaName, tableName),
new HdfsEnvironment.HdfsContext(session, localSchemaName, tableName),
new Path(locationHandle.getJsonSerializableTargetPath())));
CarbondataMetadataUtils.createMetaDataFolderSchemaFile(hdfsEnvironment, session, columnHandles, absoluteTableIdentifier, partitionedBy,
CarbondataMetadataUtils.createMetaDataFolderSchemaFile(hdfsEnvironment, session, columnHandles, finalAbsoluteTableIdentifier, partitionedBy,
sortBy.stream().map(s -> s.getColumnName().toLowerCase(Locale.ENGLISH)).collect(toList()), targetPath.toString(), initialConfiguration);
this.tableStorageLocation = Optional.of(targetPath.toString());
try {
Map<String, String> serdeParameters = initSerDeProperties(tableName);
Table table = buildTableObject(
Table localTable = buildTableObject(
session.getQueryId(),
schemaName,
localSchemaName,
tableName,
session.getUser(),
columnHandles,
@ -981,11 +981,11 @@ public class CarbondataMetadata
true, // carbon table is set as external table
prestoVersion,
serdeParameters);
PrincipalPrivileges principalPrivileges = MetastoreUtil.buildInitialPrivilegeSet(table.getOwner());
HiveBasicStatistics basicStatistics = table.getPartitionColumns().isEmpty() ? HiveBasicStatistics.createZeroStatistics() : HiveBasicStatistics.createEmptyStatistics();
PrincipalPrivileges principalPrivileges = MetastoreUtil.buildInitialPrivilegeSet(localTable.getOwner());
HiveBasicStatistics basicStatistics = localTable.getPartitionColumns().isEmpty() ? HiveBasicStatistics.createZeroStatistics() : HiveBasicStatistics.createEmptyStatistics();
metastore.createTable(
session,
table,
localTable,
principalPrivileges,
Optional.empty(),
ignoreExisting,
@ -1092,8 +1092,8 @@ public class CarbondataMetadata
public CarbondataTableHandle getTableHandle(ConnectorSession session, SchemaTableName tableName)
{
requireNonNull(tableName, "tableName is null");
Optional<Table> table = metastore.getTable(new HiveIdentity(session), tableName.getSchemaName(), tableName.getTableName());
if (!table.isPresent()) {
Optional<Table> finalTable = metastore.getTable(new HiveIdentity(session), tableName.getSchemaName(), tableName.getTableName());
if (!finalTable.isPresent()) {
return null;
}
@ -1102,14 +1102,14 @@ public class CarbondataMetadata
throw new PrestoException(HiveErrorCode.HIVE_INVALID_METADATA, "Unexpected table present in Hive metastore: " + tableName);
}
MetastoreUtil.verifyOnline(tableName, Optional.empty(), MetastoreUtil.getProtectMode(table.get()), table.get().getParameters());
MetastoreUtil.verifyOnline(tableName, Optional.empty(), MetastoreUtil.getProtectMode(finalTable.get()), finalTable.get().getParameters());
return new CarbondataTableHandle(
tableName.getSchemaName(),
tableName.getTableName(),
table.get().getParameters(),
getPartitionKeyColumnHandles(table.get()),
HiveBucketing.getHiveBucketHandle(table.get()));
finalTable.get().getParameters(),
getPartitionKeyColumnHandles(finalTable.get()),
HiveBucketing.getHiveBucketHandle(finalTable.get()));
}
private Optional<ConnectorOutputMetadata> finishUpdateAndDelete(ConnectorSession session,
@ -1133,12 +1133,12 @@ public class CarbondataMetadata
hdfsEnvironment.doAs(user, () -> {
if (blockUpdateDetailsList.size() > 0) {
CarbonTable carbonTable = getCarbonTable(tableHandle.getSchemaName(),
CarbonTable finalCarbonTable = getCarbonTable(tableHandle.getSchemaName(),
tableHandle.getTableName(),
MetastoreUtil.getHiveSchema(table.get()),
initialConfiguration);
SegmentUpdateStatusManager statusManager = new SegmentUpdateStatusManager(carbonTable);
SegmentUpdateStatusManager statusManager = new SegmentUpdateStatusManager(finalCarbonTable);
SegmentUpdateDetails[] segementDetailsList = statusManager.getUpdateStatusDetails();
for (SegmentUpdateDetails segementDetails : segementDetailsList) {
segementDetails.getDeletedRowsInBlock();
@ -1179,26 +1179,26 @@ public class CarbondataMetadata
List<HiveColumnHandle> columnHandles,
Map<String, String> tableProperties)
{
SchemaTableName schemaTableName = tableMetadata.getTable();
String schemaName = schemaTableName.getSchemaName();
String tableName = schemaTableName.getTableName();
SchemaTableName finalSchemaTableName = tableMetadata.getTable();
String finalSchemaName = finalSchemaTableName.getSchemaName();
String finalTableName = finalSchemaTableName.getTableName();
partitionedBy.addAll(CarbondataTableProperties.getPartitionedBy(tableMetadata.getProperties()));
sortBy.addAll(CarbondataTableProperties.getSortedBy(tableMetadata.getProperties()));
Optional<HiveBucketProperty> bucketProperty = Optional.empty();
columnHandles.addAll(getColumnHandles(tableMetadata, ImmutableSet.copyOf(partitionedBy), typeTranslator));
tableProperties.putAll(getEmptyTableProperties(tableMetadata, bucketProperty, new HdfsEnvironment.HdfsContext(session, schemaName, tableName)));
tableProperties.putAll(getEmptyTableProperties(tableMetadata, bucketProperty, new HdfsEnvironment.HdfsContext(session, finalSchemaName, finalTableName)));
}
@Override
public CarbondataOutputTableHandle beginCreateTable(ConnectorSession session, ConnectorTableMetadata tableMetadata, Optional<ConnectorNewTableLayout> layout)
{
// get the root directory for the database
SchemaTableName schemaTableName = tableMetadata.getTable();
String schemaName = schemaTableName.getSchemaName();
String tableName = schemaTableName.getTableName();
SchemaTableName finalSchemaTableName = tableMetadata.getTable();
String finalSchemaName = finalSchemaTableName.getSchemaName();
String finalTableName = finalSchemaTableName.getTableName();
this.user = session.getUser();
this.schemaName = schemaName;
this.schemaName = finalSchemaName;
currentState = State.CREATE_TABLE_AS;
List<String> partitionedBy = new ArrayList<String>();
@ -1206,7 +1206,7 @@ public class CarbondataMetadata
List<HiveColumnHandle> columnHandles = new ArrayList<HiveColumnHandle>();
Map<String, String> tableProperties = new HashMap<String, String>();
getParametersForCreateTable(session, tableMetadata, partitionedBy, sortBy, columnHandles, tableProperties);
metastore.getDatabase(schemaName).orElseThrow(() -> new SchemaNotFoundException(schemaName));
metastore.getDatabase(finalSchemaName).orElseThrow(() -> new SchemaNotFoundException(finalSchemaName));
// to avoid type mismatch between HiveStorageFormat & Carbondata StorageFormat this hack no option
HiveStorageFormat tableStorageFormat = HiveStorageFormat.valueOf("CARBON");
@ -1222,29 +1222,29 @@ public class CarbondataMetadata
// it will get final path to create carbon table
LocationHandle locationHandle = getCarbonDataTableCreationPath(session, tableMetadata, HiveWriteUtils.OpertionType.CREATE_TABLE_AS);
Path targetPath = locationService.getTableWriteInfo(locationHandle, false).getTargetPath();
AbsoluteTableIdentifier absoluteTableIdentifier = AbsoluteTableIdentifier.from(targetPath.toString(),
new CarbonTableIdentifier(schemaName, tableName, UUID.randomUUID().toString()));
AbsoluteTableIdentifier finalAbsoluteTableIdentifier = AbsoluteTableIdentifier.from(targetPath.toString(),
new CarbonTableIdentifier(finalSchemaName, finalTableName, UUID.randomUUID().toString()));
hdfsEnvironment.doAs(session.getUser(), () -> {
initialConfiguration = ConfigurationUtils.toJobConf(this.hdfsEnvironment.getConfiguration(
new HdfsEnvironment.HdfsContext(session, schemaName, tableName),
new HdfsEnvironment.HdfsContext(session, finalSchemaName, finalTableName),
new Path(locationHandle.getJsonSerializableTargetPath())));
// Create Carbondata metadata folder and Schema file
CarbondataMetadataUtils.createMetaDataFolderSchemaFile(hdfsEnvironment, session, columnHandles, absoluteTableIdentifier, partitionedBy,
CarbondataMetadataUtils.createMetaDataFolderSchemaFile(hdfsEnvironment, session, columnHandles, finalAbsoluteTableIdentifier, partitionedBy,
sortBy.stream().map(s -> s.getColumnName().toLowerCase(Locale.ENGLISH)).collect(toList()), targetPath.toString(), initialConfiguration);
this.tableStorageLocation = Optional.of(targetPath.toString());
Path outputPath = new Path(locationHandle.getJsonSerializableTargetPath());
Properties schema = readSchemaForCarbon(schemaName, tableName, targetPath, columnHandles, partitionColumns);
Properties schema = readSchemaForCarbon(finalSchemaName, finalTableName, targetPath, columnHandles, partitionColumns);
// Create committer object
setupCommitWriter(schema, outputPath, initialConfiguration, false);
});
try {
CarbondataOutputTableHandle result = new CarbondataOutputTableHandle(
schemaName,
tableName,
finalSchemaName,
finalTableName,
columnHandles,
metastore.generatePageSinkMetadata(new HiveIdentity(session), schemaTableName),
metastore.generatePageSinkMetadata(new HiveIdentity(session), finalSchemaTableName),
locationHandle,
tableStorageFormat,
partitionStorageFormat,
@ -1255,7 +1255,7 @@ public class CarbondataMetadata
EncodedLoadModel, jobContext.getConfiguration().get(LOAD_MODEL)));
LocationService.WriteInfo writeInfo = locationService.getQueryWriteInfo(locationHandle);
metastore.declareIntentionToWrite(session, writeInfo.getWriteMode(), writeInfo.getWritePath(), schemaTableName);
metastore.declareIntentionToWrite(session, writeInfo.getWriteMode(), writeInfo.getWritePath(), finalSchemaTableName);
return result;
}
catch (RuntimeException ex) {
@ -1386,7 +1386,7 @@ public class CarbondataMetadata
List<Segment> segmentFilesToBeUpdated = blockUpdateDetailsList.stream()
.map(SegmentUpdateDetails::getSegmentName)
.map(Segment::new).collect(Collectors.toList());
List<Segment> segmentFilesToBeUpdatedLatest = new ArrayList<>();
List<Segment> finalSegmentFilesToBeUpdatedLatest = new ArrayList<>();
List<Segment> segmentFilesToBeDeleted = blockUpdateDetailsList.stream()
.filter(segmentUpdateDetails -> segmentUpdateDetails.getSegmentStatus() != null &&
segmentUpdateDetails.getSegmentStatus().equals(SegmentStatus.MARKED_FOR_DELETE))
@ -1396,12 +1396,12 @@ public class CarbondataMetadata
for (Segment segment : segmentFilesToBeUpdated) {
String file =
SegmentFileStore.writeSegmentFile(carbonTable, segment.getSegmentNo(), timeStamp.toString());
segmentFilesToBeUpdatedLatest.add(new Segment(segment.getSegmentNo(), file));
finalSegmentFilesToBeUpdatedLatest.add(new Segment(segment.getSegmentNo(), file));
}
if (!(updateSegmentStatusSuccess &&
CarbonUpdateUtil.updateTableMetadataStatus(new HashSet<>(segmentFilesToBeUpdated),
carbonTable, timeStamp.toString(), true, segmentFilesToBeDeleted,
segmentFilesToBeUpdatedLatest, ""))) {
finalSegmentFilesToBeUpdatedLatest, ""))) {
CarbonUpdateUtil.cleanStaleDeltaFiles(carbonTable, timeStamp.toString());
}
}
@ -1463,11 +1463,10 @@ public class CarbondataMetadata
Properties hiveschema = MetastoreUtil.getHiveSchema(table);
Configuration configuration = jobContext.getConfiguration();
configuration.set(SET_OVERWRITE, "false");
CarbonLoadModel carbonLoadModel =
HiveCarbonUtil.getCarbonLoadModel(hiveschema, configuration);
LoadMetadataDetails loadMetadataDetails = carbonLoadModel.getCurrentLoadMetadataDetail();
carbonLoadModel.setSegmentId(loadMetadataDetails.getLoadName());
CarbonLoaderUtil.recordNewLoadMetadata(loadMetadataDetails, carbonLoadModel, false, true);
CarbonLoadModel loadModel = HiveCarbonUtil.getCarbonLoadModel(hiveschema, configuration);
LoadMetadataDetails loadMetadataDetails = loadModel.getCurrentLoadMetadataDetail();
loadModel.setSegmentId(loadMetadataDetails.getLoadName());
CarbonLoaderUtil.recordNewLoadMetadata(loadMetadataDetails, loadModel, false, true);
}
catch (IOException e) {
LOG.error("Error occurred while committing the insert job.", e);
@ -1554,14 +1553,14 @@ public class CarbondataMetadata
try {
hdfsEnvironment.doAs(session.getUser(), () -> {
metastore.dropTable(session, handle.getSchemaName(), handle.getTableName());
Configuration initialConfiguration = ConfigurationUtils.toJobConf(this.hdfsEnvironment
Configuration finalInitialConfiguration = ConfigurationUtils.toJobConf(this.hdfsEnvironment
.getConfiguration(new HdfsEnvironment.HdfsContext(session, handle.getSchemaName(),
handle.getTableName()), new Path(this.tableStorageLocation.get())));
Properties schema = MetastoreUtil.getHiveSchema(target.get());
schema.setProperty("tablePath", this.tableStorageLocation.get());
this.carbonTable = getCarbonTable(handle.getSchemaName(), handle.getTableName(),
schema, initialConfiguration);
schema, finalInitialConfiguration);
takeLocks(State.DROP_TABLE);
AbsoluteTableIdentifier identifier = this.carbonTable.getAbsoluteTableIdentifier();
if (SegmentStatusManager.isLoadInProgressInTable(carbonTable)) {
@ -1570,7 +1569,7 @@ public class CarbondataMetadata
try {
//Simultaneous case after acquiring locks we should check table exist.
//if table is not there clean the lock folders
carbonTable = getCarbonTable(handle.getSchemaName(), handle.getTableName(), schema, initialConfiguration);
carbonTable = getCarbonTable(handle.getSchemaName(), handle.getTableName(), schema, finalInitialConfiguration);
}//CarbonFileException
catch (RuntimeException e) {
try {
@ -1867,8 +1866,8 @@ public class CarbondataMetadata
{
String tableName = absoluteTableIdentifier.getTableName();
String databaseName = absoluteTableIdentifier.getDatabaseName();
TableInfo tableInfo = carbonTable.getTableInfo();
List<SchemaEvolutionEntry> evolutionEntryList = tableInfo.getFactTable().getSchemaEvolution().getSchemaEvolutionEntryList();
TableInfo finalTableInfo = carbonTable.getTableInfo();
List<SchemaEvolutionEntry> evolutionEntryList = finalTableInfo.getFactTable().getSchemaEvolution().getSchemaEvolutionEntryList();
Long updatedTime = evolutionEntryList.get(evolutionEntryList.size() - 1).getTimeStamp();
LOG.info("Reverting changes for " + databaseName + "." + tableName);
List<ColumnSchema> addedSchemas = evolutionEntryList.get(evolutionEntryList.size() - 1).getAdded();
@ -1880,7 +1879,7 @@ public class CarbondataMetadata
break;
}
case DROP_COLUMN: {
tableInfo.getFactTable().getListOfColumns().forEach(cols -> removedSchemas.forEach(removedCols -> {
finalTableInfo.getFactTable().getListOfColumns().forEach(cols -> removedSchemas.forEach(removedCols -> {
if (cols.isInvisible() && removedCols.getColumnUniqueId().equals(cols.getColumnUniqueId())) {
cols.setInvisible(false);
}
@ -2027,38 +2026,38 @@ public class CarbondataMetadata
@Override
protected ConnectorTableMetadata doGetTableMetadata(ConnectorSession session, SchemaTableName tableName)
{
Optional<Table> table = metastore.getTable(new HiveIdentity(session), tableName.getSchemaName(), tableName.getTableName());
if (!table.isPresent() || table.get().getTableType().equals(TableType.VIRTUAL_VIEW.name())) {
Optional<Table> finalTable = metastore.getTable(new HiveIdentity(session), tableName.getSchemaName(), tableName.getTableName());
if (!finalTable.isPresent() || finalTable.get().getTableType().equals(TableType.VIRTUAL_VIEW.name())) {
throw new TableNotFoundException(tableName);
}
Function<HiveColumnHandle, ColumnMetadata> metadataGetter = columnMetadataGetter(table.get(), typeManager);
Function<HiveColumnHandle, ColumnMetadata> metadataGetter = columnMetadataGetter(finalTable.get(), typeManager);
ImmutableList.Builder<ColumnMetadata> columns = ImmutableList.builder();
for (HiveColumnHandle columnHandle : hiveColumnHandles(table.get())) {
for (HiveColumnHandle columnHandle : hiveColumnHandles(finalTable.get())) {
columns.add(metadataGetter.apply(columnHandle));
}
// External location property
ImmutableMap.Builder<String, Object> properties = ImmutableMap.builder();
properties.put(LOCATION_PROPERTY, table.get().getStorage().getLocation());
properties.put(LOCATION_PROPERTY, finalTable.get().getStorage().getLocation());
// Storage format property
properties.put(HiveTableProperties.STORAGE_FORMAT_PROPERTY, CarbondataStorageFormat.CARBON);
// Partitioning property
List<String> partitionedBy = table.get().getPartitionColumns().stream()
List<String> partitionedBy = finalTable.get().getPartitionColumns().stream()
.map(Column::getName)
.collect(toList());
if (!partitionedBy.isEmpty()) {
properties.put(HiveTableProperties.PARTITIONED_BY_PROPERTY, partitionedBy);
}
Optional<String> comment = Optional.ofNullable(table.get().getParameters().get(TABLE_COMMENT));
Optional<String> comment = Optional.ofNullable(finalTable.get().getParameters().get(TABLE_COMMENT));
// add partitioned columns into immutableColumns
ImmutableList.Builder<ColumnMetadata> immutableColumns = ImmutableList.builder();
for (HiveColumnHandle columnHandle : hiveColumnHandles(table.get())) {
for (HiveColumnHandle columnHandle : hiveColumnHandles(finalTable.get())) {
if (columnHandle.getColumnType().equals(HiveColumnHandle.ColumnType.PARTITION_KEY)) {
immutableColumns.add(metadataGetter.apply(columnHandle));
}

View File

@ -191,7 +191,7 @@ public class CarbondataMetadataFactory
@Override
public HiveMetadata get()
{
SemiTransactionalHiveMetastore metastore =
SemiTransactionalHiveMetastore semiTransactionalHiveMetastore =
new SemiTransactionalHiveMetastore(this.hdfsEnvironment,
CachingHiveMetastore.memoizeMetastore(this.metastore, this.perTransactionCacheMaximumSize),
this.renameExecution,
@ -200,7 +200,7 @@ public class CarbondataMetadataFactory
this.hiveTransactionHeartbeatInterval,
this.heartbeatService, hiveMetastoreClientService, hmsWriteBatchSize);
return new CarbondataMetadata(metastore,
return new CarbondataMetadata(semiTransactionalHiveMetastore,
this.hdfsEnvironment,
this.partitionManager,
this.writesToNonManagedTablesEnabled,
@ -212,8 +212,8 @@ public class CarbondataMetadataFactory
this.segmentInfoCodec,
this.typeTranslator,
this.hetuVersion,
new MetastoreHiveStatisticsProvider(metastore, statsCache, samplePartitionCache),
this.accessControlMetadataFactory.create(metastore),
new MetastoreHiveStatisticsProvider(semiTransactionalHiveMetastore, statsCache, samplePartitionCache),
this.accessControlMetadataFactory.create(semiTransactionalHiveMetastore),
carbondataTableReader,
this.carbondataTableStore,
this.carbondataMajorVacuumSegmentSize,

View File

@ -167,9 +167,9 @@ public class CarbondataPageSink
{
//set flag here if called and change finish accordingly.
isCompactionCalled = true;
HdfsEnvironment hdfsEnvironment = connectorPageSource.getHdfsEnvironment();
HdfsEnvironment finalHdfsEnvironment = connectorPageSource.getHdfsEnvironment();
hdfsEnvironment.doAs(session.getUser(), () -> {
finalHdfsEnvironment.doAs(session.getUser(), () -> {
try {
// Worker part: each thread to run this code
boolean mergeStatus = false;

View File

@ -163,6 +163,7 @@ public class CarbondataPageSinkProvider
ImmutableMap.of(), handle.getAdditionalConf(), false);
}
@Override
public ConnectorPageSink createPageSink(ConnectorTransactionHandle transaction, ConnectorSession session, ConnectorOutputTableHandle tableHandle)
{
CarbondataOutputTableHandle handle = (CarbondataOutputTableHandle) tableHandle;

View File

@ -232,15 +232,15 @@ public class CarbondataPageSource
nanoStart = System.nanoTime();
}
CarbondataVectorBatch columnarBatch = null;
int batchSize = 0;
int columnBatchSize = 0;
try {
batchId++;
if (vectorReader.nextKeyValue()) {
Object vectorBatch = vectorReader.getCurrentValue();
if (vectorBatch instanceof CarbondataVectorBatch) {
columnarBatch = (CarbondataVectorBatch) vectorBatch;
batchSize = columnarBatch.numRows();
if (batchSize == 0) {
columnBatchSize = columnarBatch.numRows();
if (columnBatchSize == 0) {
close();
return null;
}
@ -256,9 +256,9 @@ public class CarbondataPageSource
Block[] blocks = new Block[columnHandles.size()];
for (int column = 0; column < blocks.length; column++) {
blocks[column] = new LazyBlock(batchSize, new CarbondataBlockLoader(column));
blocks[column] = new LazyBlock(columnBatchSize, new CarbondataBlockLoader(column));
}
Page page = new Page(batchSize, blocks);
Page page = new Page(columnBatchSize, blocks);
return page;
}
catch (PrestoException e) {

View File

@ -90,16 +90,17 @@ public class CarbondataTableProperties
private static SortingColumn sortingColumnFromString(String name)
{
String finalName = name;
SortingColumn.Order order = SortingColumn.Order.ASCENDING;
String lower = name.toUpperCase(ENGLISH);
if (lower.endsWith(" ASC")) {
name = name.substring(0, name.length() - 4).trim();
finalName = name.substring(0, name.length() - 4).trim();
}
else if (lower.endsWith(" DESC")) {
name = name.substring(0, name.length() - 5).trim();
finalName = name.substring(0, name.length() - 5).trim();
order = SortingColumn.Order.DESCENDING;
}
return new SortingColumn(name, order);
return new SortingColumn(finalName, order);
}
private static String sortingColumnToString(SortingColumn column)

View File

@ -133,6 +133,7 @@ public class CarbondataWriterFactory
{
}
@Override
protected void setAdditionalSchemaProperties(Properties schema)
{
schema.setProperty(META_TABLE_LOCATION, locationService.getTableWriteInfo(locationHandle, false).getTargetPath().toString());

View File

@ -143,14 +143,15 @@ class ColumnarVectorWrapperDirect
@Override
public void putDecimals(int rowId, int count, BigDecimal value, int precision)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (nullBitSet.get(rowId)) {
columnVector.putNull(rowId);
if (nullBitSet.get(inputRowId)) {
columnVector.putNull(inputRowId);
}
else {
columnVector.putDecimal(rowId, value, precision);
columnVector.putDecimal(inputRowId, value, precision);
}
rowId++;
inputRowId++;
}
}
@ -185,8 +186,9 @@ class ColumnarVectorWrapperDirect
@Override
public void putByteArray(int rowId, int count, byte[] value)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
columnVector.putByteArray(rowId++, value);
columnVector.putByteArray(inputRowId++, value);
}
}
@ -303,84 +305,90 @@ class ColumnarVectorWrapperDirect
@Override
public void putFloats(int rowId, int count, float[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (nullBitSet.get(rowId)) {
columnVector.putNull(rowId);
if (nullBitSet.get(inputRowId)) {
columnVector.putNull(inputRowId);
}
else {
columnVector.putFloat(rowId, src[i]);
columnVector.putFloat(inputRowId, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putShorts(int rowId, int count, short[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (nullBitSet.get(rowId)) {
columnVector.putNull(rowId);
if (nullBitSet.get(inputRowId)) {
columnVector.putNull(inputRowId);
}
else {
columnVector.putShort(rowId, src[i]);
columnVector.putShort(inputRowId, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putInts(int rowId, int count, int[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (nullBitSet.get(rowId)) {
columnVector.putNull(rowId);
if (nullBitSet.get(inputRowId)) {
columnVector.putNull(inputRowId);
}
else {
columnVector.putInt(rowId, src[i]);
columnVector.putInt(inputRowId, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putLongs(int rowId, int count, long[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (nullBitSet.get(rowId)) {
columnVector.putNull(rowId);
if (nullBitSet.get(inputRowId)) {
columnVector.putNull(inputRowId);
}
else {
columnVector.putLong(rowId, src[i]);
columnVector.putLong(inputRowId, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putDoubles(int rowId, int count, double[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (nullBitSet.get(rowId)) {
columnVector.putNull(rowId);
if (nullBitSet.get(inputRowId)) {
columnVector.putNull(inputRowId);
}
else {
columnVector.putDouble(rowId, src[i]);
columnVector.putDouble(inputRowId, src[i]);
}
rowId++;
inputRowId++;
}
}
@Override
public void putBytes(int rowId, int count, byte[] src, int srcIndex)
{
int inputRowId = rowId;
for (int i = 0; i < count; i++) {
if (nullBitSet.get(rowId)) {
columnVector.putNull(rowId);
if (nullBitSet.get(inputRowId)) {
columnVector.putNull(inputRowId);
}
else {
columnVector.putByte(rowId, src[i]);
columnVector.putByte(inputRowId, src[i]);
}
rowId++;
inputRowId++;
}
}

View File

@ -58,8 +58,9 @@ public class BooleanStreamReader
@Override
public void putBytes(int rowId, int count, byte[] src, int srcIndex)
{
int srcIdx = srcIndex;
for (int i = 0; i < count; i++) {
type.writeBoolean(builder, src[srcIndex++] == 1);
type.writeBoolean(builder, src[srcIdx++] == 1);
}
}

View File

@ -76,8 +76,9 @@ public class DecimalSliceStreamReader
@Override
public void putDecimals(int rowId, int count, BigDecimal value, int precision)
{
int id = rowId;
for (int i = 0; i < count; i++) {
putDecimal(rowId++, value, precision);
putDecimal(id++, value, precision);
}
}

View File

@ -58,8 +58,9 @@ public class IntegerStreamReader
@Override
public void putInts(int rowId, int count, int value)
{
int id = rowId;
for (int i = 0; i < count; i++) {
putInt(rowId++, value);
putInt(id++, value);
}
}

View File

@ -15,12 +15,10 @@
package io.hetu.core.plugin.carbondata.integrationtest;
import com.esotericsoftware.minlog.Log;
import com.google.gson.Gson;
import io.hetu.core.plugin.carbondata.server.HetuTestServer;
import io.prestosql.hive.$internal.au.com.bytecode.opencsv.CSVReader;
import io.prestosql.spi.PrestoException;
import io.prestosql.spi.StandardErrorCode;
import org.apache.carbondata.common.logging.LogServiceFactory;
import org.apache.carbondata.core.constants.CarbonCommonConstants;
import org.apache.carbondata.core.datastore.filesystem.CarbonFile;
@ -53,7 +51,6 @@ import java.io.FileReader;
import java.io.IOException;
import java.math.BigDecimal;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;
import java.sql.SQLException;
import java.text.DateFormat;
@ -67,10 +64,8 @@ import java.util.List;
import java.util.Map;
import java.util.TreeMap;
import java.util.stream.Collectors;
import java.util.stream.Stream;
import static io.prestosql.spi.StandardErrorCode.GENERIC_INTERNAL_ERROR;
import static io.prestosql.spi.StandardErrorCode.NOT_SUPPORTED;
import static org.testng.Assert.assertEquals;
import static org.testng.Assert.assertFalse;
import static org.testng.Assert.assertTrue;
@ -1429,7 +1424,7 @@ public class TestCarbonAllDataType
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtabledrop", false), false);
} catch (IOException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
}
@ -1474,7 +1469,7 @@ public class TestCarbonAllDataType
i++;
}
} catch (IOException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
// Step 3: convert format level TableInfo to code level TableInfo
@ -1555,7 +1550,7 @@ public class TestCarbonAllDataType
try {
date = inputFormat.parse(data);
} catch (ParseException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
String dateString = outuptformat.format(date);
dateString = "date '" + dateString + "'";
@ -1569,7 +1564,7 @@ public class TestCarbonAllDataType
try {
date = inputFormat.parse(data);
} catch (ParseException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
String dateString = outuptformat.format(date);
dateString = "date '" + dateString + "'";
@ -1578,7 +1573,7 @@ public class TestCarbonAllDataType
return "date '" + data + "'";
}
case "varchar":
{//'china'
{
return "'" + data + "'";
}
case "timestamp":
@ -1592,7 +1587,7 @@ public class TestCarbonAllDataType
try {
date = inputFormattime.parse(data);
} catch (ParseException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
dateString = outuptformattime.format(date);
}
@ -1603,7 +1598,7 @@ public class TestCarbonAllDataType
try {
date = inputFormattime.parse(data);
} catch (ParseException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
dateString = outuptformattime.format(date);
}
@ -1614,7 +1609,7 @@ public class TestCarbonAllDataType
try {
date = inputFormattime.parse(data);
} catch (ParseException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
dateString = outuptformattime.format(date);
}
@ -1625,7 +1620,7 @@ public class TestCarbonAllDataType
try {
date = inputFormattime.parse(data);
} catch (ParseException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
dateString = outuptformattime.format(date);
}
@ -1637,9 +1632,10 @@ public class TestCarbonAllDataType
return dateString;
}
case "smallint": {
// smallint '12'
return "smallint '" + data + "'";
}
default:
break;
}
return data;
}
@ -1697,7 +1693,7 @@ public class TestCarbonAllDataType
hetuServer.execute(inserData);
}
catch(Exception e) {
e.printStackTrace();
logger.error(e.getMessage());
}
}
@ -1736,7 +1732,7 @@ public class TestCarbonAllDataType
}
catch (IOException | InterruptedException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
}
@ -1785,14 +1781,15 @@ public class TestCarbonAllDataType
*/
private boolean checkStatusFileForDeleteMarked(String tableName, int updateNumber, int segmentNumber) throws SQLException
{
BufferedReader reader = null;
try {
File dir = new File(storePath + "/carbon.store/testdb/" + tableName + "/Metadata");
File[] tableUpdateStatusFiles = dir.listFiles((d, name) -> name.startsWith("tableupdatestatus"));
Arrays.sort(tableUpdateStatusFiles);
Gson gson = new Gson();
BufferedReader reader = new BufferedReader(new FileReader(tableUpdateStatusFiles[updateNumber]));
reader = new BufferedReader(new FileReader(tableUpdateStatusFiles[updateNumber]));
SegmentUpdateDetails[] segmentUpdateDetails = gson.fromJson(reader, SegmentUpdateDetails[].class);
File tableStatusFile = new File(dir.getAbsolutePath() + "/tablestatus");
File tableStatusFile = new File(dir.getCanonicalPath() + "/tablestatus");
reader = new BufferedReader(new FileReader(tableStatusFile));
LoadMetadataDetails loadMetadataDetails = gson.fromJson(reader, LoadMetadataDetails[].class)[segmentNumber];
if ((segmentUpdateDetails[0].getSegmentStatus() != null && segmentUpdateDetails[0].getSegmentStatus().toString().equals("Marked for Delete")) &&
@ -1803,6 +1800,16 @@ public class TestCarbonAllDataType
hetuServer.execute("drop table if exists testdb." + tableName);
Assert.fail("Failed to read status files");
}
finally {
if (reader != null) {
try {
reader.close();
}
catch (IOException e) {
logger.error(e.getMessage());
}
}
}
return false;
}
@ -1822,7 +1829,7 @@ public class TestCarbonAllDataType
"/carbon.store/testdb/mytesttable/Fact/Part0/Segment_0.1", false), true);
} catch (IOException e) {
hetuServer.execute("DROP TABLE if exists testdb.mytesttable");
e.printStackTrace();
logger.error(e.getMessage());
}
hetuServer.execute("DROP TABLE if exists testdb.mytesttable");
@ -1846,7 +1853,7 @@ public class TestCarbonAllDataType
}
catch (IOException e) {
hetuServer.execute("DROP TABLE if exists testdb.mytesttable2");
e.printStackTrace();
logger.error(e.getMessage());
}
hetuServer.execute("DROP TABLE if exists testdb.mytesttable2");
@ -1879,7 +1886,7 @@ public class TestCarbonAllDataType
}
catch (IOException | InterruptedException e) {
hetuServer.execute("DROP TABLE if exists testdb.myectable");
e.printStackTrace();
logger.error(e.getMessage());
}
hetuServer.execute("DROP TABLE if exists testdb.myectable");
@ -1904,7 +1911,7 @@ public class TestCarbonAllDataType
FileFactory.mkdirs( storePath + "/carbon.store/mytestDb");
}
} catch (IOException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
String location = "'" + "file:///" + storePath + "/carbon.store/mytestDb" + "')" ;
@ -2487,91 +2494,44 @@ public class TestCarbonAllDataType
@Test
public void test_writer_count() throws SQLException
{
hetuServer.execute("drop table if exists testdb.testorders1");
hetuServer.execute("drop table if exists testdb.testorders_bak");
hetuServer.execute("set session task_writer_count=32");
hetuServer.execute("set session implicit_conversion=true");
hetuServer.execute("CREATE TABLE testdb.testorders1(orderkey int, orderstatus STRING, totalprice double, orderdate date)");
hetuServer.execute("INSERT INTO testdb.testorders1 VALUES(10,'SUCCESS', 125.15, DATE'2919-05-17')");
hetuServer.execute("INSERT INTO testdb.testorders1 VALUES(20,'SUCCESS', 125.15, DATE'2919-05-17')");
hetuServer.execute("INSERT INTO testdb.testorders1 VALUES(30,'SUCCESS', 125.15, DATE'2919-05-17')");
hetuServer.execute("CREATE TABLE testdb.testorders_bak(orderkey bigint, orderstatus varchar(7), totalprice double, orderdate date)");
hetuServer.execute("insert into testdb.testorders_bak(orderkey, orderstatus, totalprice) select orderkey, orderstatus, totalprice from testorders1");
hetuServer.execute("CREATE TABLE testdb.testWriterCount_32(orderkey int, orderstatus STRING, totalprice double, orderdate date)");
hetuServer.execute("INSERT INTO testdb.testWriterCount_32 VALUES(10,'SUCCESS', 125.15, DATE'2919-05-17')");
hetuServer.execute("INSERT INTO testdb.testWriterCount_32 VALUES(20,'SUCCESS', 125.15, DATE'2919-05-17')");
hetuServer.execute("INSERT INTO testdb.testWriterCount_32 VALUES(30,'SUCCESS', 125.15, DATE'2919-05-17')");
hetuServer.execute("set session task_writer_count=1");
List<Map<String, Object>> actualResult = hetuServer.executeQuery("Select count (*) as RESULT from testdb.testorders_bak");
hetuServer.execute("INSERT INTO testdb.testWriterCount_32 SELECT * FROM testdb.testWriterCount_32");
verifyRowCount("testdb.testWriterCount_32", 6);
List<Map<String, Object>> expectedResult = new ArrayList<Map<String, Object>>() {{
add(new HashMap<String, Object>() {{ put("RESULT", 3); }});
}};
Assert.assertEquals(actualResult.toString(), expectedResult.toString());
hetuServer.execute("INSERT INTO testdb.testWriterCount_32 SELECT * FROM testdb.testWriterCount_32");
verifyRowCount("testdb.testWriterCount_32", 12);
// checking correct number of files created in Fact
File files1 = null;
files1 = new File(storePath + "/carbon.store/testdb/testorders_bak/Fact");
File[] fileListPart0 = files1.listFiles();//Part0
if (fileListPart0.length == 1) {
File[] Segment_0 = fileListPart0[0].listFiles();
if (Segment_0.length == 1) {
File[] fileListSegment_0 = Segment_0[0].listFiles();
if (fileListSegment_0.length == 6) {
assertEquals("true", "true");
}
assertEquals(fileListSegment_0.length, 6);
}
assertEquals(Segment_0.length, 1);
}
assertEquals(fileListPart0.length, 1);
//test with 16
hetuServer.execute("set session task_writer_count=16");
hetuServer.execute("insert into testdb.testorders_bak(orderkey, orderstatus, totalprice) select orderkey, orderstatus, totalprice from testorders1");
hetuServer.execute("set session task_writer_count=1");
actualResult = hetuServer.executeQuery("Select count (*) as RESULT from testdb.testorders_bak");
expectedResult = new ArrayList<Map<String, Object>>() {{
add(new HashMap<String, Object>() {{ put("RESULT", 6); }});
}};
Assert.assertEquals(actualResult.toString(), expectedResult.toString());
//test with 8
hetuServer.execute("set session task_writer_count=8");
hetuServer.execute("insert into testdb.testorders_bak(orderkey, orderstatus, totalprice) select orderkey, orderstatus, totalprice from testorders1");
hetuServer.execute("set session task_writer_count=1");
actualResult = hetuServer.executeQuery("Select count (*) as RESULT from testdb.testorders_bak");
expectedResult = new ArrayList<Map<String, Object>>() {{
add(new HashMap<String, Object>() {{ put("RESULT", 9); }});
}};
Assert.assertEquals(actualResult.toString(), expectedResult.toString());
//test with 2
hetuServer.execute("set session task_writer_count=2");
hetuServer.execute("insert into testdb.testorders_bak(orderkey, orderstatus, totalprice) select orderkey, orderstatus, totalprice from testorders1");
hetuServer.execute("set session task_writer_count=1");
actualResult = hetuServer.executeQuery("Select count (*) as RESULT from testdb.testorders_bak");
expectedResult = new ArrayList<Map<String, Object>>() {{
add(new HashMap<String, Object>() {{ put("RESULT", 12); }});
}};
Assert.assertEquals(actualResult.toString(), expectedResult.toString());
hetuServer.execute("INSERT INTO testdb.testWriterCount_32 SELECT * FROM testdb.testWriterCount_32");
verifyRowCount("testdb.testWriterCount_32", 24);
String filePath = storePath + "/carbon.store/testdb/testwritercount_32";
try {
hetuServer.execute("set session task_writer_count=32");
hetuServer.execute("VACUUM TABLE testdb.testorders_bak AND WAIT");
assertEquals(FileFactory.isFileExist(storePath +
"/carbon.store/testdb/testorders_bak/Fact/Part0/Segment_0.1", false), true);
} catch (IOException e) {
hetuServer.execute("VACUUM TABLE testdb.testWriterCount_32 AND WAIT");
assertTrue(FileFactory.isFileExist(filePath + "/Fact/Part0/Segment_0.1"));
}
catch (IOException e) {
assertTrue(false, "Unable to read file from table path");
}
finally {
hetuServer.execute("drop table if exists testdb.testWriterCount_32");
hetuServer.execute("set session task_writer_count=1");
}
}
hetuServer.execute("drop table if exists testdb.testorders1");
hetuServer.execute("drop table if exists testdb.testorders_bak");
hetuServer.execute("set session task_writer_count=1");
private void verifyRowCount(String tableName, int rowCount) throws SQLException
{
String query = String.format("SELECT COUNT(*) AS result FROM %s", tableName);
List<Map<String, Object>> actualResult = hetuServer.executeQuery(query);
List<Map<String, Object>> expectedResult = new ArrayList<Map<String, Object>>() {{
add(new HashMap<String, Object>() {{ put("result", rowCount); }});
}};
Assert.assertEquals(actualResult.toString(), expectedResult.toString());
}
private TableInfo getTableInfoFromSchemaFile(String tableName)

View File

@ -40,6 +40,7 @@ import java.util.List;
import java.util.Map;
import static org.testng.Assert.assertTrue;
import static org.testng.AssertJUnit.assertNotNull;
public class TestCarbonAutoVacuum
{
@ -99,7 +100,6 @@ public class TestCarbonAutoVacuum
@AfterClass
public void tearDown() throws SQLException, IOException, InterruptedException
{
//hetuServer.execute("drop table if exists hive.default.demotable");
logger.info("TearDown begin: " + this.getClass().getSimpleName());
hetuServer.stopServer();
CarbonUtil.deleteFoldersAndFiles(FileFactory.getCarbonFile(storePath));
@ -145,11 +145,12 @@ public class TestCarbonAutoVacuum
try {
CarbondataAutoVacuumThread.enableTracingVacuumTask(true);
assertNotNull(catalog);
connector = catalog.getConnector(catalog.getConnectorCatalogName());
connectorMetadata = connector.getConnectorMetadata();
connectorMetadata.getTablesForVacuum();
} catch (Exception e) {
logger.debug(e.getMessage());
}
CarbondataAutoVacuumThread.waitForSubmittedVacuumTasksFinish();
@ -215,7 +216,7 @@ public class TestCarbonAutoVacuum
connectorMetadata = connector.getConnectorMetadata();
connectorMetadata.getTablesForVacuum();
} catch (Exception e) {
logger.debug(e.getMessage());
}
CarbondataAutoVacuumThread.waitForSubmittedVacuumTasksFinish();

View File

@ -62,7 +62,6 @@ public class TestCarbondataAutoCleanup
public void setup() throws Exception
{
logger.info("Setup begin: " + this.getClass().getSimpleName());
String dataPath = rootPath + "/src/test/resources/alldatatype.csv";
CarbonProperties.getInstance().addProperty(CarbonCommonConstants.CARBON_WRITTEN_BY_APPNAME, "HetuTest");
CarbonProperties.getInstance().addProperty(CarbonCommonConstants.MAX_QUERY_EXECUTION_TIME, "0");
@ -131,7 +130,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup1/Fact/Part0/Segment_3", false), false);
}
catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -160,7 +159,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup2/Fact/Part0/Segment_2", false), false);
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup2/Fact/Part0/Segment_3", false), false);
} catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -190,7 +189,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup3/Fact/Part0/Segment_2", false), false);
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup3/Fact/Part0/Segment_3", false), false);
} catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -222,7 +221,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanupwithpushdown/Fact/Part0/Segment_2", false), false);
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanupwithpushdown/Fact/Part0/Segment_3", false), false);
} catch (IOException exception) {
logger.debug(exception.getMessage());
}
}
finally {
@ -255,7 +254,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup4/Fact/Part0/Segment_2", false), false);
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup4/Fact/Part0/Segment_3", false), false);
} catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -286,7 +285,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup5/Fact/Part0/Segment_3", false), false);
}
catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -315,7 +314,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup6/Fact/Part0/Segment_2", false), false);
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup6/Fact/Part0/Segment_3", false), false);
} catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -345,7 +344,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup7/Fact/Part0/Segment_2", false), false);
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup7/Fact/Part0/Segment_3", false), false);
} catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -375,7 +374,7 @@ public class TestCarbondataAutoCleanup
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup8/Fact/Part0/Segment_2", false), false);
assertEquals(FileFactory.isFileExist(storePath + "/carbon.store/testdb/testtableautocleanup8/Fact/Part0/Segment_3", false), false);
} catch (IOException exception) {
logger.debug(exception.getMessage());
}
CarbondataMetadata.enableTracingCleanupTask(false);
@ -404,7 +403,7 @@ public class TestCarbondataAutoCleanup
content = content.replaceFirst(modificationOrdeletionTimesStamp, replace);
Files.write(path, content.getBytes(charset));
} catch (IOException e) {
e.printStackTrace();
logger.error(e.getMessage());
}
}
}

View File

@ -99,7 +99,7 @@ public class TestCarbondataMinorConfig
"/carbon.store/mytestdb/mytesttable/Fact/Part0/Segment_0.1", false), true);
} catch (IOException e) {
hetuServer.execute("DROP TABLE if exists mytestdb.mytesttable");
e.printStackTrace();
logger.error(e.getMessage());
}
hetuServer.execute("DROP TABLE if exists mytestdb.mytesttable");

View File

@ -15,6 +15,7 @@
package io.hetu.core.plugin.carbondata.integrationtest;
import io.airlift.log.Logger;
import io.hetu.core.plugin.carbondata.server.HetuTestServer;
import org.apache.carbondata.core.constants.CarbonCommonConstants;
import org.apache.carbondata.core.datastore.impl.FileFactory;
@ -36,6 +37,7 @@ import static org.testng.Assert.assertTrue;
public class TestsWithHiveConnector
{
private static final Logger log = Logger.get(TestsWithHiveConnector.class);
private String rootPath = new File(this.getClass().getResource("/").getPath() + "../..")
.getCanonicalPath();
@ -114,7 +116,7 @@ public class TestsWithHiveConnector
assertEquals(FileFactory.isFileExist(storePath +
"hive.store/default/parttable/year=2013", false), false);
} catch (IOException exception) {
exception.printStackTrace();
log.error(exception.getMessage());
}
hetuServer.execute("DROP TABLE hive.default.parttable");
}
@ -141,7 +143,7 @@ public class TestsWithHiveConnector
assertEquals(FileFactory.isFileExist(storePath +
"hive.store/default/parttable2/year=2013", false), false);
} catch (IOException exception) {
exception.printStackTrace();
log.error(exception.getMessage());
}
hetuServer.execute("insert into hive.default.parttable2 values (4,2014)");
@ -150,7 +152,7 @@ public class TestsWithHiveConnector
assertEquals(FileFactory.isFileExist(storePath +
"hive.store/default/parttable2/year=2014", false), false);
} catch (IOException exception) {
exception.printStackTrace();
log.error(exception.getMessage());
}
hetuServer.execute("DROP TABLE hive.default.parttable2");
@ -178,7 +180,7 @@ public class TestsWithHiveConnector
assertEquals(FileFactory.isFileExist(storePath +
"/hive.store/default/parttable3/year=2013", false), true);
} catch (IOException exception) {
exception.printStackTrace();
log.error(exception.getMessage());
}
hetuServer.execute("DROP TABLE hive.default.parttable3");
}

View File

@ -91,11 +91,11 @@ public class HetuTestServer
carbonProperties.putAll(properties);
logger.info("------------ Starting Presto Server -------------");
DistributedQueryRunner queryRunner = createQueryRunner(hetuProperties);
DistributedQueryRunner distributedQueryRunner = createQueryRunner(hetuProperties);
Connection connection = createJdbcConnection(dbName);
statement = (PrestoStatement) connection.createStatement();
logger.info("STARTED SERVER AT :" + queryRunner.getCoordinator().getBaseUrl());
logger.info("STARTED SERVER AT :" + distributedQueryRunner.getCoordinator().getBaseUrl());
}
public void stopServer() throws SQLException
@ -123,14 +123,20 @@ public class HetuTestServer
public List<Map<String, Object>> executeQuery(String query) throws SQLException
{
logger.info(">>>>> Executing Query: " + query);
ResultSet rs = null;
try {
ResultSet rs = statement.executeQuery(query);
rs = statement.executeQuery(query);
return convertResultSetToList(rs);
}
catch (SQLException e) {
logger.error("Exception Occured: " + e.getMessage() + "\n Failed Query: " + query);
throw e;
}
finally {
if (rs != null) {
rs.close();
}
}
}
private List<Map<String, Object>> convertResultSetToList(ResultSet rs) throws SQLException
@ -186,7 +192,7 @@ public class HetuTestServer
{
try {
queryRunner.installPlugin(new CarbondataPlugin());
Map<String, String> carbonProperties = ImmutableMap.<String, String>builder()
Map<String, String> carbonPropertiesMap = ImmutableMap.<String, String>builder()
.putAll(this.carbonProperties)
.put("carbon.unsafe.working.memory.in.mb", "512")
.build();
@ -197,7 +203,7 @@ public class HetuTestServer
.build();
// CreateCatalog will create a catalog for CarbonData in etc/catalog.
queryRunner.createCatalog(carbonDataCatalog, carbonDataConnector, carbonProperties);
queryRunner.createCatalog(carbonDataCatalog, carbonDataConnector, carbonPropertiesMap);
queryRunner.createCatalog(carbonDataCatalogLocationDisabled, carbonDataConnector, carbonPropertiesLocationDisabled);
}
catch (RuntimeException e) {

View File

@ -5,7 +5,7 @@
<parent>
<groupId>io.hetu.core</groupId>
<artifactId>presto-root</artifactId>
<version>1.4.0</version>
<version>1.7.0-SNAPSHOT</version>
</parent>
<artifactId>hetu-clickhouse</artifactId>

View File

@ -345,8 +345,9 @@ public class ClickHouseClient
}
@Override
public void renameColumn(JdbcIdentity identity, JdbcTableHandle handle, JdbcColumnHandle jdbcColumn, String newColumnName)
public void renameColumn(JdbcIdentity identity, JdbcTableHandle handle, JdbcColumnHandle jdbcColumn, String inputNewColumnName)
{
String newColumnName = inputNewColumnName;
try (Connection connection = connectionFactory.openConnection(identity)) {
if (connection.getMetaData().storesUpperCaseIdentifiers()) {
newColumnName = newColumnName.toUpperCase(ENGLISH);

View File

@ -41,6 +41,7 @@ public class ClickHouseApplyRemoteFunctionPushDown
/**
* rewrite the remote function to a executable function in the data source.
*/
@Override
public Optional<String> rewriteRemoteFunction(CallExpression callExpression, BaseJdbcRowExpressionConverter rowExpressionConverter, JdbcConverterContext jdbcConverterContext)
{
if (!isConnectorSupportedRemoteFunction(callExpression)) {

View File

@ -37,8 +37,9 @@ public class ClickHouseSqlStatementWriter
}
@Override
public String aggregation(String functionName, List<String> arguments, boolean isDistinct)
public String aggregation(String inputFunctionName, List<String> arguments, boolean isDistinct)
{
String functionName = inputFunctionName;
if (functionName.toUpperCase(Locale.ENGLISH).equals("VARIANCE")) {
functionName = "varPop";
}

View File

@ -205,7 +205,7 @@ public final class ClickHouseServerTest
{
String actualTable = tablePattern;
for (String table : tables) { //tableName + _ + UUID
for (String table : tables) {
int lastIndex = table.lastIndexOf("_");
if (lastIndex == -1) {
continue;

View File

@ -5,7 +5,7 @@
<parent>
<groupId>io.hetu.core</groupId>
<artifactId>presto-root</artifactId>
<version>1.4.0</version>
<version>1.7.0-SNAPSHOT</version>
</parent>
<artifactId>hetu-common</artifactId>

View File

@ -47,16 +47,6 @@ public class SslSocketUtil
if (!tlsEnabled) {
return Optional.empty();
}
// https://docs.oracle.com/javase/8/docs/technotes/guides/security/jsse/JSSERefGuide.html#CustomizingStores
// as per link above, the default SSLContext will be constructed using the default KeyManager and
// default TrustManager. Those can be configured using the following system properties:
// javax.net.ssl.keyStore
// javax.net.ssl.keyStorePassword
// javax.net.ssl.keyStoreType
// javax.net.ssl.trustStore
// javax.net.ssl.trustStorePassword
// see link above for more details
return Optional.of(SSLContext.getDefault());
}

View File

@ -13,6 +13,7 @@
*/
package io.hetu.core.common.util;
import io.airlift.log.Logger;
import io.airlift.security.pem.PemReader;
import javax.security.auth.x500.X500Principal;
@ -29,6 +30,8 @@ import java.util.Optional;
public class TrustStore
{
private static final Logger LOGGER = Logger.get(TrustStore.class);
private TrustStore() {}
public static KeyStore loadTrustStore(File trustStorePath, Optional<String> trustStorePassword)
@ -48,6 +51,7 @@ public class TrustStore
}
}
catch (IOException | GeneralSecurityException ignored) {
LOGGER.error("loadTrustStore error : %s", ignored.getMessage());
}
try (InputStream in = new FileInputStream(trustStorePath)) {

View File

@ -34,9 +34,9 @@ public class TestTempFolder
root = folder.getRoot();
assertTrue(root.exists());
File newFile = folder.newFile("aNewFile");
assertEquals(newFile.getAbsolutePath(), folder.getRoot().getAbsolutePath() + "/aNewFile");
assertEquals(newFile.getCanonicalPath(), folder.getRoot().getCanonicalPath() + "/aNewFile");
File newFolder = folder.newFile("aNewFolder");
assertEquals(newFolder.getAbsolutePath(), folder.getRoot().getAbsolutePath() + "/aNewFolder");
assertEquals(newFolder.getCanonicalPath(), folder.getRoot().getCanonicalPath() + "/aNewFolder");
}
assertFalse(root.exists());
}

View File

@ -5,7 +5,7 @@
<parent>
<groupId>io.hetu.core</groupId>
<artifactId>presto-root</artifactId>
<version>1.4.0</version>
<version>1.7.0-SNAPSHOT</version>
</parent>
<artifactId>hetu-cube</artifactId>

View File

@ -36,8 +36,7 @@ public class CubeFilter
public CubeFilter(String sourceTablePredicate)
{
this.sourceTablePredicate = sourceTablePredicate;
this.cubePredicate = null;
this(sourceTablePredicate, null);
}
public String getSourceTablePredicate()

View File

@ -140,13 +140,13 @@ public class CubeStatement
return this;
}
public Builder groupBy(String column)
public Builder groupByAddString(String column)
{
this.groupBy.add(column);
return this;
}
public Builder groupBy(String... columns)
public Builder groupByAddStringList(String... columns)
{
this.groupBy.addAll(Arrays.asList(columns));
return this;

View File

@ -34,8 +34,8 @@ public class TestCubeStatement
.select("name", "address", "nationkey")
.aggregate(AggregationSignature.count())
.from("tpch.tiny.customer")
.groupBy("address")
.groupBy("name", "nationkey")
.groupByAddString("address")
.groupByAddStringList("name", "nationkey")
.build();
assertEquals(statement.getFrom(), "tpch.tiny.customer", "incorrect from table");

View File

@ -4,7 +4,7 @@
<parent>
<groupId>io.hetu.core</groupId>
<artifactId>presto-root</artifactId>
<version>1.4.0</version>
<version>1.7.0-SNAPSHOT</version>
</parent>
<artifactId>hetu-datacenter</artifactId>

View File

@ -55,6 +55,7 @@ public final class DataCenterColumnHandle
}
@JsonProperty
@Override
public String getColumnName()
{
return columnName;

View File

@ -56,11 +56,11 @@ public final class DataCenterTableHandle
*/
public DataCenterTableHandle(String catalogName, String schemaName, String tableName, OptionalLong limit)
{
this.catalogName = catalogName;
this.schemaName = requireNonNull(schemaName, "schemaName is null");
this.tableName = requireNonNull(tableName, "tableName is null");
this.limit = requireNonNull(limit, "limit is null");
this.pushDownSql = "";
this(catalogName,
requireNonNull(schemaName, "schemaName is null"),
requireNonNull(tableName, "tableName is null"),
requireNonNull(limit, "limit is null"),
"");
}
/**
@ -125,6 +125,7 @@ public final class DataCenterTableHandle
return new SchemaTableName(schemaName, tableName);
}
@Override
public String getSchemaPrefixedTableName()
{
return catalogName + SPLIT_DOT + schemaName + SPLIT_DOT + tableName;

View File

@ -179,7 +179,7 @@ public class DataCenterPlanOptimizer
List<RowExpression> pushable = new ArrayList<>();
List<RowExpression> nonPushable = new ArrayList<>();
for (RowExpression conjunct : logicalRowExpressions.extractConjuncts(node.getPredicate())) {
for (RowExpression conjunct : LogicalRowExpressions.extractConjuncts(node.getPredicate())) {
try {
conjunct.accept(queryGenerator.getConverter(), new JdbcConverterContext());
pushable.add(conjunct);

View File

@ -116,6 +116,20 @@ public class DataCenterQueryGenerator
.setSchemaTableName(Optional.of(new SchemaTableName(dcTableHandle.getSchemaName(), dcTableHandle.getTableName())))
.setSelections(selections)
.setFrom(Optional.of(table.toString()));
String catalogName = dcTableHandle.getCatalogName();
if (catalogName != null) {
contextBuilder.setRemoteCatalogName(catalogName);
}
String schemaName = dcTableHandle.getSchemaName();
if (schemaName != null) {
contextBuilder.setRemoteSchemaName(schemaName);
}
String tableName = dcTableHandle.getTableName();
if (tableName != null) {
contextBuilder.setRemoteTablename(tableName);
}
// If LIMIT has been push down, add it to context
if (dcTableHandle.getLimit().isPresent()) {
contextBuilder.setLimit(dcTableHandle.getLimit());

View File

@ -1392,24 +1392,24 @@ public class TestCrossRegionDynamicFilter
hetuServer.installPlugin(new StateStoreManagerPlugin());
hetuServer.loadStateSotre();
DistributedQueryRunner queryRunner = null;
DistributedQueryRunner distributedQueryRunner = null;
try {
queryRunner = DistributedQueryRunner.builder(testSessionBuilder().build())
distributedQueryRunner = DistributedQueryRunner.builder(testSessionBuilder().build())
.setNodeCount(1)
.build();
Map<String, String> connectorProperties = new HashMap<>(properties);
connectorProperties.putIfAbsent("connection-url", hetuServer.getBaseUrl().toString());
connectorProperties.putIfAbsent("connection-user", "root");
queryRunner.installPlugin(new DataCenterPlugin());
queryRunner.createDCCatalog("dc", "dc", connectorProperties);
queryRunner.installPlugin(new TpchPlugin());
queryRunner.createCatalog("tpch", "tpch", properties);
distributedQueryRunner.installPlugin(new DataCenterPlugin());
distributedQueryRunner.createDCCatalog("dc", "dc", connectorProperties);
distributedQueryRunner.installPlugin(new TpchPlugin());
distributedQueryRunner.createCatalog("tpch", "tpch", properties);
return queryRunner;
return distributedQueryRunner;
}
catch (Throwable e) {
closeAllSuppress(e, queryRunner);
closeAllSuppress(e, distributedQueryRunner);
throw e;
}
}

View File

@ -212,31 +212,31 @@ public class TestDataCenterClient
@Test(expectedExceptions = RuntimeException.class)
public void testPasswordWithoutSSL()
{
DataCenterConfig config = new DataCenterConfig().setConnectionUrl(this.baseUri)
DataCenterConfig dataCenterConfig = new DataCenterConfig().setConnectionUrl(this.baseUri)
.setConnectionUser("root")
.setConnectionPassword("root")
.setSsl(false);
DataCenterStatementClientFactory.newHttpClient(config);
DataCenterStatementClientFactory.newHttpClient(dataCenterConfig);
}
@Test(expectedExceptions = RuntimeException.class)
public void testKerberosWithoutSSL()
{
DataCenterConfig config = new DataCenterConfig().setConnectionUrl(this.baseUri)
DataCenterConfig dataCenterConfig = new DataCenterConfig().setConnectionUrl(this.baseUri)
.setConnectionUser("root")
.setKerberosRemoteServiceName("kerberos")
.setSsl(false);
DataCenterStatementClientFactory.newHttpClient(config);
DataCenterStatementClientFactory.newHttpClient(dataCenterConfig);
}
@Test(expectedExceptions = RuntimeException.class)
public void testAccessTokenWithoutSSL()
{
DataCenterConfig config = new DataCenterConfig().setConnectionUrl(this.baseUri)
DataCenterConfig dataCenterConfig = new DataCenterConfig().setConnectionUrl(this.baseUri)
.setConnectionUser("root")
.setAccessToken("token")
.setSsl(false);
DataCenterStatementClientFactory.newHttpClient(config);
DataCenterStatementClientFactory.newHttpClient(dataCenterConfig);
}
@Test(expectedExceptions = RuntimeException.class)

View File

@ -1,6 +1,6 @@
# Audit Log
openLooKeng audit logging functionality is a custom event listener that is invoked for query creation and query completion (success or failure)
openLooKeng audit logging functionality is a custom event listener, which monitors the start and stop of openLooKeng cluster and the dynamic addition and deletion of nodes in the cluster; Listen to WebUi user login and exit events; Listen for query events and call when the query is created and completed (success or failure).
An audit log contains the following information:
1. time when an event occurs
@ -24,15 +24,17 @@ To enable audit logging feature, the following configs must be present in `etc/e
hetu.event.listener.type=AUDIT
hetu.event.listener.listen.query.creation=true
hetu.event.listener.listen.query.completion=true
hetu.auditlog.logoutput=/var/log/
hetu.auditlog.logconversionpattern=yyyy-MM-dd.HH
```
Other audit logging properties include:
The following is a detailed description of audit logging properties:
`hetu.event.listener.audit.file`: Optional property to define absolute file path for the audit file. Ensure the process running the openLooKeng server has write access to this directory.
`hetu.event.listener.type`: property to define logging type for audit files. Allowed values are AUDIT and LOGGER.
`hetu.event.listener.audit.filecount`: Optional property to define the number of files to use
`hetu.auditlog.logoutput`: property to define absolute file directory for audit files. Ensure the process running the openLooKeng server has write access to this directory.
`hetu.event.listener.audit.limit`: Optional property to define the maximum number of bytes to write to any one file
`hetu.auditlog.logconversionpattern`: property to define the conversion pattern of audit files. Allowed values are yyyy-MM-dd.HH and yyyy-MM-dd.
Example configuration file:
@ -41,7 +43,6 @@ event-listener.name=hetu-listener
hetu.event.listener.type=AUDIT
hetu.event.listener.listen.query.creation=true
hetu.event.listener.listen.query.completion=true
hetu.event.listener.audit.file=/var/log/hetu/hetu-audit.log
hetu.event.listener.audit.filecount=1
hetu.event.listener.audit.limit=100000
hetu.auditlog.logoutput=/var/log/
hetu.auditlog.logconversionpattern=yyyy-MM-dd.HH
```

View File

@ -0,0 +1,25 @@
#Extension Physical Execution Planner
This section describes how to add an extension physical execution planner in openLooKeng. With the extension physical execution planner, openLooKeng can utilize other operator acceleration libraries to speed up the execution of SQL statements.
##Configuration
To enable extension physical execution feature, the following configs must be added in
`config.properties`
``` properties
extension_execution_planner_enabled=true
extension_execution_planner_jar_path=file:///xxPath/omni-openLooKeng-adapter-1.6.1-SNAPSHOT.jar
extension_execution_planner_class_path=nova.hetu.olk.OmniLocalExecutionPlanner
```
The above attributes are described below:
- `extension_execution_planner_enabled`: Enable extension physical execution feature.
- `extension_execution_planner_jar_path`: Set the file path of the extension physical execution jar package.
- `extension_execution_planner_class_path`: Set the package path of extension physical execution generated class in jar。
##Usage
The below command can control the enablement of extension physical execution feature in WebUI or Cli while running openLooKeng:
```
set session extension_execution_planner_enabled=true/false
```

View File

@ -118,6 +118,20 @@ This section describes the most important config properties that may be used to
>
> This is the amount of memory set aside as headroom/buffer in the JVM heap for allocations that are not tracked by openLooKeng.
### `query.suspend-query-enabled`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables running query temporary suspension when system is in low resource situation.
### `query.max-suspended-queries`
> - **Type:** `integer`
> - **Default value:** `10`
>
> Maximum number of queries to attempt suspension before starting of killing the queries. This property comes in effect only if `query.suspend-query-enabled` is configured `true`
## Spilling Properties
### `experimental.spill-enabled`
@ -161,6 +175,30 @@ This section describes the most important config properties that may be used to
>
> This config property can be overridden by the `spill_window_operator` session property.
### `experimental.spill-build-for-outer-join-enabled`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables spill feature for right-outer and full-outer join operations.
>
>
>
> This config property can be overridden by the `spill_build_for_outer_join_enabled` session property.
### `experimental.inner-join-spill-filter-enabled`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables bloom filter based build-side spill matching for probe side spill decision.
>
>
>
> This config property can be overridden by the `inner_join_spill_filter_enabled` session property.
### `experimental.spill-reuse-tablescan`
> - **Type:** `boolean`
@ -179,7 +217,7 @@ This section describes the most important config properties that may be used to
>
> Directory where spilled content will be written. It can be a comma separated list to spill simultaneously to multiple directories, which helps to utilize multiple drives installed in the system.
>
>
> When `experimental.spiller-spill-to-hdfs` is to `true`, `experimental.spiller-spill-path` must contain only a single directory.
>
> It is not recommended to spill to system drives. Most importantly, do not spill to the drive on which the JVM logs are written, as disk overutilization might cause JVM to pause for lengthy periods, causing queries to fail.
@ -240,6 +278,71 @@ This section describes the most important config properties that may be used to
>
> Enables using a randomly generated secret key (per spill file) to encrypt and decrypt data spilled to disk
### `experimental.spill-direct-serde-enabled`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables to serialize/read the page directly to/from the stream.
### `experimental.spill-prefetch-read-pages`
> - **Type:** `integer`
> - **Default value:** `1`
>
> Sets number of pages prefetched while reading from spilled files.
### `experimental.spill-use-kryo-serialization`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables Kryo based serialization for spill to disk, instead of default java serializer.
### `experimental.revocable-memory-selection-threshold`
> - **Type:** `data size`
> - **Default value:** `512 MB`
>
> Sets memory selection threshold for revocable memory of operator to directly allocate revocable memory for remaining bytes ready to revoke.
### `experimental.prioritize-larger-spilts-memory-revoke`
> - **Type:** `boolean`
> - **Default value:** `true`
>
> Enables to prioritize splits with larger revocable memory.
### `experimental.spill-non-blocking-orderby`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables order by operator to use asynchronous mechanism to spill, i.e it can accumulate input even when a spill is in progress and initiate a secondary spill when the secondary data accumulate exceeds a threshold or when the primary spill is completed, the default value of the threshold is the minimum between 20MB and 5% of available free memory. This property must be used in conjunction with the `experimental.spill-enabled` property.
>
>
>
> This config property can be overridden by the `spill_non_blocking_orderby` session property.
### `experimental.spiller-spill-to-hdfs`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables spilling into HDFS. When this property is set to `true` the property `experimental.spiller-spill-profile` must be set and also `experimental.spiller-spill-path` must contain only a single path.
### `experimental.spiller-spill-profile`
> - **Type:** `string`
> - **No default value.** Must be set when spilling to hdfs is enabled
>
>
> This property defines the [filesystem](../develop/filesystem.md) profile used to spill. The corresponding profile must exist in `etc/filesystem`. For example, if this property is set as `experimental.spiller-spill-profile=spill-hdfs`, a profile describing this filesystem `spill-hdfs.properties` must be created in `etc/filesystem` with necessary information including authentication type, config, and keytabs (if applicable, refer [filesystem](../develop/filesystem.md) for details).
>
> This property is required when `experimental.spiller-spill-to-hdfs` is set to `true`. It must be included in configuration files for all coordinators and all workers. The specified file system must be accessible by all workers, and they must be able to read from and write to the path declared in `experimental.spiller-spill-path` folder in the specified file system.
## Exchange Properties
Exchanges transfer data between openLooKeng nodes for different stages of a query. Adjusting these properties may help to resolve inter-node communication issues or improve network utilization.
@ -275,19 +378,8 @@ Exchanges transfer data between openLooKeng nodes for different stages of a quer
>
> Maximum size of a response returned from an exchange request. The response will be placed in the exchange client buffer which is shared across all concurrent requests for the exchange.
>
>
>
> Increasing the value may improve network throughput if there is high latency. Decreasing the value may improve query performance for large clusters as it reduces skew due to the exchange client buffer holding responses for more tasks (rather than hold more data from fewer tasks).
### `exchange.max-error-duration`
> - **Type:** `duration`
> - **Minimum value:** `1m`
> - **Default value:** `7m`
>
> The maximum amount of time coordinator waits for inter-task related errors to be resolved before it's considered a failure.
### `sink.max-buffer-size`
> - **Type:** `data size`
@ -295,6 +387,113 @@ Exchanges transfer data between openLooKeng nodes for different stages of a quer
>
> Output buffer size for task data that is waiting to be pulled by upstream tasks. If the task output is hash partitioned, then the buffer will be shared across all of the partitioned consumers. Increasing this value may improve network throughput for data transferred between stages if the network has high latency or if there are many nodes in the cluster.
## Failure Recovery handling Properties
### Failure Retry Policies
### `failure.recovery.retry.profile`
> - **Type:** `String`
> - **Default value:** `default`
>
> This property defines the failure detection profile used to determine if failure has happened for a http client. The value `<profile-name>` set for this property has to correspond to `<profile-name>.properties` file in `etc/failure-retry-policy/`. In case no such profile is available, and this property is not set, "default" profile is used.
> For example, `failure.recovery.retry.profile="test"` requires `test.properties` file to be present in `etc/failure-retry-policy`.
> The file `test.properties` must contain `failure.recovery.retry.type` specified.
### `failure.recovery.retry.type`
> - **Type:** `String`
> - **Default value:** `timeout`
>
> The failure detection mechanism in use. Default is timeout based failure detection.
>
#### `timeout` based failure detection.
> Using this mechanism, HTTP client failures are retried for a specific duration before considering it as a permanent failure.
> Additional properties `max.error.duration` can be defined for this type of failure detection.
>
#### `max-retry` based failure detection.
> Using this mechanism, HTTP client failures are retried for a specific number of times before considering it as a permanent failure.
> Additional properties `max.retry.count` and `max.error.duration` can be defined for this type of failure detection.
> Using this type of failure detection is configured to be used, `max.retry.count` times retry is performed before consulting the failure detector module. When the remote node is failed as per the failure detector module, HTTP client considers it a permanent failure. Otherwise, i.e. When remote worker node is alive but not sending response, retry happens for `max.error.duration` before considering it as permanent failure.
### `max.error.duration`
> - **Type:** `duration`
> - **Default value:** `300s`
>
> The maximum amount of time coordinator waits for inter-task related errors to be resolved before it's considered a permanent failure.
### `max.retry.count`
> - **Type:** `integer`
> - **Default value:** `100`
>
> The maximum number of retry for failed task performed by the coordinator before consulting the failure detector module about the remote node status.
> This parameter is the minimum count before consulting the failure detection module. Hence, the actual number of failures may vary slightly based on the cluster size, and load on the cluster.
> This property is used only for `max-retry` based failure detection profiles.
> The minimum value for this parameter is 100.
### Gossip Protocol Configurations for Failure Detection
### `failure-detection-protocol`
>- **Type:** String
>- **Default value:** `heartbeat`
>
> This property defines the type of failure detector in use. Default configuration is `heartbeat` failure detector.
> Gossip protocol can be enabled by specifying this parameter in `config.properties` file, with the value `gossip`.
> All nodes (i.e. coordinator as well as workers) in a cluster should have this property specified in their respective `etc/config.properties` file.
### `failure-detector.heartbeat-interval`
>- **Type:** Duration
>- **Default value:** `500ms` (500 miliseconds)
>
> This is the interval of gossip between two nodes in the cluster.
> In gossip protocol, two workers are expected to gossip with higher frequency than the coordinator and a worker.
> In `config.properties` for the coordinator, this property can be set with a reasonably higher value, such as `5s` (5 seconds).
> In workers, this property can be left to use the default value.
>
### `failure-detector.worker-gossip-probe-interval`
>
> - **Type:** Duration
>- **Default value:** `5s` (5 seconds)
>
> Gossip protocol uses monitoring tasks (same as the heartbeat failure detector) to keep tab on the other nodes.
> This property specifies the interval of refreshing the monitoring tasks to trigger worker to worker gossip.
> This property, if needed to be configured with any other value than the default, should be specified only for the worker nodes.
> This parameter should have higher value than `failure-detector.heartbeat-interval`.
>
### `failure-detector.coordinator-gossip-probe-interval`
>
> - **Type:** Duration
>- **Default value:** `5s` (5 seconds)
>
> Gossip protocol uses monitoring tasks (same as the heartbeat failure detector) to keep tab on the other nodes.
> This property specifies the interval of refreshing the monitoring tasks to trigger coordinator to worker gossip.
> This property, if needed to be configured with any other value than the default, should be specified only for the coordinator.
> This parameter should have higher value than `failure-detector.heartbeat-interval` and `failure-detector.worker-gossip-probe-interval`.
>
### `failure-detector.coordinator-gossip-collate-interval`
>
> - **Type:** Duration
>- **Default value:** `2s` (2 seconds)
>
> This property specifies the interval in which the coordinator collates all the gossips it obtained from all the workers.
> This property has to be specified only for the coordinator.
> This parameter should have higher value than `failure-detector.heartbeat-interval`.
>
### `failure-detector.gossip-group-size`
>
> - **Type:** Integer
>- **Default value:** `Integer.MAX_VALUE`
>
> A worker should gossip with how many other workers in the cluster, is defined by this parameter.
> Any value higher than the cluster-size (i.e. the number of workers) implies all-to-all gossip.
> To keep the network overhead low, this value should be reasonably low for a big cluster (e.g. 10 for a cluster size of 100).
> On each refresh of the worker-monitoring tasks at the coordinator, the coordinator defines the list of worker URIs of size `failure-detector.gossip-group-size` to trigger worker-to-worker gossip.
## Task Properties
### `task.concurrency`
@ -679,7 +878,7 @@ helps with cache affinity scheduling.
> Auto-Vacuum enables the system to automatically manage vacuum jobs by constantly monitoring the tables which needs vacuum in order to maintain optimal performance.
> Engine gets the tables from data sources that are eligible for vacuum and trigger vacuum operation for those tables.
### `auto-vacuum.enabled:`
### `auto-vacuum.enabled`
> - **Type:** `boolean`
> - **Default value:** `false`
@ -749,15 +948,25 @@ helps with cache affinity scheduling.
> - **Default value:** `5m`
>
> The maximum time coordinator waits for remote-task related error to be resolved before it's considered a failure.
>
> Note:
> For snapshot recovery `query.remote-task.max-error-duration` should be greater than `exchange.max-error-duration`.
## Distributed Snapshot
## Query Recovery
### `recovery_enabled`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> This session property is used to enable or disable the recovery framework, which enables to restart/resume the query in case of failure.
### `snapshot_enabled`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> This session property is used to enable or disable the distributed snapshot functionality.
> This session property is enabled to capture snapshots during query execution, when recovery framework is enabled. Without recovery framework enabled this flag has no significance
### `hetu.experimental.snapshot.profile`
@ -769,20 +978,67 @@ helps with cache affinity scheduling.
>
> This is an experimental property. In the future it may be allowed to store snapshots in non-file-system locations, e.g. in a connector.
### `hetu.snapshot.maxRetries`
### `hetu.recovery.maxRetries`
> - **Type:** `int`
> - **Default value:** `10`
>
> This property defines the maximum number of error recovery attempts for a query. When the limit is reached, the query fails.
>
> This can also be specified on a per-query basis using the `snapshot_max_retries` session property.
> This can also be specified on a per-query basis using the `recovery_max_retries` session property.
### `hetu.snapshot.retryTimeout`
### `hetu.recovery.retryTimeout`
> - **Type:** `duration`
> - **Default value:** `10m` (10 minutes)
>
> This property defines the maximum amount of time for the system to wait until all tasks are successfully restored. If any task is not ready within this timeout, then the recovery attempt is considered a failure, and the query will try to resume from an earlier snapshot if available.
>
> This can also be specified on a per-query basis using the `snapshot_retry_timeout` session property.
> This can also be specified on a per-query basis using the `recovery_retry_timeout` session property.
### `hetu.snapshot.useKryoSerialization`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables Kryo based serialization for snapshot, instead of default java serializer.
### `experimental.eliminate-duplicate-spill-files`
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Enables elimination of duplicate spill files storage as part of snapshot capture.
## HTTP Client Configurations
### `http.client.idle-timeout`
> - **Type:** `duration`
> - **Default value:** `30s` (30 seconds)
>
> This property defines the time for which a given http client shall stay connected without any operations performed over it.
> After the specified time elapse with no activity, then the client connection is closed and related resources are released.
>
> (Note: this parameter should be configured with higher time when in high load environment)
### `http.client.request-timeout`
> - **Type:** `duration`
> - **Default value:** `10s` (10 seconds)
>
> This property defines the time threshold for a given http client for which response should be received.
> After the configured time elapsed and no response received, then client connection consider that to be failure in submission of request.
>
> (Note: this parameter should be configured with higher time when in high load environment)
## Connector Properties configuration
### `case-insensitive-name-matching`
>
> - **Type:** `boolean`
> - **Default value:** `false`
>
> Case-insensitive matching between database and collection names. The default is case sensitive.

View File

@ -11,9 +11,9 @@ To achieve better performance while maintaining execution reliability, the *dist
As of release 1.2.0, openLooKeng supports recovery of tasks and worker node failures.
## Enable Distributed Snapshot
## Enable Recovery framework
Distributed snapshot is most useful for long running queries. It is disabled by default, and must be enabled and disabled via a session property [`snapshot_enabled`](properties.md#snapshot_enabled). It is recommended that the feature is only enabled for complex queries that require high reliability.
Recovery framework is most useful for long running queries. It is disabled by default, and must be enabled and disabled via a session property [`recovery_enabled`](properties.md#recovery_enabled). It is recommended that the feature is only enabled for complex queries that require high reliability.
## Requirements
@ -37,7 +37,7 @@ When a query that does not meet the above requirements is submitted with distrib
## Detection
Error recovery is triggered when communication between the coordinator and a remote task fails for an extended period of time, as controlled by the [`query.remote-task.max-error-duration`](properties.md#queryremote-taskmax-error-duration) configuration.
Error recovery is triggered when communication between the coordinator and a remote task fails for an extended period of time, as controlled by the [`Failure Recovery handling Properties`](properties.md#Failure Recovery handling Properties) configuration.
## Storage Considerations
@ -55,8 +55,20 @@ Each query execution may produce multiple snapshots. Contents of these snapshots
The ability to recover from an error and resume from a snapshot does not come for free. Capturing a snapshot, depending on complexity, takes time. Thus it is a trade-off between performance and reliability.
It is suggested to only turn on distributed snapshot when necessary, i.e. for queries that run for a long time. For these types of workloads, the overhead of taking snapshots becomes negligible.
It is suggested to turn on snapshot capture when necessary, i.e. for queries that run for a long time. For these types of workloads, the overhead of taking snapshots becomes negligible.
## Snapshot statistics
Snapshot capture and restore statistics are displayed in CLI along with query result when CLI is launched in debug mode
Snapshot capture statistics includes number of snapshots captured, size of snapshots captured, CPU Time taken for capturing the snapshots and Wall Time taken for capturing the snapshots during the query. These statistics are displayed for all snapshots and for last snapshot separately.
Snapshot restore statistics covers number of times restored from snapshots during query, Size of the snapshots loaded for restoring, CPU Time taken for restoring from snapshots and Wall Time taken for restoring from snapshots. Restore statistics are displayed only when there is restore(recovery) happened during the query.
Additionally, while query is in progress number of capturing snapshots and id of the restoring snapshot will be displayed. Refer below picture for more details
![](../images/snapshot_statistics.png)
## Configurations
Configurations related to distributed snapshot feature can be found in [Properties Reference](properties.md#distributed-snapshot).
Configurations related to recovery framework feature can be found in [Properties Reference](properties.md#Query Recovery).

View File

@ -34,6 +34,10 @@ saturation of the configured spill paths.
openLooKeng treats spill paths as independent disks (see [JBOD](https://en.wikipedia.org/wiki/Non-RAID_drive_architectures#JBOD)), so there is no need to use RAID for spill.
## Spill To HDFS
Spilling directly into HDFS is also possible for that `experimental.spiller-spill-to-hdfs` needs to be set to `true`, `experimental.spiller-spill-profile` needs to be set and `spiller-spill-path` must contain only a single directory when we intend to spill into HDFS. (refer `experimental.spiller-spill-to-hdfs` and `experimental.spiller-spill-profile` properties for more details )
## Spill Compression
@ -61,14 +65,17 @@ When the build table is partitioned, the spill-to-disk mechanism can decrease th
With this mechanism, the peak memory used by the join operator can be decreased to the size of the largest build table partition. Assuming no data skew, this will be `1 / task.concurrency` times the size of the whole build table.
Note: spill-to-disk is not supported for Cross Join.
### Aggregations
Aggregation functions perform an operation on a group of values and return one value. If the number of groups you\'re aggregating over is large, a significant amount of memory may be needed. When spill-to-disk
is enabled, if there is not enough memory, intermediate cumulated aggregation results are written to disk. They are loaded back and merged with a lower memory footprint.
is enabled, if there is not enough memory, intermediate accumulated aggregation results are written to disk. They are loaded back and merged with a lower memory footprint.
### Order By
If you're trying to sort a larger amount of data, a significant amount of memory may be needed. When spill to disk for order by is enabled, if there is not enough memory, intermediate sorted results are written to disk. They are loaded back and merged with a lower memory footprint.
Generally when a spill is in progress the operator is blocked from taking inputs, but when `experimental.spill-non-blocking-orderby` is set to `true` order by uses asynchronous mechanism to spill (see`experimental.spill-non-blocking-orderby`).
### Window functions

View File

@ -33,4 +33,54 @@ and statistics about the query is available by clicking the *JSON* link. These v
> - **Allowed values:** `true`, `false`
> - **Default value:** `false`
>
> Insecure authentication over HTTP is disabled by default. This could be overridden via "hetu.queryeditor-ui.allow-insecure-over-http" property of "etc/config.properties" (e.g. hetu.queryeditor-ui.allow-insecure-over-http=true).
> Insecure authentication over HTTP is disabled by default. This could be overridden via `hetu.queryeditor-ui.allow-insecure-over-http` property of `etc/config.properties` (e.g. hetu.queryeditor-ui.allow-insecure-over-http=true).
### `hetu.queryeditor-ui.execution-timeout`
> - **Type:** `duration`
> - **Default value:** `100 DAYS`
>
> UI Execution timeout is set to 100 days as default. This could be overridden via `hetu.queryeditor-ui.execution-timeout` of `etc/config.properties`
### `hetu.queryeditor-ui.max-result-count`
> - **Type:** `int`
> - **Default value:** `1000`
>
> UI max result count is set to 1000 as default. This could be overridden via `hetu.queryeditor-ui.max-result-count` of `etc/config.properties`
### `hetu.queryeditor-ui.max-result-size-mb`
>- **Type:** `size`
>- **Default value:** `1GB`
>
> UI max result size is set to 1 GB as default. This could be overridden via `hetu.queryeditor-ui.max-result-size-mb` of `etc/config.properties`
### `hetu.queryeditor-ui.session-timeout`
> - **Type:** `duration`
> - **Default value:** `1 DAYS`
>
> UI session timeout is set to 1 day as default. This could be overridden via `hetu.queryeditor-ui.session-timeout` of `etc/config.properties`
### `hetu.queryhistory.max-count`
> - **Type:** `int`
> - **Default value:** `1000`
>
> The maximum number of query history stored by openLooKeng. This could be overridden via "hetu.queryhistory.max-count" of "etc/config.properties".
### `hetu.collectionsql.max-count`
> - **Type:** `int`
> - **Default value:** `100`
>
> The Maximum number of SQL collected by each user. This could be overridden via "hetu.collectionsql.max-count" of "etc/config.properties".
## Remarks
The max length of the favorite SQL is 600 by default. You can modify it through the following steps:
1. Login MySQL database according to the JDBC configuration of `hetu-metastore.properties`
2. Select table hetu_favorite, execute script `alter table hetu_favorite modify query varchar(2000) not null;` to modify the max length of the favorite SQL.

View File

@ -1,5 +1,8 @@
# Hudi Connector
### Release Notes
Currently Hudi only supports version 0.7.0.
### Hudi Introduction
Apache Hudi is a fast growing data lake storage system that helps organizations build and manage petabyte-scale data lakes. Hudi enables storing vast amounts of data on top of existing DFS compatible storage while also enabling stream processing in addition to typical batch-processing. This is made possible by providing two new primitives. Specifically,

View File

@ -192,4 +192,4 @@ The array fields of this structure can be defined by using the following command
## Limitations
1. openLooKeng does not support to query table in Elasticsearch which has duplicated columns, such as column "name" and "NAME";
2. opneLooKeng does not support to query the table in Elasticsearch that the name has special characters, such as '-', '.', etc.
2. openLooKeng does not support to query the table in Elasticsearch that the name has special characters, such as '-', '.', etc.

View File

@ -620,6 +620,36 @@ The following operations are not supported when `avro_schema_url` is set:
columns are not supported in `CREATE TABLE`.
- `ALTER TABLE` commands modifying columns are not supported.
## Drop Column Behavior
Syntax supported for Drop Column is as follows:
```sql
ALTER TABLE 'name' DROP COLUMN 'column_name'
```
In case of Hive connector, DROP COLUMN drops column which is at the end of existing columns. Hive doesn't support DROP COLUMN syntax, however closest DDL supported by Hive which enumerates the above is REPLACE COLUMNS. REPLACE COLUMNS removes all existing columns and adds the new set of columns whereas DROP COLUMN does remove all existing columns and add the same set of columns excluding column specified in the query, manifesting as it dropped the column.
For example, consider a table with columns **a**, **b** and **c**.
```sql
lk:default> SELECT * FROM hive_table;
a | b | c
---+----+-----
1 | 10 | 100
(1 row)
```
On dropping column **a**:
```sql
lk:default> SELECT * FROM hive_table;
b | c
---+----
1 | 10
(1 row)
```
## Procedures
@ -716,7 +746,7 @@ Drop a schema:
DROP SCHEMA hive.web
```
## Metastore Cache:
## Metastore Cache
Hive connector maintains a metastore cache to service the metastore request faster to various operations. Loading, reloading and retention times of the cache entries can be configured in `hive.properties`.
@ -740,7 +770,7 @@ REFRESH META CACHE
Additionally, metadata cache refresh command can be used to reload the metastore cache by user.
## Performance tuning notes:
## Performance tuning notes
#### INSERT

View File

@ -565,4 +565,4 @@ lk:default> SELECT created_at, raw_date FROM (
(5 rows)
```
The Kafka connector contains converters for ISO 8601, RFC 2822 text formats and for number-based timestamps using seconds or miilliseconds since the epoch. There is also a generic, text-based formatter which uses Joda-Time format strings to parse text columns.
The Kafka connector contains converters for ISO 8601, RFC 2822 text formats and for number-based timestamps using seconds or milliseconds since the epoch. There is also a generic, text-based formatter which uses Joda-Time format strings to parse text columns.

View File

@ -37,7 +37,7 @@ memory.spill-path=/opt/hetu/data/spill
hetu.metastore.cache.type=local
```
##### Multi-Node Setup
- This section will give an example configuration for Memory Connector an a cluster with more than one node.
- This section will give an example configuration for Memory Connector and a cluster with more than one node.
- Create a file `etc/catalog/memory.properties` with the following information:
``` properties
connector.name=memory
@ -80,10 +80,10 @@ memory.spill-path=/opt/hetu/data/spill
**Note:**
- `spill-path` should be set to a directory with enough free space to hold
the table data.
- See **Configuration Properties** section for additional properties and
- See [**Configuration Properties**](#configuration-properties) section for additional properties and
details.
- In `etc/config.properties` ensure that `task.writer-count` is set to
`>=` number of nodes in the cluster running openLooKeng. This will help
- In `etc/config.properties` ensure that `task.writer-count` is set
`>=` to number of nodes in the cluster running openLooKeng. This will help
distribute the data uniformly between all the workers.
Examples
@ -112,17 +112,45 @@ Create a table using the Memory Connector with sorting, indices and spill compre
CREATE TABLE memory.default.nation
WITH (
sorted_by=array['nationkey'],
index_columns=array['name', 'regionkey'],
partitioned_by=array['regionkey'],
index_columns=array['name'],
spill_compression=true
)
AS SELECT * from tpch.tiny.nation;
After table creation completes, the Memory Connector will start building indices and sorting data in the background. Once the processing is complete any queries using the sort or index columns will be faster and more efficient.
For now, `sorted_by` only accepts a single column.
For now, `sorted_by` and `partitioned_by` only accepts a single column.
Memory and Disk Usage via JMX
-----------------------------
JMX can be used to show memory and disk usage of memory connector tables
Please refer to [JMX Connector](./jmx.md) for setup
Configuration Properties
The `io.prestosql.plugin.memory.data:name=MemoryTableManager` table of `jmx.current` contains information on all the tables' memory and disk usage size in bytes
SELECT * FROM jmx.current."io.prestosql.plugin.memory.data:name=MemoryTableManager";
```
currentbytes | alltablesdiskbyteusage | alltablesmemorybyteusage | node | object_name
--------------+------------------------+--------------------------+----------+---------------------------------------------------------
23 | 3456 | 23 | example1 | io.prestosql.plugin.memory.data:name=MemoryTableManager
253 | 8713 | 667 | example2 | io.prestosql.plugin.memory.data:name=MemoryTableManager
```
Not all tables will be in memory since some may have been spilled to disk. `currentbytes` column will show the current memory occupied by tables which are in memory now.
The usage for each node is shown as a separate row, aggregation functions can be utilized to show total usage across the cluster. For example, to view total disk or memory usage on all nodes, run:
SELECT sum(alltablesdiskbyteusage) as totaldiskbyteusage, sum(alltablesmemorybyteusage) as totalmemorybyteusage FROM jmx.current."io.prestosql.plugin.memory.data:name=MemoryTableManager";
```
totaldiskbyteusage | totalmemorybyteusage
-------------------+---------------------
12169 | 690
```
# Configuration Properties
------------------------
| Property Name | Default Value | Required| Description |
@ -133,31 +161,50 @@ Configuration Properties
| `memory.max-page-size ` | 512KB | No | Memory limit for each page. Default value is recommended.|
| `memory.logical-part-processing-delay` | 5s | No | The delay between when the table is created/updated and LogicalPart processing starts. Default value is recommended.|
| `memory.thread-pool-size ` | Half of threads available to the JVM | No | Maximum threads to allocate for background processing (e.g. sorting, index creation, cleanup, etc)|
| `memory.table-statistics-enabled` | False | No | When enabled, user can run analyze to collect statistics and leverage that information for accelerating queries.|
Path whitelist`["/tmp", "/opt/hetu", "/opt/openlookeng", "/etc/hetu", "/etc/openlookeng", current workspace]`
Path whitelist: `["/tmp", "/opt/hetu", "/opt/openlookeng", "/etc/hetu", "/etc/openlookeng", current workspace]`
Additional WITH properties
--------------
--------------------------
Use these properties when creating a table with the Memory Connector to make queries faster.
| Property Name | Argument type | Requirements | Description|
|--------------------------|---------------------------|----------------------------------|------------|
| sorted_by | `array['col']` | Maximum of one column. Column type must be comparable. | Sort and create indexes on the given column|
| partitioned_by | `array['col']` | Maximum of one column. | Partition the table on the given column|
| index_columns | `array['col1', 'col2']` | None | Create indexes on the given column|
| spill_compression | `boolean` | None | Compress data when spilling to disk|
Index Types
--------------
These are the types of indices that are built on the columns you specify in `sorted_by` or `index_columns`. If a query operator is not supported by a particular index you can still
use that operator, but the query will not benefit from the index.
These are the types of indices that are built on the columns you specify in `sorted_by` or `index_columns`.
If a query operator is not supported by a particular index, you can still use that operator, but the query will not benefit from the index.
| Index ID |Built for Columns In | Supported query operators |
|-----------------------------------|----------------------------------------|---------------------------------------|
| Bloom | `sorted_by,index_columns` | `=` `IN` |
| Bloom | `index_columns` | `=` `IN` |
| MinMax | `sorted_by,index_columns` | `=` `>` `>=` `<` `<=` `IN` `BETWEEN` |
| Sparse | `sorted_by` | `=` `>` `>=` `<` `<=` `IN` `BETWEEN` |
Using statistics
-----------------
If the statistic configuration is enabled, you can refer to the example below to use it.
Create a table using the Memory Connector:
CREATE TABLE memory.default.nation AS
SELECT * from tpch.tiny.nation;
Run Analyze to collect the statistic information:
ANALYZE memory.default.nation;
And then run the queries. Note that currently we do not support automatic statistic update, so you will need to run ANALYZE again if the table is updated.
Developer Information
----------------------------
@ -166,20 +213,11 @@ This section outlines the overall design of the Memory Connector, as shown in th
![Memory Connector Overall Design](../images/memory-connector-design.png)
### Scheduling Process
The data to be processed are stored in pages, which are distributed to different worker nodes in openLooKeng.
In the Memory Connector, each worker has a number of LogicalParts.
During table creation, LogicalParts in the workers are filled with the input pages in a round-robin fashion.
Table data will be automatically spilled to disk as part of a background process as well.
If there is not enough memory to hold the entire data, the tables can be released from memory according to LRU rule.
HetuMetastore is used to persist table metadata.
At query time, when Tablescan operation is scheduled, the LogicalParts will be scheduled.
The data to be processed are stored in pages, which are distributed to different worker nodes in openLooKeng. In the Memory Connector, each worker has several LogicalParts. During table creation, LogicalParts in the workers are filled with the input pages in a round-robin fashion. Table data will be automatically spilled to disk as part of a background process as well. If there is not enough memory to hold the entire data, the tables can be released from memory according to LRU rule. HetuMetastore is used to persist table metadata. At query time, when Tablescan operation is scheduled, the LogicalParts will be scheduled.
### LogicalPart
As shown in the lower part of the design figure, LogicalPart is the data structure that contains both indexes and data.
The sorting and indexing are handled in a background process allowing faster querying,
but the table is still queriable during processing.
LogicalParts have a maximum configurable size (default 256 MB).
New LogicalParts are created once the previous one is full.
As shown in the lower part of the design figure, LogicalPart is the data structure that contains both indexes and data. The sorting and indexing are handled in a background process allowing faster querying,
but the table is still queriable during processing. LogicalParts have a maximum configurable size (default 256 MB). New LogicalParts are created once the previous one is full.
### Indices
@ -213,4 +251,5 @@ Limitations and known Issues
- Without State Store and Hetu Metastore with global cache, after `DROP TABLE`, memory is not released immediately on the workers. It is released on the next `CREATE TABLE` operation.
- Currently only a single column in ascending order is supported by `sorted_by`
- If a CTAS (CREATE TABLE AS) query fails or is cancelled, an invalid table will remain. This table must be dropped manually.
- If a CTAS (CREATE TABLE AS) query fails or is cancelled, an invalid table will remain. This table must be dropped manually.
- And we support BOOLEAN, All INT Types, CHAR, VARCHAR, DOUBLE, REAL, DECIMAL, DATE, TIME, UUID types as partition keys.

View File

@ -0,0 +1,103 @@
OmniData Connector
==============
## Overview
The OmniData connector allows querying data stored in the remote Hive data warehouse.
It pushes the operators of openLooKeng down to the storage node to achieve near-data calculation, thereby reducing the amount of network transmission data and improving computing performance.
For more information, please see: [OmniData](https://www.hikunpeng.com/en/developer/boostkit/big-data?accelerated=3) and [OmniData connector](https://github.com/kunpengcompute/omnidata-openlookeng-connector).
## Supported File Types
The following file types are supported for the OmniData connector:
- ORC
- Parquet
- Text
## Configuration
Create `etc/catalog/omnidata.properties` with the following configurations, replacing `example.net:9083` with the correct host and port for your Hive metastore Thrift service:
``` properties
connector.name=omnidata-openlookeng
hive.metastore.uri=thrift://example.net:9083
```
### HDFS Configuration
For basic setups, openLooKeng configures the HDFS client automatically and does not require any configuration files. In some cases, such as when using federated HDFS or NameNode high availability, it is necessary to specify additional HDFS client options in order to access your HDFS cluster. To do so, add the `hive.config.resources` property to reference your HDFS config files:
``` properties
hive.config.resources=/etc/hadoop/conf/core-site.xml,/etc/hadoop/conf/hdfs-site.xml
```
Only specify additional configuration files if necessary for your setup. We also recommend reducing the configuration files to have the minimum set of required properties, as additional properties may cause problems.
The configuration files must exist on all openLooKeng nodes. If you are referencing existing Hadoop config files, make sure to copy them to any openLooKeng nodes that are not running Hadoop.
## OmniData Configuration Properties
| Property Name | Description | Default |
| ------------------------------- | ------------------------------------------------------------ | ------- |
| hive.metastore | The type of Hive metastore | thrift |
| hive.config.resources | An optional comma-separated list of HDFS configuration files. These files must exist on the machines running openLooKeng. Only specify this if absolutely necessary to access HDFS. Example: `/etc/hdfs-site.xml` | |
| hive.omnidata-enabled | Allows push-down operators to execute on the storage side. If disabled, all operators will not be pushed down. | true |
| hive.min-offload-row-number | If the number of rows in the table is less than the threshold, all operators of the table will not be pushed down. | 500 |
| hive.filter-offload-enabled | Allows the filter operator to be pushed down to the storage side. If disabled, the filter operator will not be pushed down. | true |
| hive.filter-offload-factor | Only when the selection rate of the filter operator is less than the threshold, it will be pushed down. | 0.25 |
| hive.aggregator-offload-enabled | Allows the aggregator operator to be pushed down to the storage side. If disabled, the aggregator operator will not be pushed down. | true |
| hive.aggregator-offload-factor | Only when the aggregation rate of the aggregator operator is less than the threshold, it will be pushed down. | 0.25 |
For more configuration, please refer to the [Hive Configuration Properties](./hive.md#Hive Configuration Properties) chapter.
### Querying OmniData
The SQL query plan after some operators are pushed down:
```sql
lk:tpch_flat_orc_date_1000> explain select sum(l_extendedprice * l_discount) as revenue
-> from
-> lineitem
-> where
-> l_shipdate >= DATE '1993-01-01'
-> and l_shipdate < DATE '1994-01-01'
-> and l_discount between 0.06 - 0.01 and 0.06 + 0.01
-> and l_quantity < 25;
Query Plan
------------------------------------------------------------------------------------------------------
Output[revenue]
│ Layout: [sum:double]
│ Estimates: {rows: 4859991664 (40.74GB), cpu: 246.43G, memory: 86.00GB, network: 45.26GB}
│ revenue := sum
└─ Aggregate(FINAL)
│ Layout: [sum:double]
│ Estimates: {rows: 4859991664 (40.74GB), cpu: 246.43G, memory: 86.00GB, network: 45.26GB}
│ sum := sum(sum_4)
└─ LocalExchange[SINGLE] ()
│ Layout: [sum_4:double]
│ Estimates: {rows: 5399990738 (45.26GB), cpu: 201.17G, memory: 45.26GB, network: 45.26GB}
└─ RemoteExchange[GATHER]
│ Layout: [sum_4:double]
│ Estimates: {rows: 5399990738 (45.26GB), cpu: 201.17G, memory: 45.26GB, network: 45.26GB}
└─ Aggregate(PARTIAL)
│ Layout: [sum_4:double]
│ Estimates: {rows: 5399990738 (45.26GB), cpu: 201.17G, memory: 45.26GB, network: 0B}
│ sum_4 := sum(expr)
└─ ScanProject[table = hive:tpch_flat_orc_date_1000:lineitem offload={ filter=[AND(AND(BETWEEN(l_discount, 0.05, 0.07), LESS_THAN(l_quantity, 25.0)), AND(GREATER_THAN_OR_EQUAL(l_shipdate, 8401), LESS_THAN(l_shipdate, 8766)))]} ]
Layout: [expr:double]
Estimates: {rows: 5999989709 (50.29GB), cpu: 100.58G, memory: 0B, network: 0B}/{rows: 5999989709 (50.29GB), cpu: 150.87G, memory: 0B, network: 0B}
expr := (l_extendedprice) * (l_discount)
l_extendedprice := l_extendedprice:double:5:REGULAR
l_discount := l_discount:double:6:REGULAR
```
## OmniData Connector Limitations
- The OmniData service needs to be deployed on the storage node.
- Only the pushdown of Filter, Aggregator, and Limit operators are supported.

View File

@ -45,6 +45,95 @@ Finally, you can access the `hetutb` table in the `public` schema:
If you used a different name for your catalog properties file, use that catalog name instead of `opengauss` in the above examples.
## openGauss Update/Delete Support
### Create openGauss Table
Example
```sql
CREATE TABLE opengauss_table (
id int,
name varchar(255));
```
### INSERT on openGauss tables
Example
```sql
INSERT INTO opengauss_table
VALUES
(1, 'Jack'),
(2, 'Bob');
```
### UPDATE on openGauss tables
Example
```sql
UPDATE opengauss_table
SET name='Tim'
WHERE id=1;
```
Above example updates the column `name`'s value to `Tim` of rows with column `id` having value `1`.
SELECT result before UPDATE:
```sql
lk:default> SELECT * FROM opengauss_table;
id | name
----+------
1 | Jack
2 | Bob
(2 rows)
```
SELECT result after UPDATE
```sql
lk:default> SELECT * FROM opengauss_table;
id | name
----+------
2 | Bob
1 | Tim
(2 rows)
```
### DELETE on openGauss tables
Example
```sql
DELETE FROM opengauss_table
WHERE id=2;
```
Above example delete the rows with column `id` having value `2`.
SELECT result before DELETE:
```sql
lk:default> SELECT * FROM opengauss_table;
id | name
----+------
2 | Bob
1 | Tim
(2 rows)
```
SELECT result after DELETE:
```sql
lk:default> SELECT * FROM opengauss_table;
id | name
----+------
1 | Tim
(1 row)
```
****Note:****
> - When the compatibility type of the openGuass database is O (DBCOMPATIBILITY = A), the `Date` data type is not supported.
@ -64,4 +153,4 @@ openGauss Connector Limitations
The following SQL statements are not yet supported:
[DELETE](../sql/delete.md), [GRANT](../sql/grant.md), [REVOKE](../sql/revoke.md), [SHOW GRANTS](../sql/show-grants.md), [SHOW ROLES](../sql/show-roles.md), [SHOW ROLE GRANTS](../sql/show-role-grants.md)
[GRANT](../sql/grant.md), [REVOKE](../sql/revoke.md), [SHOW GRANTS](../sql/show-grants.md), [SHOW ROLES](../sql/show-roles.md), [SHOW ROLE GRANTS](../sql/show-role-grants.md)

View File

@ -46,9 +46,98 @@ Finally, you can access the `clicks` table in the `web` schema:
If you used a different name for your catalog properties file, use that catalog name instead of `postgresql` in the above examples.
## PostgreSQL Update/Delete Support
### Create PostgreSQL Table
Example
```sql
CREATE TABLE postgresql_table (
id int,
name varchar(255));
```
### INSERT on PostgreSQL tables
Example
```sql
INSERT INTO postgresql_table
VALUES
(1, 'Jack'),
(2, 'Bob');
```
### UPDATE on PostgreSQL tables
Example
```sql
UPDATE postgresql_table
SET name='Tim'
WHERE id=1;
```
Above example updates the column `name`'s value to `Tim` of rows with column `id` having value `1`.
SELECT result before UPDATE:
```sql
lk:default> SELECT * FROM postgresql_table;
id | name
----+------
1 | Jack
2 | Bob
(2 rows)
```
SELECT result after UPDATE
```sql
lk:default> SELECT * FROM postgresql_table;
id | name
----+------
2 | Bob
1 | Tim
(2 rows)
```
### DELETE on PostgreSQL tables
Example
```sql
DELETE FROM postgresql_table
WHERE id=2;
```
Above example delete the rows with column `id` having value `2`.
SELECT result before DELETE:
```sql
lk:default> SELECT * FROM postgresql_table;
id | name
----+------
2 | Bob
1 | Tim
(2 rows)
```
SELECT result after DELETE:
```sql
lk:default> SELECT * FROM postgresql_table;
id | name
----+------
1 | Tim
(1 row)
```
PostgreSQL Connector Limitations
--------------------------------
The following SQL statements are not yet supported:
[DELETE](../sql/delete.md), [GRANT](../sql/grant.md), [REVOKE](../sql/revoke.md), [SHOW GRANTS](../sql/show-grants.md), [SHOW ROLES](../sql/show-roles.md), [SHOW ROLE GRANTS](../sql/show-role-grants.md)
[GRANT](../sql/grant.md), [REVOKE](../sql/revoke.md), [SHOW GRANTS](../sql/show-grants.md), [SHOW ROLES](../sql/show-roles.md), [SHOW ROLE GRANTS](../sql/show-role-grants.md)

View File

@ -0,0 +1,173 @@
Redis Connector
====================
Overview
--------
this connector allows the use of Redis key/value pair is presented as a single row in openLooKeng.
**Note**
*In Redis,key/value pair can only be mapped to string or hash value types.keys can be stored in a zset,then keys can split into multiple slice*
*Support Redis 2.8.0 or higher*
Configuration
-------------
To configure the Redis connector, create a catalog properties file `etc/catalog/redis.properties` with the following contents, replacing the properties as appropriate:
``` properties
connector.name=redis
redis.table-names=schema1.table1,schema1.table2
redis.nodes=host1:port
```
### Multiple Redis Servers
You can have as many catalogs as you need. If you have additional
Redis servers, simply add another properties file to ``etc/catalog``
with a different name, making sure it ends in ``.properties``.
For example, if you name the property file `sales.properties`, openLooKeng will create a catalog named `sales` using the configured connector.
Configuration properties
------------------------
The following configuration properties are available:
| Property Name | Description |
|:-----------------------------------------------------------------|:----------------------------------------------------------------------------------------------------------------------|
| `redis.table-names` | List of all tables provided by the catalog |
| `redis.default-schema` | Default schema name for tables (default `default`) |
| `redis.nodes` | List of nodes in the Redis server |
| `redis.connect-timeout` | Timeout for connecting to the Redis server (ms) (default 2000) |
| `redis.scan-count` | The number of keys obtained from each scan for string and hash value types (default 100) |
| `redis.key-prefix-schema-table` | Redis keys have schema-name:table-name prefix (default false) |
| `redis.key-delimiter` | Delimiter separating schema_name and table_name if redis.key-prefix-schema-table is used (default `:`) |
| `redis.table-description-dir` | Directory containing table description files (default `etc/redis/`) |
| `redis.hide-internal-columns` | Whether internal columns are shown in table metadata or not. (default true) |
| `redis.database-index` | Redis database index (default 0) |
| `redis.password` | Redis server password (default null) |
| `redis.table-description-interval` | the interval of flush description files (ms) (default no flush,table description will be memoized without expiration) |
Internal columns
----------------
| Column name | Type | Description |
|:-------------------| :------ |:-----------------------------------------------------------------------------------------------------------------------------------------|
| `_key` | VARCHAR | Redis key. |
| `_value` | VARCHAR | Redis value corresponding to the key |
| `_key_length` | BIGINT | Number of bytes in the key. |
| `_key_corrupt` | BOOLEAN | True if the decoder could not decode the key for this row. When true, data columns mapped from the key should be treated as invalid. |
| `_value_corrupt` | BOOLEAN | True if the decoder could not decode the value for this row. When true, data columns mapped from the value should be treated as invalid. |
Table Definition Files
----------------------
For openLooKeng, every key/value pair must be mapped into columns to allow queries against the data. It is like kafka conntector,so you can refer to kafka-tutorial
A table definition file consists of a JSON definition for a table. The name of the file can be arbitrary but must end in `.json`.
for example,there is a nation.json
``` json
{
"tableName": "nation",
"schemaName": "tpch",
"key": {
"dataFormat": "raw",
"fields": [
{
"name": "redis_key",
"type": "VARCHAR(64)",
"hidden": "true"
}
]
},
"value": {
"dataFormat": "json",
"fields": [
{
"name": "nationkey",
"mapping": "nationkey",
"type": "BIGINT"
},
{
"name": "name",
"mapping": "name",
"type": "VARCHAR(25)"
},
{
"name": "regionkey",
"mapping": "regionkey",
"type": "BIGINT"
},
{
"name": "comment",
"mapping": "comment",
"type": "VARCHAR(152)"
}
]
}
}
```
In redis,such data exists
```shell
127.0.0.1:6379> keys tpch:nation:*
1) "tpch:nation:2"
2) "tpch:nation:4"
3) "tpch:nation:16"
4) "tpch:nation:18"
5) "tpch:nation:10"
6) "tpch:nation:17"
7) "tpch:nation:1"
```
```shell
127.0.0.1:6379> get tpch:nation:1
"{\"nationkey\":1,\"name\":\"ARGENTINA\",\"regionkey\":1,\"comment\":\"al foxes promise slyly according to the regular accounts. bold requests alon\"}"
```
Now we can use redis connector get data from redis,(redis_key don't show,because we set "hidden": "true" )
```shell
lk> select * from redis.tpch.nation;
nationkey | name | regionkey | comment
-----------+----------------+-----------+--------------------------------------------------------------------------------------------------------------------
3 | CANADA | 1 | eas hang ironic, silent packages. slyly regular packages are furiously over the tithes. fluffily bold
9 | INDONESIA | 2 | slyly express asymptotes. regular deposits haggle slyly. carefully ironic hockey players sleep blithely. carefull
19 | ROMANIA | 3 | ular asymptotes are about the furious multipliers. express dependencies nag above the ironically ironic account
2 | BRAZIL | 1 | y alongside of the pending deposits. carefully special packages are about the ironic forges. slyly special
```
**Note**
*if redis.key-prefix-schema-table is false (default is false),all keys in redis will be mapped to table's key,no matching occurs*
Please refer to the `kafka-tutorial` for the description of the ``dataFormat`` as well as various available decoders.
In addition to the above Kafka types, the Redis connector supports ``hash`` type for the ``value`` field which represent data stored in the Redis hash.
Redis connector use `hgetall key` to get data.
``` json
{
"tableName": ...,
"schemaName": ...,
"value": {
"dataFormat": "hash",
"fields": [
...
]
}
}
```
the Redis connector supports ``zset`` type for the ``key`` field which represent key stored in the Redis zset.
if and only if ``zset`` is used as key datafomart,the split is truly supported , because we can use `zrange zsetkey split.start split.end` to get keys of a split.
``` json
{
"tableName": ...,
"schemaName": ...,
"key": {
"dataFormat": "zset",
"name": "zsetkey", //zadd zsetkey score member
"fields": [
...
]
}
}
```
Redis Connector Limitations
---------------------------
only support read operation,don't support write operation.

View File

@ -1,5 +1,5 @@
# Developer Guide
openLooKeng is based on Trino(formerly known as PrestoSQL), and has been forked from the Trino open source project. openLooKeng has additional optimizations, and enhanced features to allow in-situ analytics on any data, anywhere, including geographically remote data sources. This guide is intended for openLooKeng contributors and plugin developers.
openLooKeng is based on Trino 316(formerly known as PrestoSQL), and has been forked from the Trino open source project. openLooKeng has additional optimizations, and enhanced features to allow in-situ analytics on any data, anywhere, including geographically remote data sources. This guide is intended for openLooKeng contributors and plugin developers.

View File

@ -21,10 +21,10 @@ This interface is too big to list in this documentation, but if you are interest
connector. If your underlying data source supports schemas, tables and columns, this interface should be straightforward to implement. If you are attempting to adapt something that is not a relational database (as
the Example HTTP connector does), you may need to get creative about how you map your data source to openLooKeng\'s schema, table, and column concepts.
### ConnectorSplitManger
### ConnectorSplitManager
The split manager partitions the data for a table into the individual chunks that openLooKeng will distribute to workers for processing. For example, the Hive connector lists the files for each Hive partition and creates
one or more split per file. For data so urces that don\'t have partitioned data, a good strategy here is to simply return a single split for the entire table. This is the strategy employed by the Example HTTP connector.
one or more split per file. For data sources that don\'t have partitioned data, a good strategy here is to simply return a single split for the entire table. This is the strategy employed by the Example HTTP connector.
### ConnectorRecordSetProvider

View File

@ -3,7 +3,7 @@ External Function Registration and Push Down
Introduction
------------
The connector can register `external function` into openLooKeng. In Jdbc connector, openLookeng can push them down to data source which support to execute those functions.
The connector can register `external function` into openLooKeng. In Jdbc connector, openLooKeng can push them down to data source which support to execute those functions.
Function Registration in Connector
----------------------------------

View File

@ -216,7 +216,7 @@ An in-depth look at the various annotations relevant to writing an aggregation f
- `@CombineFunction`:
The `@CombineFunction` annotation declares the function used to ombine two state objects. This function is used to merge all the partial aggregation states. It takes two state objects, and merges
The `@CombineFunction` annotation declares the function used to combine two state objects. This function is used to merge all the partial aggregation states. It takes two state objects, and merges
the results into the first one (in the above example, just by adding them together).
- `@OutputFunction`:

View File

@ -5,7 +5,7 @@
* Mac OS X or Linux
* Java 8 Update 161 or higher (8u161+), 64-bit. Both Oracle JDK and OpenJDK are supported.
* AArch64 ([Bisheng JDK 11 or higher](https://www.hikunpeng.com/developer/devkit/compiler?data=JDK))
* AArch64 ([Bisheng JDK 1.8.262 or higher](https://www.hikunpeng.com/developer/devkit/compiler?data=JDK))
* Maven 3.3.9+ (for building)
* Python 2.4+ (for running with the launcher script)

View File

@ -13,7 +13,7 @@ Below is a high level overview of the `Type` interface, for more details see the
- Native encoding:
The interpretation of a value in its native container type form is defined by its `Type`. For some types, such as `BigintType`, it matches the Java interpretation of the native container type (64bit 2\'s complement). However, for other types such as `TimestampWithTimeZoneType`, which also uses `long` for its native container type, the value stored in the `long` is a 8byte binary value combining the timezone and the milliseconds since the unix epoch. In particular, this means that you cannot compare two native values and expect a meaningful result, without knowing the native encoding.
The interpretation of a value in its native container type form is defined by its `Type`. For some types, such as `BigintType`, it matches the Java interpretation of the native container type (64bit 2\'s complement). However, for other types such as `TimestampWithTimeZoneType`, which also uses `long` for its native container type, the value stored in the `long` is a 8byte binary value combining the timezone and the milliseconds since the Unix epoch. In particular, this means that you cannot compare two native values and expect a meaningful result, without knowing the native encoding.
- Type signature:

View File

@ -40,6 +40,10 @@
>
> openLooKeng Bilibili channel: https://space.bilibili.com/627629884
8. Which version of Trino is openLooKeng developed on?
> Based on Trino 316 version development.
## Functions
1. What connectors does the openLooKeng support?

View File

@ -106,7 +106,7 @@ Obviously, not all data types are compatible with each other, below table lists
**Note:**
- Y or Y(#): standard for support implicit convert. But there might be some limitation need your attention. please refer to below item.
- Y or Y(#): standard for support implicit convert. But there might be some limitation need your attention. Please refer to below item.
- N: standard for not support implicit convert
(1): BOOLEAN-\>NUMBER the converted result can be only 0 or 1
@ -123,23 +123,23 @@ Obviously, not all data types are compatible with each other, below table lists
(7): VARCHAR-\>BOOLEAN only \'0\',\'1\',\'TRUE\',\'FALSE\' can be converted. Others will be failed
(8): VARCHAR-\>DECIMAL conversion will fail when its not an numeric or the converted value is out of range of DECIMAL. Scale will be cut off when out of range.
(8): VARCHAR-\>DECIMAL conversion will fail when it's not an numeric or the converted value is out of range of DECIMAL. Scale will be cut off when out of range.
(9): VARCHAR-\>CHAR if length of VARCHAR is larger than CHAR, it will be cut off.
(10): VARCHAR-\>DATE The VARCHAR can only be formatted like:\'YYYY-MM-DD\', e.g. 2000-01-01
(10): VARCHAR-\>DATE The VARCHAR can only be formatted like: \'YYYY-MM-DD\', e.g. 2000-01-01
(11): VARCHAR-\>TIME The VARCHAR can only be formatted like:\'HH:MM:SS.XXX\'
(11): VARCHAR-\>TIME The VARCHAR can only be formatted like: \'HH:MM:SS.XXX\'
(12): VARCHAR-\>TIME ZONE The VARCHAR can only be formatted like:\'HH:MM:SS.XXX XXX\', e.g. 01:02:03.456 America/Los\_Angeles
(12): VARCHAR-\>TIME ZONE The VARCHAR can only be formatted like: \'HH:MM:SS.XXX XXX\', e.g. 01:02:03.456 America/Los\_Angeles
(13): VARCHAR-\>TIMESTAMP The VARCHAR can only be formatted like:YYYY-MM-DD HH:MM:SS.XXX
(13): VARCHAR-\>TIMESTAMP The VARCHAR can only be formatted like: YYYY-MM-DD HH:MM:SS.XXX
(14): DATE-\>TIMESTAMP will auto padding the time with 0. e.g.\'2010-01-01\' -> 2010-01-01 00:00:00.000
(14): DATE-\>TIMESTAMP will auto padding the time with 0. e.g. \'2010-01-01\' -> 2010-01-01 00:00:00.000
(15): TIME-\>TIME WITH TIME ZONE will auto padding the default time zone
(16): TIME-\>TIMESTAMP will auto add the default date:1970-01-01
(16): TIME-\>TIMESTAMP will auto add the default date: 1970-01-01
Miscellaneous
-------------

View File

@ -245,7 +245,7 @@ Returns `true` if this Geometry is an empty geometrycollection, polygon, point e
**ST\_IsSimple(Geometry)** -\> boolean
Returns `true` if this Geometry has no anomalous geometric points, such as self intersection or self tangency.
Returns `true` if this Geometry has no anomalous geometric points, such as self-intersection or self-tangency.
**ST\_IsRing(Geometry)** -\> boolean

View File

@ -13,7 +13,7 @@ Lambda expressions are written with `->`:
x -> CAST(x AS JSON)
x -> x + TRY(1 / 0)
Most SQL expressions can be used in a lambda body, with a fewexceptions:
Most SQL expressions can be used in a lambda body, with a few exceptions:
- Subqueries are not supported. `x -> 2 + (SELECT 3)`
- Aggregations are not supported. `x -> max(y)`

View File

@ -45,7 +45,7 @@ Returns the rank of a value in a group of values. This is similar to `rank`, exc
**ntile(n)** -\> bigint
Divides the rows for each window partition into `n` buckets ranging from `1` to at most `n`. Bucket values will differ by at most `1`. If the number of rows in the partition does not divide evenly into the number
of buckets, then the remainder values are distributed one per bucket,starting with the first bucket.
of buckets, then the remainder values are distributed one per bucket, starting with the first bucket.
For example, with `6` rows and `4` buckets, the bucket values would be as follows: `1` `1` `2` `2` `3` `4`

Binary file not shown.

After

Width:  |  Height:  |  Size: 138 KiB

Binary file not shown.

After

Width:  |  Height:  |  Size: 29 KiB

View File

@ -46,6 +46,7 @@ headless: true
- [Audit Log]({{< relref "./docs/admin/audit-log.md" >}})
- [Reliable Execution]({{< relref "./docs/admin/reliable-execution.md" >}})
- [JDBC Data Source Multi-Split Management]({{< relref "./docs/admin/multi-split-for-jdbc-data-source.md" >}})
- [Extension Physical Execution Planner]({{< relref "./docs/admin/extension-execution-planner.md" >}})
- [Query Optimizer]("#")
- [Table Statistics]({{< relref "./docs/optimizer/statistics.md" >}})
@ -62,10 +63,11 @@ headless: true
- [BTree Index]({{< relref "./docs/indexer/btree.md" >}})
- [HIndex Statements]({{< relref "./docs/indexer/hindex-statements.md" >}})
- [New Index]({{< relref "./docs/indexer/new-index.md" >}})
-
- [Star Tree Cubes](#)
- [Overview] ({{< relref "./docs/preagg/overview.md" >}}>)
- [Statements] ({{< "./docs/preagg/statements.md" >}})
- [Overview]({{< relref "./docs/preagg/overview.md" >}})
- [Join Support]({{< relref "./docs/preagg/join-queries.md" >}})
- [Statements]({{< relref "./docs/preagg/statements.md" >}})
- [Connectors]({{< relref "./docs/connector/_index.md" >}})
- [Carbondata]({{< relref "./docs/connector/carbondata.md" >}})
@ -82,6 +84,7 @@ headless: true
- [JMX]({{< relref "./docs/connector/jmx.md" >}})
- [Kafka]({{< relref "./docs/connector/kafka.md" >}})
- [Kafka Connector Tutorial]({{< relref "./docs/connector/kafka-tutorial.md" >}})
- [Redis] ({{< relref "./docs/connector/redis.md" >}})
- [Local File]({{< relref "./docs/connector/localfile.md" >}})
- [Memory]({{< relref "./docs/connector/memory.md" >}})
- [MongoDB]({{< relref "./docs/connector/mongodb.md" >}})
@ -96,7 +99,8 @@ headless: true
- [TPCH]({{< relref "./docs/connector/tpch.md" >}})
- [VDM]({{< relref "./docs/connector/vdm.md" >}})
- [Kylin]({{< relref "./docs/connector/kylin.md" >}})
- [OmniData]({{< relref "./docs/connector/omnidata.md" >}})
- [Functions and Operators]("#")
- [Logical Operators]({{< relref "./docs/functions/logical.md" >}})
- [Comparison Functions and Operators]({{< relref "./docs/functions/comparison.md" >}})
@ -158,8 +162,6 @@ headless: true
- [GRANT ROLES]({{< relref "./docs/sql/grant-roles.md" >}})
- [INSERT]({{< relref "./docs/sql/insert.md" >}})
- [INSERT OVERWRITE]({{< relref "./docs/sql/insert-overwrite.md" >}})
- [INSERT CUBE]({{< relref "./docs/sql/insert-cube.md" >}})
- [INSERT OVERWRITE CUBE]({{< relref "./docs/sql/insert-overwrite-cube.md" >}})
- [JMX]({{< relref "./docs/sql/jmx.md" >}})
- [PREPARE]({{< relref "./docs/sql/prepare.md" >}})
- [RESET SESSION]({{< relref "./docs/sql/reset-session.md" >}})
@ -172,9 +174,9 @@ headless: true
- [SHOW CACHE]({{< relref "./docs/sql/show-cache.md" >}})
- [SHOW CATALOGS]({{< relref "./docs/sql/show-catalogs.md" >}})
- [SHOW COLUMNS]({{< relref "./docs/sql/show-columns.md" >}})
- [SHOW CREATE CUBE]({{< relref "./docs/sql/show-create-cube.md" >}})
- [SHOW CREATE TABLE]({{< relref "./docs/sql/show-create-table.md" >}})
- [SHOW CREATE VIEW]({{< relref "./docs/sql/show-create-view.md" >}})
- [SHOW CUBES]({{< relref "./docs/sql/show-cubes.md" >}})
- [SHOW FUNCTIONS]({{< relref "./docs/sql/show-functions.md" >}})
- [SHOW EXTERNAL FUNCTION]({{< relref "./docs/sql/show-external-function.md" >}})
- [SHOW GRANTS]({{< relref "./docs/sql/show-grants.md" >}})
@ -217,6 +219,11 @@ headless: true
- [Task Resource]({{< relref "./docs/rest/task.md" >}})
- [Release Notes]("#")
- [1.6.1 (27 Apr 2022)]({{< relref "./docs/releasenotes/releasenotes-1.6.1.md" >}})
- [1.6.0 (30 Mar 2022)]({{< relref "./docs/releasenotes/releasenotes-1.6.0.md" >}})
- [1.5.0 (30 Dec 2021)]({{< relref "./docs/releasenotes/releasenotes-1.5.0.md" >}})
- [1.4.1 (12 Nov 2021)]({{< relref "./docs/releasenotes/releasenotes-1.4.1.md" >}})
- [1.4.0 (15 Oct 2021)]({{< relref "./docs/releasenotes/releasenotes-1.4.0.md" >}})
- [1.3.0 (30 Jun 2021)]({{< relref "./docs/releasenotes/releasenotes-1.3.0.md" >}})
- [1.2.0 (31 Mar 2021)]({{< relref "./docs/releasenotes/releasenotes-1.2.0.md" >}})
- [1.1.0 (30 Dec 2020)]({{< relref "./docs/releasenotes/releasenotes-1.1.0.md" >}})

View File

@ -31,7 +31,7 @@ CREATE INDEX index_name USING bloom ON hive.schema.table (column1) WITH ("bloom.
CREATE INDEX index_name USING bloom ON hive.schema.table (column1) WHERE p in (part1, part2, part3);
```
**Note:** If the table is multi-partitioned (for example, partitioned by colA and colB), only index creation on the **first** level is supported (colA).
**Note:** If the table is multi-partitioned (for example, partitioned by colA and colB), for BTree index, only index creation on the **first** level is supported (colA). Bloom, Bitmap and Minmax index creation on either (colA or colB) is supported.
## SHOW

View File

@ -10,17 +10,13 @@ The `getID()` method in the `Index` interface returns the ID of this index type
### Level
A heuristic index stores additional and usually partial information of a dataset in a more compact way to speed up lookups in various ways. Therefore, each index must
have a domain on which it is applied. For instance, if an index marks the max value of a data set, we must know how big the data set is when we define the "max" value (i.e.
it can be the max of a group of rows, a data partition, or even a whole table). When a new index type is created, it must implement a method `Set<Level> getSupportedIndexLevels();`
which returns the data set level it can support. The levels are defined as an enum in `Index` interface.
A heuristic index stores additional and usually partial information of a dataset in a more compact way to speed up lookups in various ways. Therefore, each index must have a domain on which it is applied. For instance, if an index marks the max value of a data set, we must know how big the data set is when we define the "max" value (i.e. it can be the max of a group of rows, a data partition, or even a whole table). When a new index type is created, it must implement a method `Set<Level> getSupportedIndexLevels();` which returns the data set level it can support. The levels are defined as an enum in `Index` interface.
## Interface overlook
### Indexing methods
Apart from the methods metioned above, this section gives a quick guide on the most important methods needed to create a new index type. For the complete
document on `Index` interface, please refer to the Java Doc of the source code.
Apart from the methods mentioned above, this section gives a quick guide on the most important methods needed to create a new index type. For the complete document on `Index` interface, please refer to the Java Doc of the source code.
There are two main functionalities in the `Index` interface:
@ -46,7 +42,7 @@ boolean addValues(Map<String, List<Object>> values) throws IOException;
Index deserialize(InputStream in) throws IOException;
void serialize(OutputStream out) throws IOException;
```
```
The usage of them are pretty straightforward. A good example to help understand their usage is the source code of `MinMaxIndex`, where adding values is just to
update the `max` and `min` variables according to the input number, and `serialize()/deserialize()`

View File

@ -3,13 +3,13 @@
In addition to the manual deployment of openLooKeng Sever, you can follow below guide to complete deployment faster and easier. The script is friendly to most of linux OS. However, to Ubuntu, you need to manually install the following dependencies:
In addition to the manual deployment of openLooKeng Sever, you can follow below guide to complete the deployment faster and easier. The script is friendly to most of Linux OS. However, to Ubuntu, you need to manually install the following dependencies:
> sshpass1.06 or above
## Deploying openLooKeng on a Single Node
Execute below command can help you download the necessary packages and deploy openLooKeng server in one-click:
Executing the below command can help you download the necessary packages and deploy openLooKeng server in one-click:
```shell
bash <(wget -qO- https://download.openlookeng.io/install.sh)
@ -21,7 +21,7 @@ or:
wget -O - https://download.openlookeng.io/install.sh|bash
```
Normally, you don\'t need to do any thing, except for the installation to complete. It will automatically start the service.
Normally, you don\'t need to do anything, except waiting for the installation to complete. It will automatically start the service.
Execute below command to stop openLooKeng service.:
@ -51,13 +51,13 @@ or:
bash <(wget -qO- https://download.openlookeng.io/install.sh) --multi-node
```
First of all, this command will download scripts and packages required by openLooKeng service. After the download is completed, it will check whether the dependent packages `expect` and `sshpass` are installed. If not, those dependencies will be installed automatically.
First, this command will download scripts and packages required by openLooKeng service. After the download is completed, it will check whether the dependent packages `expect` and `sshpass` are installed. If not, those dependencies will be installed automatically.
Besides, jdk version is required to be greater than 1.8.0\_151. If not, jdk1.8.0\_201 will be installed in the cluster. It is recommended to manually install these dependencies before installing openLooKeng service.
Secondly, the script will download openLooKeng-server tarball and copy that tarball to all the nodes in the cluster. Then install the openLooKeng-server by using this tarball.
Lastly, the script will setup openLooKeng server with the standard configurations, includes configurations for JVM, Node and also for build-in catalogs like `tpch`, `tpcds`, `memory connector`.
Lastly, the script will setup openLooKeng server with the standard configurations, includes configurations for JVM, Node and for built-in catalogs like `tpch`, `tpcds`, `memory connector`.
By design, the script will check if there are existing configuration under directory:
`/home/openlkadmin/.openlkadmin/cluster_node_info`
@ -79,9 +79,9 @@ The general configurations for openLooKeng\'s coordinator, workers are taken fro
`/home/openlkadmin/.openlkadmin/cluster_config_info` and configurations for connectors are taken from the directory `/home/openlkadmin/.openlkadmin/catalog` respectively. If these directories or any required configuration files are absent during the deploy script running, default configuration files will be generated
automatically and deployed to all nodes.
Which means, alternatively, you can add those configuration files before running this deploy script, if you want to customized the deployment.
Which means, alternatively, you can add those configuration files before running this deploy script, if you want to customize the deployment.
If above process all succeed, the deploy script will automatically start the openLooKeng Service for you. Execute below command to stop openLooKeng service.:
If all the above processes succeed, the deploy script will automatically start the openLooKeng Service for you. Execute below command to stop openLooKeng service.:
```shell
/opt/openlookeng/bin/stop.sh
@ -108,7 +108,7 @@ bash <(wget -qO- https://download.openlookeng.io/install.sh) --file <cluster_nod
```
For more help message,execute below command to deploy single node cluster:
For more help, execute below command to deploy single node cluster:
```shell
bash <(wget -qO- https://download.openlookeng.io/install.sh) -h
```
@ -150,9 +150,9 @@ execute below command to deploy the configurations to openLooKeng cluster:
bash /opt/openlookeng/bin/configuration_deploy.sh
```
Note, if you want to add more configrations or customize the configurations, you can add properties to the templates into file located at `/home/openlkadmin/.openlkadmin/.etc_template/coordinator` or `/home/openlkadmin/.openlkadmin/.etc_template/worker`.
Note, if you want to add more configurations or customize the configurations, you can add properties to the templates into file located at `/home/openlkadmin/.openlkadmin/.etc_template/coordinator` or `/home/openlkadmin/.openlkadmin/.etc_template/worker`.
The property format has to be key=\<value\>, where value is wrapped with \'\<\' and \'\>\', which means it it a dynamic value. For example:
The property format must be key=\<value\>, where value is wrapped with \'\<\' and \'\>\', which means it is a dynamic value. For example:
``` properties
http-server.http.port=<http-server.http.port>
@ -174,7 +174,7 @@ It is very easy and straight forward to uninstall openLooKeng Service, simply ru
bash /opt/openlookeng/bin/uninstall.sh
```
This will uninstall openLooKeng Service by removing directory `/opt/openlookeng` and all files inside it. However, the `openlkadmin` user and all the configuration files under`/home/openlkadmin/` will not be removed. If you wan to delete user and configuration files, you need to run the below command:
This will uninstall openLooKeng Service by removing directory `/opt/openlookeng` and all files inside it. However, the `openlkadmin` user and all the configuration files under`/home/openlkadmin/` will not be removed. If you want to delete user and configuration files, you need to run the below command:
```shell
bash /opt/openlookeng/bin/uninstall.sh --all
@ -190,7 +190,7 @@ If you can't access the download URL from the machine where you want to install
1. Also save third party dependencies under `/opt/openlookeng/resource`. That is, download all files from either `https://download.openlookeng.io/auto-install/third-resource/x86/` or `https://download.openlookeng.io/auto-install/third-resource/aarch64/`, depending on the machine's architecture. This should include 1 `OpenJDK` file and 2 `sshpass` files.
1. If you plan to perform multi-node installation, and some nodes in the cluster have a different architecture type from the current machine, then also download the `OpenJDK` file for the other architecture, and save it under `/opt/openlookeng/resource/<arch>`, where `<arch>` is either `x86` or `aarch64`, corresponding to the other architecture.
1. If you plan to perform multi-node installation, and some nodes in the cluster have a different architecture type from the current machine, then also download the `OpenJDK` file for the other architecture and save it under `/opt/openlookeng/resource/<arch>`, where `<arch>` is either `x86` or `aarch64`, corresponding to the other architecture.
After all resources are available, execute below command to deploy single node cluster:
@ -217,7 +217,7 @@ bash /opt/openlookeng/bin/install_offline.sh --help
## Adding Node to Cluster
If you want to add node to make the cluster bigger,execute the below command:
If you want to add node to make the cluster bigger, execute the below command:
```shell
bash /opt/openlookeng/bin/add_cluster_node.sh -n <ip_address_1,ip_address_N>
@ -238,11 +238,11 @@ or:
bash /opt/openlookeng/bin/add_cluster_node.sh --file <add_nodes_file_path>
```
If there are multiple nodes, separated by commas(,). add_ nodes_ File example: ip_address_1,ip_address_2……,ip_address_N.
If there are multiple nodes, separated by commas (,). add_ nodes_ File example: ip_address_1,ip_address_2……,ip_address_N.
## Removing Node to Cluster
If you want to remove node to make the cluster smaller,execute the below command:
If you want to remove node to make the cluster smaller, execute the below command:
```shell
bash /opt/openlookeng/bin/remove_cluster_node.sh -n <ip_address_1,ip_address_N>
@ -263,7 +263,7 @@ or:
bash /opt/openlookeng/bin/remove_cluster_node.sh --file <remove_nodes_file_path>
```
If there are multiple nodes, separate them with commas(,). add_ nodes_ File example: ip_address_1,ip_address_2……,ip_address_N.
If there are multiple nodes, separate them with commas (,). add_ nodes_ File example: ip_address_1,ip_address_2……,ip_address_N.
## See Also

View File

@ -25,7 +25,7 @@ The above properties are described below:
- `hetu.multiple-coordinator.enabled`: Enable multiple coordinators.
- `hetu.embedded-state-store.enabled`: Enable coordinators to start embedded state store.
Note: It is suggested to enable embedded state store on all coordinators(or at least 3) to guarantee the high availability of service when node/network is down.
Note: It is suggested to enable embedded state store on all coordinators (or at least 3) to guarantee the high availability of service when node/network is down.
###Configuring State Store
Please refer to the section [State Store](../admin/state-store.md) to configure state store.

View File

@ -61,7 +61,7 @@ Before an application uses the openLooKeng ODBC driver, the data source DSN must
### Opening the ODBC Data Source Administrator (64-bit)
1. Click **Start**, and choose **Control Panel**.
1. Click **Start** and choose **Control Panel**.
2. In **Control Panel**, click **System and Security**, and then click **Administrative Tools**.
@ -171,6 +171,6 @@ You can obtain the details about data types by calling **SQLGetTypInfo** in **Ca
The openLooKeng ODBC driver supports **both ANSI and Unicode** applications. The default connection character set is the system default character set for ANSI applications and utf8 for Unicode applications. If the character set used by the application is different from the above-mentioned character set, it may cause garbled characters. For this, the user should specify the connection character set to adapt to the character set required by the application. The corresponding configuration of the connection character set is described as follows.
When calling the ODBC API to retrieve data, if bound to the SQL_C_WCHAR C data type buffer, the driver will return the Unicode encoded result for both ANSI and Unicode applications. When bound to the SQL_C_CHAR C data type buffer, by deafult, the driver will return to the ANSI application the result encoded in system default character set, and for Unicode application the driver will return the result encoded in utf8. If the encoding character set used by the application does not match the default, the result may be garbled. To this end, the user should configure the connection character set to specify the encoding of the result. For example, if the application has garbled Chinese characters, you can try to configure the connection character set to GBK or GB2312.
When calling the ODBC API to retrieve data, if bound to the SQL_C_WCHAR C data type buffer, the driver will return the Unicode encoded result for both ANSI and Unicode applications. When bound to the SQL_C_CHAR C data type buffer, by default, the driver will return to the ANSI application the result encoded in system default character set, and for Unicode application the driver will return the result encoded in utf8. If the encoding character set used by the application does not match the default, the result may be garbled. To this end, the user should configure the connection character set to specify the encoding of the result. For example, if the application has garbled Chinese characters, you can try to configure the connection character set to GBK or GB2312.
While configuring data source all connection character sets supported by the openLooKeng ODBC driver can be set in the **Character Set** drop-down box on the page 3 of the User interface. User can select the connection character from the drop-down box after the **Test DSN** is success.

View File

@ -0,0 +1,65 @@
## Join Query Support
StarTree Cube can help optimize aggregation over join queries as well. The optimizer looks for aggregation subtree pattern in the logical plan that typically looks like following.
```
AggregationNode
|- ProjectNode[Optional]
. |- ProjectNode[Optional]
. . |- FilterNode[Optional]
. . . |- JoinNode
. . . . [More Joins]
. . . . .
. . . . |- ProjectNode[Optional] - Left
. . . . . |- TableScanNode [Fact Table]
. . . . |- ProjectNode[Optional] - Right
. . . . . |- TableScanNode [Dim Table]
```
If the query matches the pattern, the optimizer rewrites the logical plan by replacing the Fact TableScanNode with Cube TableScanNode. This is similar to the single
table rewrite.
### Star Schema Support
Join Query optimizer supports star schema only. A star schema is a data warehousing architecture model where one fact table references multiple dimension tables, which, when viewed as a diagram,
looks like a star with the fact table in the center and the dimension tables radiating from it. All kinds of joins are supported.
![star-schema](../images/star-schema.png "star schema")
### Cube Management
`Create Cube` can be still be used to define Cubes to optimize Join queries as well. The difficult part is identifying GROUP construct while building the Cubes. With single table
queries, the GROUP BY clause will contain columns only from same the table. But with join queries, especially star schema queries, the GROUP BY contain columns from Dimension tables and not the Fact table.
Let's analyze more with following query
```sql
SELECT SUM(lo_revenue) AS lo_revenue, d_year, p_brand
FROM lineorder
LEFT JOIN dates ON lo_orderdate = d_datekey
LEFT JOIN part on lo_partkey = p_partkey
LEFT JOIN supplier on lo_suppkey = s_suppkey
WHERE p_category = 'MFGR#12' AND s_region = 'AMERICA'
GROUP BY d_year, p_brand
ORDER BY d_year, p_brand;
```
Here `lineorder` is the Fact table and `dates`, `part`, `supplier` are the Dimension tables. Cubes will be defined on the `lineorder` table. The group by columns `d_year`, `p_brand` are part
of the dimension tables `dates` and `part` appropriately. They cannot be used directly in `CREATE CUBE` statement. The proper solution is to use the foreign key columns of `lineorder` table in
the GROUP construct while building Cubes.
```sql
CREATE CUBE lineorder_cube ON lineorder WITH(
AGGREGATIONS = (sum(lo_revenue)),
GROUP = (lo_orderdate, lo_partkey, lo_suppkey));
```
The optimizer parses the join conditions and uses those columns to identify the matching Cubes. The performance gain is realized if Cube size is smaller than fact table.
### Limitations
* Only star schema is supported.
* Count distinct not supported because Cube does not store actual dimension values.
* Queries won't be optimized if Cubes are defined on both Fact and Dimension as the optimizer does not have capability to differentiate between two.
* If Cubes are defined on more than one table of the Join query - then Optimizer does not work. Assumption is that Cubes are defined only the Fact table.
* Supports only simple aggregation like SUM, COUNT, AVG, MIN, MAX - defined on Single column. Cube does not support SUM(revenue - supplycost) aggregation. The following query cannot be optimized using Cube.
```
SELECT sum(lo_extendedprice * lo_discount) AS revenue
FROM lineorder
WHERE toYear(lo_orderdate) = 1993 AND lo_discount BETWEEN 1 AND 3 AND lo_quantity < 25;
```
### Future
* Support for snowflake schema
* Building a single cube over multiple tables

View File

@ -15,8 +15,8 @@ Few of the Cube properties are
- Query latency is reduced by rewriting the logical plan to use Cube instead of the original table.
## Cube Optimizer Rule
As part of logical plan optimization, Cube optimizer rule analyzes and optimizes the aggregation sub-tree of the logical plan with Cubes.
The rule looks for the aggregation sub-tree that typically looks like the following
As part of logical plan optimization, Cube optimizer rule analyzes and optimizes the aggregation subtree of the logical plan with Cubes.
The rule looks for the aggregation subtree that typically looks like the following
```
AggregationNode
@ -27,17 +27,17 @@ AggregationNode
|- TableScanNode
```
The rule parses through the sub-tree and identifies the table name, aggregate functions, where clause, group by clause that is matched with Cube metadata
The rule parses through the subtree and identifies the table name, aggregate functions, where clause, group by clause that is matched with Cube metadata
to identify any Cube that can help optimize the query. In case of multiple match, recently created Cube is selected for optimization. If any match found, entire
aggregation sub-tree is rewritten using the Cube. This optimizer uses the TupleDomain construct to match if predicates provided in the Query can be supported by the
Cubes.
aggregation subtree is rewritten using the Cube. This optimizer uses the TupleDomain construct to match if predicates provided in the Query can be supported by the
Cubes.
The following picture depicts the change in the logical plan after the optimization.
![img](../images/cube-logical-plan-optimizer.png)
## Recommended Usage
1. Cubes are mose useful for iceberg queries that takes huge input and produces small input
1. Cubes are most useful for iceberg queries that takes huge input and produces small output.
2. Query performance is best when size of the Cube is less that on the actual table on which Cube was built.
3. Cubes need to be rebuilt if the source table is updated.
@ -47,7 +47,11 @@ operation on the update is considered as a change in the existing data even if o
can't be differentiated, Cubes can't be used as it might result in incorrect result. We are working on a solution to overcome this limitation.
## Supported Connectors
The following are supported Connectors for storing a cube
Star Tree Cube can be stored in following Connectors
1. Hive
2. Memory
Tables from following Connectors can be used as source to build a StarTree Cube.
1. Hive
2. Memory
3. Clickhouse
@ -58,8 +62,6 @@ The following are supported Connectors for storing a cube
2.1. Overcome the limitation of Creating Cube for larger dataset.
2.2. Update cube if source table has been updated.
## Enabling and Disabling StarTree Cube
To enable:
```sql
@ -73,13 +75,13 @@ SET SESSION enable_star_tree_index=false;
## Configuration Properties
| Property Name | Default Value | Required| Description|
|---------------------------------------------------|---------------------|---------|--------------|
| optimizer.enable-star-tree-index | false | No | Enables StarTree Cube|
| cube.metadata-cache-size | 50 | No | The maximum number of metadata for StarTree Cubes that could be loaded into cache before eviction happens|
| optimizer.enable-star-tree-index | false | No | Enables StarTree Cube |
| cube.metadata-cache-size | 50 | No | The maximum number of metadata for StarTree Cubes that could be loaded into cache before eviction happens |
| cube.metadata-cache-ttl | 1h | No | The maximum time to live of StarTree Cubes that are be loaded into cache before eviction happens |
## Dependencies
StarTree Cube relies on Hetu metastore to store the Cube related metadata.
StarTree Cube relies on Hetu Metastore to store the Cube related metadata.
Please check [Hetu Metastore](../admin/meta-store.md) for more information.
## Examples
@ -118,11 +120,22 @@ SELECT nationkey, avg(nationkey), max(regionkey) FROM nation WHERE nationkey >=
Since the data inserted into the Cube was for `nationkey >= 5`, only queries matching this condition will utilize the Cube.
Queries not matching the condition would continue to work but won't use the Cube.
## Building Cube for Large dataset
If the source table of a Cube gets updated, the corresponding Cube gets expired automatically. In order to overcome
this issue, we have added support in openLooKeng CLI by introducing **RELOAD CUBE** command. The user will have the
ability to manually reload a cube if the status of the Cube becomes INACTIVE or EXPIRED. The syntax to reload the
Cube nation_cube is as follows,
```sql
RELOAD CUBE nation_cube
```
Please note that this feature is only supported via the CLI. During this reload process if an unexpected error occurs, the user will get to see the original SQL statement
to recreate the cube manually.
## Building Cube for Large Dataset
One of the limitations with the current implementation is that Cube cannot be built for a larger dataset at once. This is due to the cluster memory limitation.
Processing large number of rows requires more memory than cluster is configured with. This results in query failing with message **Query exceeded per-node user memory
limit**. To overcome this issue, **INSERT INTO CUBE** sql support was added. The user has ability to build a Cube for larger data by executing multiple
insert into cube statements. The insert statement accepts a where clause, and it can be used to limit the number of processed and inserted into Cube.
limit**. To overcome this issue, **INSERT INTO CUBE** SQL support was added. The user has ability to build a Cube for larger data by executing multiple
insert into Cube statements. The insert statement accepts a where clause, and it can be used to limit the number of processed and inserted into Cube.
This section explains the steps to build a Cube for larger dataset.
@ -146,7 +159,7 @@ INSERT INTO CUBE store_sales_cube WHERE ss_sold_date_sk BETWEEN 2451911 AND 2422
```
### Solution 1)
To overcome this issue, multiple insert statements can be used into process rows and insert into cube and the number of rows can be limited by using where clause;
To overcome this issue, multiple insert statements can be used to process rows and insert into Cube and the number of rows can be limited by using where clause;
```sql
INSERT INTO CUBE store_sales_cube WHERE ss_sold_date_sk BETWEEN 2451911 AND 2452010;
@ -157,8 +170,8 @@ INSERT INTO CUBE store_sales_cube WHERE ss_sold_date_sk BETWEEN 2452211 AND 2452
### Solution 2)
CLI has been modified to support creating Cubes for larger dataset and without need for multiple insert statements. CLI internally handles this process.
Once the user runs create cube statement with where clause, the CLI takes care of creating the cube as well as inserting the data into it. This process improves the user experience and
improves the memory footprint based on the cluster memory limits. CLI internally parses the converts the statement into one create cube statement followed by
Once the user runs create Cube statement with where clause, the CLI takes care of creating the Cube as well as inserting the data into it. This process improves the user experience and
improves the memory footprint based on the cluster memory limits. CLI internally parses the converts the statement into one create Cube statement followed by
one or more insert statements. This change is only works if user executes the command from CLI and not via any other means i.e. JDBC, etc...
```sql
@ -176,46 +189,73 @@ SHOW CUBES;
```
**Note:**
1. The system will try to rewrite all type of Predicates into a Range to see if they can be merged together.
1. The system will try to rewrite all type of Predicates into a Range to see if they can be merged together.
All continuous predicates will be merged into a single range predicate and remaining predicates are untouched.
Only the following types are supported and can be merged together.
`Integer, TinyInt, SmallInt, BigInt, Date`
For other data types, it is difficult to identify if two predicates are continuous therefore they cannot be merged together. And because of this issue, there is
possibility that particular cube may not be used during query optimization even if the cube has all the required data. For example,
Only the following types are supported and can be merged together.
`Integer, TinyInt, SmallInt, BigInt, Date, String`
For String data type, predicate merge logic functionally works only if the Strings are ending with a digit and all are of same length.
For example,
```sql
INSERT INTO CUBE store_sales_cube WHERE store_id BETWEEN 'A01' AND 'A10';
INSERT INTO CUBE store_sales_cube WHERE store_id BETWEEN 'A11' AND 'A20';
```
Here these two predicates cannot be merged into store_id BETWEEN 'A01' AND 'A20'; So the cube won't be used
for queries that are spanning over two the predicates;
After the insertion, the two predicates will be merged into `'A01' AND 'A20'`
```sql
SELECT ss_store_id, sum(ss_sales_price) WHERE ss_store_id BETWEEN 'A05' AND 'A15'; - Cube won't be used for optimizing this query. This is a limitation as of now.
SELECT ss_store_id, sum(ss_sales_price) WHERE ss_store_id BETWEEN 'A05' AND 'A15'; - Cube would be used for this query.
```
Because of the predicate rewrite some of the following queries can't be supported
Consider the following example where `store_id` values are not of same length.
```sql
INSERT INTO CUBE store_sales_cube WHERE store_id = 'A1';
INSERT INTO CUBE store_sales_cube WHERE store_id = 'A2'
```
store_id predicate will be rewritten as `store_id >= 'A1' and store < 'A3'` as per the varchar predicate merge logic;
```sql
INSERT INTO CUBE store_sales_cube WHERE store_id = 'A10'
```
The above query would fail because `A10` is subset of the range `store_id >= 'A1' and store < 'A3'`. So Users should be wary of this issue.
For other data types, it is difficult to identify if two predicates are continuous therefore they cannot be merged together. And because of this issue, there is
possibility that particular Cube may not be used during query optimization even if the Cube has all the required data.
2. Predicate rewrite has some limitations as well. Consider the following query
```sql
INSERT INTO CUBE store_sales_cube WHERE ss_sold_date_sk > 2451911;
```
The predicate is rewriten as ss_sold_date_sk >= 2451912 to be prepare for merging continous predicates.
Since the predicate is rewritten, they query using ss_sold_date_sk > 2451911 predicate will not match with Cube predicate so Cube won't be used to
optimize the query. The same is applicable for predicates with <= operator. ie. ss_sold_date_sk <= 2451911 is rewritten as ss_sold_date_sk < 2451912
```
The predicate is rewritten as ss_sold_date_sk >= 2451912 to support merging continuous predicates.
Since the predicate is rewritten, they query using ss_sold_date_sk > 2451911 predicate will not match with Cube predicate so Cube won't be used to
optimize the query. The same is applicable for predicates with <= operator. ie. ss_sold_date_sk <= 2451911 is rewritten as ss_sold_date_sk < 2451912
```sql
SELECT ss_sold_date_sk, .... FROM hive.tpcds_sf1.store_sales WHERE ss_sold_date_sk > 2451911
```
3. Only Single column predicates can be merged.
3. Only single column predicates can be merged.
## Open issues and Limitations
1. StarTree Cube is only effective when the group by cardinality is considerably fewer than the number of rows in source table.
2. A significant amount of user effort required in maintaining Cubes for large datasets.
3. Only incremental insert into cube is supported. Cannot delete specific rows from Cube.
3. Only incremental insert into Cube is supported. Cannot delete specific rows from Cube.
4. Cubes created on a transaction table may expire automatically even if the source table has not been updated. This is due to the compaction policy which
merges delta files into single large ORC file which in turn changes the last modified of time of the table. Cube status is determined by comparing last modified
timestamp of table when cube was created with the last modified time of the table when queries are executed.
5. Openlookeng CLI has been modified to ease the process of creating Cubes for larger datasets. But still there are limitations with this implementation
as the process involves merging multiple cube predicates into one. Only cube predicates defined on Integer, Long and Date types can be merged properly. Support for Char,
String types still need to be implemented.
timestamp of table when Cube was created with the last modified time of the table when queries are executed.
5. OpenLooKeng CLI has been modified to ease the process of creating Cubes for larger datasets. But still there are limitations with this implementation
as the process involves merging multiple Cube predicates into one. Only Cube predicates defined on Integer, Long and Date types can be merged properly. Support for Char,
String types still need to be implemented.
6. Varchar predicates can be merged only if the values are of same length.
## Performance Optimizations on Star Tree
1. Star Tree Query re-write optimization for same group by columns: If the group by columns of the cube and query matches, the query is
re-written internally to select the pre-aggregated data. If the group by columns does not matches, the additional aggregations are
internally applied on the re-written query.
2. Star Tree table scan optimization for Average aggregation function: If the group by columns of the cube and query matches, the select query
is re-written internally to select the startree cube's pre-aggregated Average column data. If the group by columns does not match,
the select query is re-written internally to select the startree cube's pre-aggregated Sum and Count column data, from which the
average is later calculated.

View File

@ -44,7 +44,7 @@ Create a new partitioned Cube `orders_cube`:
partitioned_by = ARRAY['orderdate']
)
Create a new Cube `orders_cube` with some source data filter
Create a new Cube `orders_cube` with some source data filter:
CREATE CUBE orders_cube ON orders WITH (
AGGREGATIONS = ( SUM(totalprice), COUNT DISTINCT(orderid) ),
@ -52,7 +52,7 @@ Create a new Cube `orders_cube` with some source data filter
FILTER = (orderdate BETWEEN 2512450 AND 2512460)
)
Create a new Cube `orders_cube` with some additional predicate on Cube columns
Create a new Cube `orders_cube` with some additional predicate on Cube columns:
CREATE CUBE orders_cube ON orders WITH (
AGGREGATIONS = ( SUM(totalprice), COUNT DISTINCT(orderid) ),
@ -60,7 +60,7 @@ Create a new Cube `orders_cube` with some additional predicate on Cube columns
FILTER = (orderdate BETWEEN 2512450 AND 2512460)
) WHERE orderstatus = 'PENDING';
This is same as following
This is same as following:
CREATE CUBE orders_cube ON orders WITH (
AGGREGATIONS = ( SUM(totalprice), COUNT DISTINCT(orderid) ),
@ -86,12 +86,12 @@ INSERT INTO CUBE cube_name [WHERE condition]
```
### Description
`CREATE CUBE` statement creates Cube without any data. To insert data into Cube, use `INSERT INTO CUBE` sql.
`CREATE CUBE` statement creates Cube without any data. To insert data into Cube, use `INSERT INTO CUBE` SQL.
The `WHERE` clause is optional. If predicate is provided, only data matching the given predicate are processed from the source table and inserted into the Cube.
Otherwise, entire data from the source table is processed and inserted into Cube.
### Examples
Insert data into the `orders_cube` Cube
Insert data into the `orders_cube` Cube:
INSERT INTO CUBE orders_cube WHERE orderdate > date '1999-01-01';
INSERT INTO CUBE order_all_cube;
@ -117,8 +117,10 @@ INSERT OVERWRITE CUBE cube_name [WHERE condition]
```
### Description
Similar to INSERT INTO CUBE statement but with this statement the existing data is overwritten. Predicates
are optional.
Similar to `INSERT INTO CUBE` statement but with this statement the existing data is overwritten. Predicates
are optional.`INSERT OVERWRITE CUBE` is not supported on partitioned cubes. Cubes are essentially stored as tables and so `INSERT OVERWRITE` only
replaces the matching partitions and does not overwrite the entire table. So this operation is blocked on partitioned cube.
Drop and recreate cube if needed.
### Examples
Insert data based on condition into the `orders_cube` Cube:
@ -148,7 +150,25 @@ Show Cubes for `orders` table:
```sql
SHOW CUBES FOR orders;
```
## RELOAD CUBE
### Synopsis
``` sql
RELOAD CUBE cube_name
```
### Description
Reloads the Cube if the source table has been updated.
### Examples
If the source table `orders` of the cube `orders_cube` gets updated then the status of the cube `orders_cube`
gets EXPIRED. Use the command `RELOAD CUBE cube_name` to overcome this issue as follows:
```sql
RELOAD CUBE orders_cube
```
## DROP CUBE
### Synopsis

View File

@ -0,0 +1,14 @@
# Release 1.4.1 (12 Nov 2021)
## Key Features
This release mainly adds the introduction of OmniData Connector and jdk8 support under arm architecture.
| Area | Feature | PR #s |
| ----------------------- | ------------------------------------------------------------ | ------------------------------------------------------------ |
| Data Source Connector | openLooKeng supports offload operators to the near data side for process through omnidata connector, reducing the transmission of invalid data on the network and effectively improving the performance of big data computing. | 1219 |
| Arm architecture | Eliminate the mandatory requirements for Java version under arm architecture caused by JDK paused problem, and support jdk1.8.262 and above under arm architecture. | 1214 |
## Obtaining the Document
For details, see [https://gitee.com/openlookeng/hetu-core/tree/1.4.1/hetu-docs/en](https://gitee.com/openlookeng/hetu-core/tree/1.4.1/hetu-docs/en)

View File

@ -0,0 +1,23 @@
# Release 1.5.0
## Key Features
| Area | Feature |
| ---------------- | ------------------------------------------------------------ |
| Star Tree | 1. Support for optimizing join queries such as star schema queries.<br/>2. Optimized the query plan by eliminating unnecessary aggregations on top of the cube since the cube already contains the rolled up results. The performance optimizations benefits queries whose group by clause exactly matches the cubes group.<br/>3. Bug fixes to further enhance the usability, and robustness of cubes. |
| Memory Connector | 1. Improved performance of memory connector by adding support for memory table partitioning to allow data skipping of entire partitions<br/>2. Collect statistics on memory table to support openLooKeng cost based optimizers. |
| Task Recovery | Fixed several important bugs to address data inconsistency issues, and query hanging issues that occasionally occur during high concurrency, and during worker failures. |
| OLK-on-Yarn | Support deploying an HA-enabled openLooKeng cluster instance on-yarn, that contains a reverse proxy (ngnix by default), and 2 or more coordinator nodes. The cluster can be horizontally scaled manually by adding and removing yarn containers to the coordinator and worker components.|
| Spill to Disk | Optimized the spill to disk mechansim to directly write Pages to disk instead of buffering it. Changed the strategy to spill those operators which can free up maximum memory. This resulted in improvement of spill to disk performance by 30% |
## Known Issues
| Category | Description | Gitee issue |
| --------------------- | ------------------------------------------------------------ | ------------------------------------------------------------ |
| Task Recovery |When the query process reaches stage1,a worker is killed result some values become smaller(occasionally). |[I4M2LW](https://e.gitee.com/open_lookeng/issues/list?issue=I4M2LW) |
| Memory Connector |When a table with the index_columns parameter is queried, an error message is displayedjava.lang.NullPointerException. | [I4NVW3](https://e.gitee.com/open_lookeng/issues/list?issue=I4NVW3)|
| |When drop and then create a same partitioned table with data type double, query result rows is greater than expected. | [I4LUDF](https://e.gitee.com/open_lookeng/issues/list?issue=I4LUDF) |
## Obtaining the Document
For details, see [https://gitee.com/openlookeng/hetu-core/tree/1.5.0/hetu-docs/en](https://gitee.com/openlookeng/hetu-core/tree/1.5.0/hetu-docs/en)

View File

@ -0,0 +1,25 @@
# Release 1.6.0
## Key Features
| Area | Feature |
| --------------------- | ------------------------------------------------------------ |
| Star Tree | Support update cube command to allow admin to easily update an existing cube when the underlying data changes |
| Bloom Index | Hindex-Optimize Bloom Index Size-Reduce bloom index size by 10X+ times |
| Task Recovery | 1. Improve failure detection time: It need take 300s to determine a task is failed and resume after that. Improving this would improve the resume & also the overall query time<br/>2. snapshotting speed & size: When sql execute takes a snapshot, now use direct Java serialization which is slow and also takes more size. Using kryo serialization would reduce size and also increase speed there by increasing the overall throughput |
| Spill to Disk | 1. Spill to disk speed & size improvement: When spill happens during HashAggregation & GroupBy, the data serialized to disk is slow and also size is more. It can improve the overall performance by reducing size and also improving the writing speed. Using kryo serialization improves both speed and reduces size<br/>2. Support spilling to hdfs: Currently data can spill to multiple disks, now support spill to hdfs to improve throughput<br/>3. Async spill/unspill: When revocable memory crosses threshold and spill is triggered, it blocks accepting the data from the downstream operators. Accepting this and adding to the existing spill would help to complete the pipeline faster<br/>4. Enable spill for right outer & full join for spilling: It dont spill the build side data when the join type is right outer or full join as it needs the entire data in memory for lookup. This leads to out of memory when the data size is more. Instead by enable spill and create a Bloom Filter to identify the data spilled and use it during join with probe side |
| Connector Enhancement | Support data update and delete operator for PostgreSQL and openGauss |
## Known Issues
| Category | Description | Gitee issue |
| ------------- | ------------------------------------------------------------ | --------------------------------------------------------- |
| Task Recovery | When a snapshot is enabled and a CTAS with transaction is executed, an error is reported in the SQL statement. | [I502KF](https://e.gitee.com/open_lookeng/issues/list?issue=I502KF) |
| | An error occurs occasionally when snapshot is enabled and exchange.is-timeout-failure-detection-enabled is disabled. | [I4Y3TQ](https://e.gitee.com/open_lookeng/issues/list?issue=I4Y3TQ) |
| Star Tree | In the memory connector, after the star tree is enabled, data inconsistency occurs during query. | [I4QQUB](https://e.gitee.com/open_lookeng/issues/list?issue=I4QQUB) |
| | When the reload cube command is executed for 10 different cubes at the same time, some cubes fail to be reloaded. | [I4VSVJ](https://e.gitee.com/open_lookeng/issues/list?issue=I4VSVJ) |
## Obtaining the Document
For details, see [https://gitee.com/openlookeng/hetu-core/tree/1.6.0/hetu-docs/en](https://gitee.com/openlookeng/hetu-core/tree/1.6.0/hetu-docs/en)

View File

@ -0,0 +1,15 @@
# Release 1.6.1 (27 Apr 2022)
## Key Features
This release is mainly about modification and enhancement of some SPIs, which are used in more scenarios.
| Area | Feature | PR #s |
| ----------------------- | ------------------------------------------------------------ | ------------------------------------------------------------ |
| Data Source statistics | The method of obtaining statistics is added so that statistics can be directly obtained from the Connector. Some operators can be pushed down to the Connector for calculation. You may need to obtain statistics from the Connector to display the amount of processed data. | 1450 |
| Operator processing extension | Users can customize the physical execution plan of worker nodes. Users can use their own operator pipelines to replace the native implementation to accelerate operator processing. | 1436 |
| HIVE UDF extension | Adds the adaptation of HIVE UDF function namespace to support the execution of UDFs (including GenericUDF) written based on the HIVE UDF framework. | 1456 |
## Obtaining the Document
For details, see [https://gitee.com/openlookeng/hetu-core/tree/1.6.1/hetu-docs/en](https://gitee.com/openlookeng/hetu-core/tree/1.6.1/hetu-docs/en)

View File

@ -1,4 +1,3 @@
Built-in System Access Control
==============================
@ -57,7 +56,10 @@ composed of the following fields:
- `user` (optional): regex to match against user name. Defaults to `.*`.
- `catalog` (optional): regex to match against catalog name. Defaults to `.*`.
- `allow` (required): boolean indicating whether a user has access to the catalog
- ``allow`` (required): string indicating whether a user has access to the catalog.
This value can be ``all``, ``read-only`` or ``none``, and defaults to ``none``.
Setting this value to ``read-only`` has the same behavior as the ``read-only``
system access control plugin.
**Note**
@ -65,7 +67,9 @@ composed of the following fields:
*By default, all users have access to the `system` catalog. You can override this behavior by adding a rule.*
For example, if you want to allow only the user `admin` to access the `mysql` and the `system` catalog, allow all users to access the `hive` catalog, and deny all other access, you can use the following rules:
For example, if you want to allow only the user ``admin`` to access the``mysql`` and the ``system`` catalog,
allow all users to access the ``hive`` catalog, allow the user ``alice`` read-only access to the ``postgresql``
catalog, and deny all other access, you can use the following rules:
``` json
{
@ -73,15 +77,20 @@ For example, if you want to allow only the user `admin` to access the `mysql` an
{
"user": "admin",
"catalog": "(mysql|system)",
"allow": true
"allow": all
},
{
"catalog": "hive",
"allow": true
"allow": all
},
{
"user": "alice",
"catalog": "postgresql",
"allow": "read-only"
},
{
"catalog": "system",
"allow": false
"allow": none
}
]
}
@ -195,6 +204,7 @@ If you want to allow users to use the extactly same name as their Kerberos prin
```
### Node State Rules
These rules govern the node state info particular users can access. The user is granted access to update a node state based on the first matching rule read from top to bottom. If no rule matches, access is denied. Each rule is
composed of the following fields:

View File

@ -7,7 +7,7 @@ Overview
Apache Ranger delivers a comprehensive approach to security for a Hadoop cluster. It provides a centralized platform to define, administer and manage security policies consistently across Hadoop components. Check [Apache Ranger Wiki](https://cwiki.apache.org/confluence/display/RANGER/Index) for detail introduction and user guide.
[openlookeng-ranger-plugin](https://gitee.com/openlookeng/openlookeng-ranger-plugin) is a ranger plugin for openLooKeng to enable, monitor and manage comprehensive data security.
[openlookeng-ranger-plugin](https://gitee.com/openlookeng/openlookeng-ranger-plugin) is developed based on Ranger 2.1.0, which is a ranger plugin for openLooKeng to enable, monitor and manage comprehensive data security.
Build Process
-------------------------

View File

@ -38,7 +38,7 @@ Cache data with complex predicate string:
Limitations
-----------
- Only Hive connector(ORC Format) support this functionality at this time. See connector documentation for more details.
- Only Hive connector (ORC Format) support this functionality at this time. See connector documentation for more details.
- Does not support `LIKE` in `WHERE` clause.
- Does not support 'OR' operator in complex predicate.

View File

@ -14,7 +14,7 @@ AS query
Description
-----------
Create a new view of a [SELECT](./select.md) query. The view is a logical table that can be referenced by future queries. Views do not contain any data. Instead, the query stored by the view is executed everytime the view is referenced by another query.
Create a new view of a [SELECT](./select.md) query. The view is a logical table that can be referenced by future queries. Views do not contain any data. Instead, the query stored by the view is executed every time the view is referenced by another query.
The optional `OR REPLACE` clause causes the view to be replaced if it already exists rather than raising an error.

View File

@ -28,7 +28,7 @@ Assume `orders` is not a partitioned table, and have 100 rows, then execute belo
INSERT OVERWRITE orders VALUES (1, 'SUCCESS', '10.25', DATA '2020-01-01');
Then the `orders` table will only have 1 rows, that is the data specified in the `VALUE` clause.
Then the `orders` table will only have 1 row that is the data specified in the `VALUE` clause.
Assume `users` has 3 columns: (`id`, `name`, `state`) and partitioned by `state`, and the existing data has follow rows:

View File

@ -508,7 +508,7 @@ _col0
**INTERSECT**
`INTERSECT` returns only the rows that are in the result sets of both the first and the second queries. The following is an example of one of the simplest possible `INTERSECT` clauses. It selects the values `13`
and `42` and combines this result set with a second query that selects the value `13`. Since `42` is only in the result set of the first query, it is not included in the final results.:
and `42` and combines this result set with a second query that selects the value `13`. Since `42` is only in the result set of the first query, it is not included in the final results:
SELECT * FROM (VALUES 13, 42)
INTERSECT
@ -524,7 +524,7 @@ _col0
**EXCEPT**
`EXCEPT` returns the rows that are in the result set of the first query, but not the second. The following is an example of one of the simplest possible `EXCEPT` clauses. It selects the values `13` and `42` and
combines this result set with a second query that selects the value `13`. Since `13` is also in the result set of the second query, it is not included in the final result.:
combines this result set with a second query that selects the value `13`. Since `13` is also in the result set of the second query, it is not included in the final result:
SELECT * FROM (VALUES 13, 42)
EXCEPT
@ -746,7 +746,7 @@ Joins allow you to combine data from multiple relations.
### CROSS JOIN
A cross join returns the Cartesian product (all combinations) of two relations. Cross joins can either be specified using the explit `CROSS JOIN` syntax or by specifying multiple relations in the `FROM` clause.
A cross join returns the Cartesian product (all combinations) of two relations. Cross joins can either be specified using the explicit `CROSS JOIN` syntax or by specifying multiple relations in the `FROM` clause.
Both of the following queries are equivalent:

View File

@ -0,0 +1,32 @@
SHOW CREATE CUBE
=================
Synopsis
--------
``` sql
SHOW CREATE CUBE cube_name
```
Description
-----------
Show the SQL statement that creates the specified cube.
Examples
--------
Create a cube `orders_cube` on `orders` table as follows
CREATE CUBE orders_cube ON orders WITH (AGGREGATIONS = (avg(totalprice), sum(totalprice), count(*)),
GROUP = (custKEY, ORDERkey), format= 'orc')
Use `SHOW CREATE CUBE` command to show the SQL statement that was used to create the cube `orders_cube`:
SHOW CREATE CUBE orders_cube;
``` sql
CREATE CUBE orders_cube ON orders WITH (AGGREGATIONS = (avg(totalprice), sum(totalprice), count(*)),
GROUP = (custKEY, ORDERkey), format= 'orc')
```

View File

@ -1,6 +1,6 @@
# 审计日志
openLooKeng审计日志记录功能是一个自定义事件监听器在查询创建和完成成功或失败时调用。审计日志包含以下信息
openLooKeng审计日志记录功能是一个自定义事件监听器监听openLooKeng集群启停与集群中节点的动态添加与删除事件监听WebUi用户登录与退出事件监听查询事件,在查询创建和完成(成功或失败)时调用。审计日志包含以下信息:
1. 事件发生时间
2. 用户ID
@ -23,15 +23,17 @@ openLooKeng审计日志记录功能是一个自定义事件监听器在查询
hetu.event.listener.type=AUDIT
hetu.event.listener.listen.query.creation=true
hetu.event.listener.listen.query.completion=true
hetu.auditlog.logoutput=/var/log/
hetu.auditlog.logconversionpattern=yyyy-MM-dd.HH
```
其他审计日志记录属性包括:
`hetu.event.listener.audit.file`可选属性用于定义审计文件的绝对文件路径。确保运行openLooKeng服务器的进程对该目录有写权限
`hetu.event.listener.type`用于定义审计日志的记录类型允许的值为AUDIT和LOGGER
`hetu.event.listener.audit.filecount`:可选属性,用于定义要使用的文件数
`hetu.auditlog.logoutput`用于定义审计文件的绝对目录路径。确保运行openLooKeng服务器的进程对该目录有写权限
`hetu.event.listener.audit.limit`:可选属性,用于定义写入任一文件的最大字节数
`hetu.auditlog.logconversionpattern`用于定义审计日志的轮转模式。允许的值为yyyy-MM-dd.HH和yyyy-MM-dd
配置文件示例:
@ -43,4 +45,6 @@ hetu.event.listener.listen.query.completion=true
hetu.event.listener.audit.file=/var/log/hetu/hetu-audit.log
hetu.event.listener.audit.filecount=1
hetu.event.listener.audit.limit=100000
hetu.auditlog.logoutput=/var/log/
hetu.auditlog.logconversionpattern=yyyy-MM-dd.HH
```

View File

@ -0,0 +1,24 @@
#扩展物理执行计划
本节介绍openLooKeng如何添加扩展物理执行计划。通过物理执行计划的扩展openLooKeng可以使用其他算子加速库来加速SQL语句的执行。
##配置
在配置文件`config.properties`增加如下配置:
``` properties
extension_execution_planner_enabled=true
extension_execution_planner_jar_path=file:///xxPath/omni-openLooKeng-adapter-1.6.1-SNAPSHOT.jar
extension_execution_planner_class_path=nova.hetu.olk.OmniLocalExecutionPlanner
```
上述属性说明如下:
- `extension_execution_planner_enabled`:是否开启扩展物理执行计划特性。
- `extension_execution_planner_jar_path`指定扩展jar包的文件路径。
- `extension_execution_planner_class_path`指定扩展jar包中执行计划生成类的包路径。
##使用
当运行openLooKeng时可在WebUI或Cli中通过如下命令控制扩展物理执行计划的开启:
```
set session extension_execution_planner_enabled=true/false
```

View File

@ -115,6 +115,20 @@
>
> 此属性是在JVM堆中为openLooKeng不跟踪的分配留作裕量/缓冲区的内存量。
### `query.suspend-query-enabled`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 系统资源不足时,临时挂起运行中的查询。
### `query.max-suspended-queries`
> - **类型:** `integer`
> - **默认值:** `10`
>
> 终止查询之前,查询挂起尝试的最大次数。仅当`query.suspend-query-enabled`设置为`true`时,此属性才生效。
## 溢出属性
### `experimental.spill-enabled`
@ -148,6 +162,24 @@
>
> 此配置属性可由`spill_window_operator`会话属性重写。
### `experimental.spill-build-for-outer-join-enabled`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 为右外连接和全外连接操作启用溢出功能。
>
> 此config属性可被`spill_build_for_outer_join_enabled`会话属性覆盖。
### `experimental.inner-join-spill-filter-enabled`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 启用基于布隆过滤器的构建侧溢出匹配,以进行探查侧溢出决策。
>
> 此config属性可被`inner_join_spill_filter_enabled`会话属性覆盖。
### `experimental.spill-reuse-tablescan`
> - **类型**`boolean`
@ -157,13 +189,13 @@
>
> 此配置属性可由`spill_reuse_tablescan`会话属性重写。
### experimental.spiller-spill-path`
### `experimental.spiller-spill-path`
> - **类型:** `string`
> - **无默认值。** 启用溢出时必须设置。
>
> 溢出内容写入的目录。该属性可以是一个逗号分隔的列表,以同时溢出到多个目录,这有助于利用系统中安装的多个驱动器。
>
> 当`experimental.spiller-spill-to-hdfs`为`true`时,`experimental.spiller-spill-path`必须只包含一个目录。
> 不建议溢出到系统驱动器上。最重要的是不要溢出到写入JVM日志的驱动器因为磁盘过度使用可能导致JVM长时间暂停从而导致查询失败。
### `experimental.spiller-max-used-space-threshold`
@ -208,7 +240,7 @@
>
> 用于在Reuse Exchange中缓存页面的内存限制。
### experimental.spill-compression-enabled`
### `experimental.spill-compression-enabled`
> - **类型:** `boolean`
> - **默认值:** `false`
@ -222,6 +254,69 @@
>
> 允许使用随机生成的密钥(每个溢出文件)来加密和解密溢出到磁盘的数据。
### `experimental.spill-direct-serde-enabled`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 允许将页面直接序列化/读取到流中或从流中序列化/读取页面。
### `experimental.spill-prefetch-read-pages`
> - **类型:** `integer`
> - **默认值:** `1`
>
> 设置从溢出文件读取时预取的页数。
### `experimental.spill-use-kryo-serialization`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 启用基于Kryo的序列化以溢出到磁盘而不使用默认的Java序列化器。
### `experimental.revocable-memory-selection-threshold`
> - **类型:** `data size`
> - **默认值:** `512 MB`
>
> 设置运算符可撤销内存的内存选择阈值,直接为准备撤销的剩余字节分配可撤销内存。
### `experimental.prioritize-larger-spilts-memory-revoke`
> - **类型:** `boolean`
> - **默认值:** `true`
>
> 启用对具有较大可撤销内存的Split进行优先级排序。
### `experimental.spill-non-blocking-orderby`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 开启按照运算符排序使用异步机制溢出。即使在溢出正在进行时也可以累积输入并在次要数据累积超过阈值或主溢出完成时启动次溢出。阈值的默认值是20MB到可用内存的5%之间的最小值。此属性必须与`experimental.spill-enabled`属性结合使用。
>
> 此config属性可被`spill_non_blocking_orderby`会话属性覆盖。
### `experimental.spiller-spill-to-hdfs`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 启用溢出到HDFS。当此属性设置为`true`时,必须设置`experimental.spiller-spill-profile`属性,并且`experimental.spiller-spill-path`必须仅包含单个路径。
### `experimental.spiller-spill-profile`
> - **类型:** `string`
> - **无默认值。** 启用溢出到HDFS时必须设置此属性。
>
>
> 此属性定义用于溢出的[filesystem](../develop/filesystem.md)配置文件。对应的配置文件必须存在于`etc/filesystem`中。例如,如果此属性设置为`experimental.spiller-spill-profile=spill-hdfs`,则必须在`etc/filesystem`中创建描述此文件系统的配置文件`spill-hdfs.properties`其中包含必要的信息包括身份验证类型、config和keytab如果适用详情请参见[filesystem](../develop/filesystem.md)。
>
> 当`experimental.spiller-spill-to-hdfs`设置为`true`时必须配置此属性。所有Coordinator和Worker的配置文件中必须包含此属性。指定的文件系统必须可由所有Worker访问并且Worker必须能够读取和写入指定文件系统中`experimental.spiller-spill-path`文件夹中指明的路径。
## 交换属性
在openLooKeng节点之间为查询的不同阶段交换数据。调整这些属性可有助于解决节点间通信问题或提高网络利用率。
@ -259,21 +354,124 @@
>
> 如果网络延迟较高,增大该值可以提高网络吞吐量。减小该值可以提高大型集群的查询性能,因为它减少了由于交换客户端缓冲区保存了较多任务(而不是保存较少任务中的较多数据)的响应而导致的倾斜。
### `exchange.max-error-duration`
> - **类型:** `duration`
> - **最小值:** `1m`
> - **默认值:** `7m`
>
> 交换错误最大缓冲时间,超过该时限则查询失败。
### `sink.max-buffer-size`
> - **类型:** `data size`
> - **默认值:** `32MB`
>
> 上游任务等待拉取任务数据的输出缓冲区大小。如果任务输出是经过哈希分区的,那么缓冲区将在所有分区的使用者之间共享。如果网络延迟较高或集群中有多个节点,增加此值可以提高在阶段之间传输的数据的网络吞吐量。
>等待上游任务拉取的任务数据的输出缓冲区大小。如果任务输出是哈希分区的,则缓冲区将在所有分区的消费者之间共享。如果网络延迟高或集群中有许多节点,则增加此值可以提高阶段之间传输数据的网络吞吐量。
## 故障恢复处理属性
### 失败重试策略
### `failure.recovery.retry.profile`
> - **类型:** `string`
> - **默认值:** `default`
>
> 此属性定义用于确定HTTP客户端上是否发生故障的故障检测配置文件。此属性的值`<profile-name>`必须对应`etc/failure-retry-policy/`路径中的`<profile-name>.properties`文件。如果没有此类配置文件可用并且未设置此属性则使用“default”配置文件。
> 例如,`failure.recovery.retry.profile="test"`要求`test.properties`文件存在于`etc/failure-retry-policy`路径中。
> `test.properties`文件必须包含指定的`failure.recovery.retry.type`。
### `failure.recovery.retry.type`
> - **类型:** `string`
> - **默认值:** `timeout`
>
> 此属性用来设置正在使用的故障检测机制。默认值是基于`timeout`的故障检测。
#### 基于`timeout`的故障检测
> 如果使用此机制HTTP客户端故障将在指定时间段内重试重试失败则被视为永久故障。
>
> 可以为此类故障检测定义`max.error.duration`属性。
#### 基于`max-retry`的故障检测
> 如果使用此机制HTTP客户端故障将在被视为永久故障之前重试指定次数。
> 可以为此类故障检测定义`max.retry.count`和`max.error.duration`属性。
> 在这种类型的故障检测中,在查询故障检测模块之前,会执行`max.retry.count`次重试。当故障检测器模块检测到远程节点发生故障时HTTP客户端将此故障视为永久故障。否则例如当远程工作节点处于活动状态但没有响应时在`max.error.duration`指定的时间段内重试,重试失败则被视为永久故障。
### `max.error.duration`
> - **类型:** `duration`
> - **默认值:** `300s`
>
> 被视为永久故障前,协调器等待解决任务间相关错误的最长时间。
### `max.retry.count`
> - **类型:** `integer`
> - **默认值:** `100`
>
> 协调器在向故障检测器模块查询远程节点状态之前,对失败任务执行的最大重试次数。
> 此属性指定查询失败检测模块之前的最小重试次数。因此,实际故障数量可能会因为集群大小和集群负载而略有不同。
> 此属性仅用于基于`max-retry`的故障检测配置文件。
> 最小值为100。
### 故障检测Gossip协议配置
### `failure-detection-protocol`
>- **类型:** `string`
>- **默认值:** `heartbeat`
>
>此属性定义正在使用的故障检测器的类型。默认配置为`heartbeat`故障检测器。
>在`config.properties`文件中,将此属性配置为`gossip`可以启用Gossip协议。
>集群中的所有节点(即协调器和工作节点)都应在其各自的`etc/config.properties`文件中指定此属性。
### `failure-detector.heartbeat-interval`
>- **类型:** `duration`
>- **默认值:** `500ms` 500毫秒
>
>集群中两个节点之间的消息散播间隔。
>在Gossip协议中两个工作节点间的消息散播频率高于协调器和一个工作节点间。
>在协调器的`config.properties`文件中,可以为此属性配置一个较大的值,例如`5s`5秒
>在工作节点中,可以使用默认值。
### `failure-detector.worker-gossip-probe-interval`
>- **类型:** `duration`
>- **默认值:** `5s`5秒
>
>Gossip协议使用监控任务与`heartbeat`故障检测器相同)来监控其他节点。
>此属性指定监控任务刷新间隔,以触发工作节点消息散播。
>仅可以为工作节点指定默认值以外的任何其他值。
>该属性的值必须大于`failure-detector.heartbeat-interval`的值。
### `failure-detector.coordinator-gossip-probe-interval`
>- **类型:** `duration`
>- **默认值:** `5s`5秒
>
>Gossip协议使用监控任务与heartbeat故障检测器相同来监控其他节点。
>此属性指定监控任务刷新间隔,以触发协调器参与工作节点消息散播。
>仅可以为协调器指定默认值以外的任何其他值。
>该属性的值必须大于`failure-detector.heartbeat-interval`和`failure-detector.worker-gossip-probe-interval`的值。
### `failure-detector.coordinator-gossip-collate-interval`
>- **类型:** `duration`
>- **默认值:** `2s`2秒
>
>此属性指定协调器整理从所有工作节点获得的所有散播消息的间隔。
>此属性只支持为协调器配置。
>该属性的值必须大于`failure-detector.heartbeat-interval`的值。
### `failure-detector.gossip-group-size`
>- **类型:** `integer`
>- **默认值:** `Integer.MAX_VALUE`
>
>此属性定义单个工作节点在集群中散播消息的工作节点数量。
>任何大于集群大小即工作节点数量的值都意味着all-to-all消息散播。
>要保持较低的网络开销针对大型集群请将此属性设置为一个较小的值例如100工作节点的集群设置为10
>每次刷新协调器上的工作节点监视任务时协调器都会定义工作节点URI列表其大小由`failure-detector.gossip-group-size`指定以触发worker-to-worker消息散播。
>
>
## 任务属性
### `task.concurrency`
@ -517,7 +715,7 @@
## 启发式索引属性
启发式索引是外部索引模块,可用于过滤连接器级别的行。 位图Bloom和MinMaxIndex是openLooKeng提供的索引列表。 到目前为止,位图索引支持使用ORC存储格式的表支持蜂巢连接器
启发式索引是外部索引模块,可用于过滤连接器级别的行。 位图Bloom和MinMaxIndex是openLooKeng提供的索引列表。 到目前为止,位图索引支持Hive连接器的ORC存储格式的表
### `hetu.heuristicindex.filter.enabled`
@ -618,7 +816,7 @@
### `hetu.split-cache-map.enabled`
> - **类型:**`boolean`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 此属性启用分片缓存功能。 如果启用了状态存储,则分片缓存映射配置也会自动复制到状态存储中。 在具有多个协调器的HA设置的情况下状态存储用于在协调器之间共享分片的缓存映射。
@ -634,7 +832,7 @@
> 自动清空使系统能够通过持续监测需要清空的表来自动管理清空作业,以保持最佳性能。引擎从符合清空条件的数据源获取表,并触发对这些表的清空操作。
### `auto-vacuum.enabled:`
### `auto-vacuum.enabled`
> - **类型:** `boolean`
> - **默认值:** `false`
@ -661,7 +859,7 @@
>
> **注意:** 此属性只能在协调节点中配置。
## **CTE属性**
## CTE属性
### `cte.cte-max-queue-size`
@ -700,18 +898,25 @@
>
> 远程任务错误最大缓冲时间,超过该时限则查询失败。
## 分布式快照
## 查询恢复
### `recovery_enabled`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 此会话属性用于启用或禁用恢复框架,该框架在发生故障时启用或禁用查询重启/恢复。
### `snapshot_enabled`
> - 类型:`boolean`
> - **默认值**`false`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 此会话属性用于启用或禁用分布式快照功能。
> 启用恢复框架时,启用此会话属性可以在查询执行期间捕获快照。如果未启用恢复框架,则此属性不生效
### `hetu.experimental.snapshot.profile`
> - 类型:`string`
> - **类型:**`string`
>
> 此属性定义用于存储快照的[文件系统](../develop/filesystem.md)配置文件。对应的配置文件必须存在于`etc/filesystem`中。例如,如果将该属性设置为`hetu.experimental.snapshot.profile=snapshot-hdfs1`,则必须在`etc/filesystem`中创建描述此文件系统的配置文件`snapshot-hdfs1.properties`,其中包含的必要信息包括身份验证类型、配置和密钥表(如适用)。具体细节请参考[文件系统](../develop/filesystem.md)相关章节。
>
@ -719,20 +924,58 @@
>
> 作为实验性属性,或可以将快照存储在非文件系统位置,如连接器。
### `hetu.snapshot.maxRetries`
### `hetu.recovery.maxRetries`
> - 类型:`int`
> - **默认值**`10`
> - **类型:** `integer`
> - **默认值** `10`
>
> 此属性定义查询错误恢复尝试的最大次数。达到限制时,查询失败。
> 此属性定义查询错误恢复尝试的最大次数。达到限制时,查询失败。
>
> 也可以使用`snapshot_max_retries`会话属性在每个查询基础上指定
> 也可以使用`recovery_max_retries`会话属性为每个查询指定此属性
### `hetu.snapshot.retryTimeout`
### `hetu.recovery.retryTimeout`
> - 类型:`duration`
> - **默认值:**`10m`10分钟
> - **类型:** `duration`
> - **默认值:** `10m`10分钟
>
> 此属性定义系统等待所有任务成功恢复的最大时长。如果在此超时时限内任何任务未就绪,则认为恢复失败,查询将尝试从较早快照恢复(如果可用)。
> 此属性定义系统等待所有任务成功恢复的最长时间。如果在此时间内有任何任务未就绪,则恢复尝试将被视为失败,查询将尝试从较早的快照恢复(如果可用)。
>
> 也可以使用`snapshot_retry_timeout`会话属性在每个查询基础上指定。
> 也可以使用`recovery_retry_timeout`会话属性为每个查询指定此属性。
### `hetu.snapshot.useKryoSerialization`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 为快照启用基于Kryo的序列化而不是默认的Java序列化。
## HTTP客户端属性配置
### `http.client.idle-timeout`
> - **类型:** `duration`
> - **默认值:** `30s` 30秒
>
> 此参数定义了当http客户端没有任何操作时其保持连接的时间。
> 当超过指定时间还没有任何操作的话,将关闭客户端并释放相关资源。
>
> (注意:建议在高负载环境下,该参数配置大一点。)
### `http.client.request-timeout`
> - **类型:** `duration`
> - **默认值:** `10s` 10秒
>
> 此参数定义了http客户端接收响应的时间阈值。
> 当超过所配置时间,客户端没有接收到任何响应,则视为客户端的请求提交失败。
>
> (注意: 建议在高负载环境下,该参数配置大一点。)
## 连接器属性配置
### `case-insensitive-name-matching`
> - **类型:** `boolean`
> - **默认值:** `false`
>
> 不区分大小写匹配数据库和集合名称,默认区分大小写。

View File

@ -10,9 +10,9 @@
自版本1.2.0起openLooKeng支持恢复任务和工作节点故障。
## 启用分布式快照
分布式快照适用于长时间运行的查询任务。该功能默认为禁用状态,可以通过会话属性[`snapshot_enabled`](properties.md#snapshot_enabled)启用或禁用。建议仅在对可靠性要求高的复杂查询场景下启用该功能。
## 启用恢复框架
恢复框架对于长时间运行的查询最有用。默认禁用,可以使用会话属性[`recovery_enabled`](properties.md#recovery_enabled)启用和禁用恢复框架。建议仅对可靠性要求高的复杂查询启用该功能。
## 要求
@ -35,7 +35,7 @@
## 检测
协调节点与远程任务之间的通信长时间失败时,将触发错误恢复,由[`query.remote-task.max-error-duration`](properties.md#queryremote-taskmax-error-duration)配置控制。
当协调器与远程任务之间的通信长时间失败时,将触发错误恢复,由[`故障恢复处理属性`](properties.md#故障恢复处理属性)配置控制。
## 存储注意事项
@ -53,8 +53,20 @@
从错误和快照中恢复需要成本。捕获快照需要时间,时间长短取决于复杂性。因此,需要在性能和可靠性之间进行权衡。
建议仅在必要时启用分布式快照,如运行时间较长的查询任务。对于这些类型的工作负载,捕获快照的开销可以忽略不计。
建议在必要时打开快照捕获,例如对于长时间运行的查询。对于这些类型的工作负载,拍摄快照的开销可以忽略不计。
## 快照统计信息
在调试模式下启动CLI时快照捕获信息和恢复信息将与查询结果一起显示在CLI中。
快照捕获统计信息包括捕获的快照数量、捕获的快照大小、捕获快照所需的CPU时间和在查询期间捕获快照所需的挂钟时间。所有快照和最后一个快照的统计信息会分别显示。
快照恢复信息包括查询期间从快照恢复的次数、加载用于恢复的快照大小、从快照恢复所需的CPU时间和从快照恢复所需的挂钟时间。仅当查询期间发生恢复时才会显示恢复信息。
此外在查询正在进行时将显示捕获的快照数量和恢复的快照的ID。更多详细信息见下图。
![](../images/snapshot_statistics_cn.png)
## 配置
与分布式快照功能相关的配置可参见[属性参考](properties.md#分布式快照)。
恢复框架功能相关的配置,请参见[属性参考](properties.md#查询恢复)。

View File

@ -31,6 +31,10 @@
openLooKeng将溢出路径视为独立的磁盘参见[JBOD](https://en.wikipedia.org/wiki/Non-RAID_drive_architectures#JBOD )因此无需使用RAID进行溢出。
## 溢出到HDFS
操作可以直接溢出到HDFS。将`experimental.spiller-spill-to-hdfs`设置为`true`,配置`experimental.spiller-spill-profile`,并且`spiller-spill-path`必须仅包含一个目录。(更多详情请参见`experimental.spiller-spill-to-hdfs`和`experimental.spiller-spill-profile`属性)
## 溢出压缩
当启用溢出压缩(`tuning-spilling`中的`spill-compression-enabled`属性溢出页将被压缩后再写入磁盘。启用此特性可以减少磁盘I/O但会牺牲额外的CPU负载来压缩和解压缩溢出页。
@ -53,6 +57,8 @@ openLooKeng将溢出路径视为独立的磁盘参见[JBOD](https://en.wikipe
通过这种机制,联接操作符使用的峰值内存可以降低到最大构建表分区的大小。假设没有数据倾斜,这个值将是整个构建表大小的`1 / task.concurrency`倍。
注意spill-to-disk不支持交叉连接。
### 聚合
聚合函数对一组值执行操作并返回一个值。如果要聚合的组数量很大,可能需要大量内存。当启用溢出到磁盘时,如果没有足够的内存,则中间累积的聚合结果将写入磁盘。结果被重新加载回来,并以较低的内存占用量合并。
@ -60,6 +66,7 @@ openLooKeng将溢出路径视为独立的磁盘参见[JBOD](https://en.wikipe
### 排序
如果尝试对大量数据进行排序,可能需要大量内存。当启用为排序溢出到磁盘时,如果内存不足,则中间排序结果将写入磁盘。结果被重新加载回来,并以较低的内存占用量合并。
通常,当溢出正在进行时,运算符将被阻止接受输入,但当`experimental.spill-non-blocking-orderby`设置为`true`时,使用异步机制溢出(请参见`experimental.spill-non-blocking-orderby`)。
### 开窗函数

View File

@ -31,3 +31,54 @@ openLooKeng提供了一个用于监视和管理查询的Web界面。Web界面可
> - **默认值:** `false`
>
> 默认情况下基于HTTP的非安全环境禁用WEB UI。可以通过配置`etc/config.properties`文件的`hetu.queryeditor-ui.allow-insecure-over-http`属性启用(例子: hetu.queryeditor-ui.allow-insecure-over-http=true)。
### `hetu.queryeditor-ui.execution-timeout`
> - **类型:** `duration`
> - **默认值:** `100 DAYS`
>
> UI执行超时默认设置为100天。可以通过配置`etc/config.properties`文件中的`hetu.queryeditor-ui.execution-timeout`属性修改。
### `hetu.queryeditor-ui.max-result-count`
> - **类型:** `int`
> - **默认值:** `1000`
>
> UI最大结果计数默认设置为1000。可以通过配置`etc/config.properties`文件中的`hetu.queryeditor-ui.max-result-count`属性修改。
### `hetu.queryeditor-ui.max-result-size-mb`
>- **类型:** `size`
>- **默认值:** `1GB`
>
>UI最大结果大小默认设置为1 GB。可以通过配置`etc/config.properties`文件中的`hetu.queryeditor-ui.max-result-size-mb`属性修改。
### `hetu.queryeditor-ui.session-timeout`
> - **类型:** `duration`
> - **默认值:** `1 DAYS`
>
> UI会话超时默认设置为1天。可以通过配置`etc/config.properties`文件中的`hetu.queryeditor-ui.session-timeout`属性修改。
### `hetu.queryhistory.max-count`
> - **Type:** `int`
> - **Default value:** `1000`
>
> openLooKeng储存的历史查询记录最大数量。可以通过配置`etc/config.properties`文件的`hetu.queryhistory.max-count`属性修改。
### `hetu.collectionsql.max-count`
> - **Type:** `int`
> - **Default value:** `100`
>
> 每位用户收藏sql语句条数上限.可以通过配置`etc/config.properties`文件的`hetu.collectionsql.max-count`属性修改。
## 备注
收藏sql语句的最大长度默认为600可通过如下步骤对其进行修改
1. 根据hetu-metastore.properties文件中jdbc配置信息登录mysql数据库。
2. 选中hetu_favorite表使用命令`alter table hetu_favorite modify query varchar(2000) not null;`修改收藏语句最大长度为2000。

View File

@ -1,5 +1,8 @@
# Hudi连接器
### 版本说明
目前Hudi只支持0.7.0版本。
### Hudi介绍
Apache Hudi是一个快速迭代的数据湖存储系统可以帮助企业构建和管理PB级数据湖。它提供在DFS上存储超大规模数据集同时使得流式处理如果批处理一样该实现主要是通过如下两个原语实现。

Some files were not shown because too many files have changed in this diff Show More