【StoneDB子查询优化】subquery子查询-exists子查询的剔除遍历处理
来源:SegmentFault
2023-02-16 15:24:57
0浏览
收藏
IT行业相对于一般传统行业,发展更新速度更快,一旦停止了学习,很快就会被行业所淘汰。所以我们需要踏踏实实的不断学习,精进自己的技术,尤其是初学者。今天golang学习网给大家整理了《【StoneDB子查询优化】subquery子查询-exists子查询的剔除遍历处理》,聊聊MySQL、数据库,我们一起来看看吧!
摘要:
记录对exists子句进行剔除遍历的处理, 对比优化前后子查询耗时
执行的SQL语句:
/stonedb57/install/bin/mysql -D tpch -e " explain select o_orderpriority, count(*) as order_count from orders where o_orderdate >= date '1993-07-01' and o_orderdate
核心函数处理:
ParameterizedFilter::UpdateMultiIndex
void ParameterizedFilter::UpdateMultiIndex(bool count_only, int64_t limit) {
MEASURE_FET("ParameterizedFilter::UpdateMultiIndex(...)");
thd_proc_info(mind->ConnInfo().Thd(), "update multi-index");
if (descriptors.Size() ClearLocalDescFilters();
return;
}
SyntacticalDescriptorListPreprocessing();
bool empty_cannot_grow = true; // if false (e.g. outer joins), then do not
// optimize empty multiindex as empty result
for (uint i = 0; i m_conn->GetThreadID()) m_conn->GetThreadID())
m_conn->GetThreadID())
ZeroTuples() && empty_cannot_grow) {
mind->Empty();
PrepareRoughMultiIndex();
rough_mind->ClearLocalDescFilters();
return;
}
// Prepare execution - rough set part
for (uint i = 0; i Empty();
PrepareRoughMultiIndex();
rough_mind->ClearLocalDescFilters();
return;
} else {
DimensionVector dims(mind->NoDimensions());
descriptors[i].attr.vc->MarkUsedDims(dims);
mind->MakeCountOnly(0, dims);
}
}
}
}
}
if (rough_mind) rough_mind->ClearLocalDescFilters();
PrepareRoughMultiIndex();
nonempty = RoughUpdateMultiIndex(); // calculate all rough conditions,
if ((!nonempty && empty_cannot_grow) || mind->m_conn->Explain()) {
mind->Empty(); // nonempty==false if the whole result is empty (outer joins
// considered)
rough_mind->ClearLocalDescFilters();
return;
}
PropagateRoughToMind(); // exclude common::RSValue::RS_NONE from mind
// count other types of conditions, e.g. joins (i.e. conditions using
// attributes from two
// dimensions)
int no_of_join_conditions = 0; // count also one-dimensional outer join conditions
int no_of_delayed_conditions = 0;
for (uint i = 0; i 1) {
STONEDB_LOG(LogCtl_Level::INFO, "UpdateMultiIndex: descriptorsNum : %d",
descriptorsNum);
}
int desc_no = 0;
for (uint i = 0; i GetDim();
}
if (last_desc_dim != -1 && cur_dim != -1 && last_desc_dim != cur_dim) {
// Make all possible projections to other dimensions
RoughMakeProjections(cur_dim, false);
}
++ApplyDescriptorNum;
// limit should be applied only for the last descriptor
ApplyDescriptor(i, (desc_no != no_desc || no_of_delayed_conditions > 0 || no_of_join_conditions) ? -1 : limit);
if (!descriptors[i].attr.vc) continue; // probably desc got simplified and is true or false
if (cur_dim >= 0 && mind->GetFilter(cur_dim) && mind->GetFilter(cur_dim)->IsEmpty() && empty_cannot_grow) {
mind->Empty();
if (rccontrol.isOn()) {
rccontrol.lock(mind->m_conn->GetThreadID())
ClearLocalDescFilters();
return;
}
last_desc_dim = cur_dim;
}
}
if (ApplyDescriptorNum > 1) {
auto diff =
std::chrono::duration_cast<:chrono::duration>>(std::chrono::high_resolution_clock::now() - start);
STONEDB_LOG(LogCtl_Level::INFO, "Timer %f : UpdateMultiIndex: ApplyDescriptorNum : %d", diff.count(),
ApplyDescriptorNum);
}
rough_mind->UpdateReducedDimension();
mind->UpdateNoTuples();
for (int i = 0; i NoDimensions(); i++)
if (mind->GetFilter(i))
table->SetVCDistinctVals(i,
mind->GetFilter(i)->NoOnes()); // distinct values - not more than the
// number of rows after WHERE
rough_mind->ClearLocalDescFilters();
// Some displays
if (rccontrol.isOn()) {
int pack_full = 0, pack_some = 0, pack_all = 0;
rccontrol.lock(mind->m_conn->GetThreadID())
NoDimensions(); i++)
if (mind->GetFilter(i)) {
Filter *f = mind->GetFilter(i);
pack_full = 0;
pack_some = 0;
pack_all = (int)((mind->OrigSize(i) + ((1 NoPower()) - 1)) >> mind->NoPower());
for (int b = 0; b IsFull(b))
pack_full++;
else if (!f->IsEmpty(b))
pack_some++;
}
rccontrol.lock(mind->m_conn->GetThreadID())
m_conn->GetThreadID()) m_conn->GetThreadID())
m_conn->GetThreadID())
ZeroTuples() && empty_cannot_grow) {
// set all following descriptors to done
for (uint j = i; j GetDim()) &&
!descriptors[i].attr.vc->IsNullsPossible()) {
for (int j = 0; j do not materialize multiindex
// (just counts tuples). WARNING: in this case cannot use multiindex for
// any operations other than NoTuples().
if (count_only) join_tips.count_only = true;
join_tips.limit = limit;
// only one dim used in distinct context?
int distinct_dim = table->DimInDistinctContext();
int dims_in_output = 0;
for (int dim = 0; dim NoDimensions(); dim++)
if (mind->IsUsedInOutput(dim)) dims_in_output++;
if (distinct_dim != -1 && dims_in_output == 1) join_tips.distinct_only[distinct_dim] = true;
}
// Optimization: Check whether all dimensions are really used
DimensionVector dims_used(mind->NoDimensions());
for (uint jj = 0; jj NoDimensions(); dim++)
if (!mind->IsUsedInOutput(dim) && dims_used[dim] == false) join_tips.forget_now[dim] = true;
}
// Joining itself
UpdateJoinCondition(join_desc, join_tips);
}
}
// Execute all delayed conditions
for (uint i = 0; i m_conn->GetThreadID()) MakeDimensionSuspect(); // no common::RSValue::RS_ALL packs
mind->UpdateNoTuples();
}
导致exists子查询遍历处理的逻辑:
for (uint i = 0; i GetDim();
}
if (last_desc_dim != -1 && cur_dim != -1 && last_desc_dim != cur_dim) {
// Make all possible projections to other dimensions
RoughMakeProjections(cur_dim, false);
}
++ApplyDescriptorNum;
// limit should be applied only for the last descriptor
ApplyDescriptor(i, (desc_no != no_desc || no_of_delayed_conditions > 0 || no_of_join_conditions) ? -1 : limit);
if (!descriptors[i].attr.vc) continue; // probably desc got simplified and is true or false
if (cur_dim >= 0 && mind->GetFilter(cur_dim) && mind->GetFilter(cur_dim)->IsEmpty() && empty_cannot_grow) {
mind->Empty();
if (rccontrol.isOn()) {
rccontrol.lock(mind->m_conn->GetThreadID())
ClearLocalDescFilters();
return;
}
last_desc_dim = cur_dim;
}
}优化策略:
思路:
- exists子句与join都涉及到内表与外表, 本质上可以理解为要处理同样的数据量
- mysql可以将in查询转换为exists, 两者在语义上可以对等 https://www.jb51.net/article/236338.htm#_label3_2_0_1
- 在出现join判断的地方,将exists条件做相同对待
优化后的遍历核心代码:
for (uint i = 0; i GetDim();
}
if (last_desc_dim != -1 && cur_dim != -1 && last_desc_dim != cur_dim) {
// Make all possible projections to other dimensions
RoughMakeProjections(cur_dim, false);
}
++ApplyDescriptorNum;
// limit should be applied only for the last descriptor
ApplyDescriptor(i, (desc_no != no_desc || no_of_delayed_conditions > 0 || no_of_join_conditions) ? -1 : limit);
if (!descriptors[i].attr.vc) continue; // probably desc got simplified and is true or false
if (cur_dim >= 0 && mind->GetFilter(cur_dim) && mind->GetFilter(cur_dim)->IsEmpty() && empty_cannot_grow) {
mind->Empty();
if (rccontrol.isOn()) {
rccontrol.lock(mind->m_conn->GetThreadID())
ClearLocalDescFilters();
return;
}
last_desc_dim = cur_dim;
}
}整个函数:
void ParameterizedFilter::UpdateMultiIndex(bool count_only, int64_t limit) {
MEASURE_FET("ParameterizedFilter::UpdateMultiIndex(...)");
thd_proc_info(mind->ConnInfo().Thd(), "update multi-index");
if (descriptors.Size() ClearLocalDescFilters();
return;
}
SyntacticalDescriptorListPreprocessing();
bool empty_cannot_grow = true; // if false (e.g. outer joins), then do not
// optimize empty multiindex as empty result
for (uint i = 0; i m_conn->GetThreadID()) m_conn->GetThreadID())
m_conn->GetThreadID())
ZeroTuples() && empty_cannot_grow) {
mind->Empty();
PrepareRoughMultiIndex();
rough_mind->ClearLocalDescFilters();
return;
}
// Prepare execution - rough set part
for (uint i = 0; i Empty();
PrepareRoughMultiIndex();
rough_mind->ClearLocalDescFilters();
return;
} else {
DimensionVector dims(mind->NoDimensions());
descriptors[i].attr.vc->MarkUsedDims(dims);
mind->MakeCountOnly(0, dims);
}
}
}
}
}
if (rough_mind) rough_mind->ClearLocalDescFilters();
PrepareRoughMultiIndex();
nonempty = RoughUpdateMultiIndex(); // calculate all rough conditions,
if ((!nonempty && empty_cannot_grow) || mind->m_conn->Explain()) {
mind->Empty(); // nonempty==false if the whole result is empty (outer joins
// considered)
rough_mind->ClearLocalDescFilters();
return;
}
PropagateRoughToMind(); // exclude common::RSValue::RS_NONE from mind
// count other types of conditions, e.g. joins (i.e. conditions using
// attributes from two
// dimensions)
int no_of_join_conditions = 0; // count also one-dimensional outer join conditions
int no_of_delayed_conditions = 0;
for (uint i = 0; i 1) {
STONEDB_LOG(LogCtl_Level::INFO, "UpdateMultiIndex: descriptorsNum : %d",
descriptorsNum);
}
int desc_no = 0;
for (uint i = 0; i GetDim();
}
if (last_desc_dim != -1 && cur_dim != -1 && last_desc_dim != cur_dim) {
// Make all possible projections to other dimensions
RoughMakeProjections(cur_dim, false);
}
++ApplyDescriptorNum;
// limit should be applied only for the last descriptor
ApplyDescriptor(i, (desc_no != no_desc || no_of_delayed_conditions > 0 || no_of_join_conditions) ? -1 : limit);
if (!descriptors[i].attr.vc) continue; // probably desc got simplified and is true or false
if (cur_dim >= 0 && mind->GetFilter(cur_dim) && mind->GetFilter(cur_dim)->IsEmpty() && empty_cannot_grow) {
mind->Empty();
if (rccontrol.isOn()) {
rccontrol.lock(mind->m_conn->GetThreadID())
ClearLocalDescFilters();
return;
}
last_desc_dim = cur_dim;
}
}
if (ApplyDescriptorNum > 1) {
auto diff =
std::chrono::duration_cast<:chrono::duration>>(std::chrono::high_resolution_clock::now() - start);
STONEDB_LOG(LogCtl_Level::INFO, "Timer %f : UpdateMultiIndex: ApplyDescriptorNum : %d", diff.count(),
ApplyDescriptorNum);
}
rough_mind->UpdateReducedDimension();
mind->UpdateNoTuples();
for (int i = 0; i NoDimensions(); i++)
if (mind->GetFilter(i))
table->SetVCDistinctVals(i,
mind->GetFilter(i)->NoOnes()); // distinct values - not more than the
// number of rows after WHERE
rough_mind->ClearLocalDescFilters();
// Some displays
if (rccontrol.isOn()) {
int pack_full = 0, pack_some = 0, pack_all = 0;
rccontrol.lock(mind->m_conn->GetThreadID())
NoDimensions(); i++)
if (mind->GetFilter(i)) {
Filter *f = mind->GetFilter(i);
pack_full = 0;
pack_some = 0;
pack_all = (int)((mind->OrigSize(i) + ((1 NoPower()) - 1)) >> mind->NoPower());
for (int b = 0; b IsFull(b))
pack_full++;
else if (!f->IsEmpty(b))
pack_some++;
}
rccontrol.lock(mind->m_conn->GetThreadID())
m_conn->GetThreadID()) m_conn->GetThreadID())
m_conn->GetThreadID())
ZeroTuples() && empty_cannot_grow) {
// set all following descriptors to done
for (uint j = i; j GetDim()) &&
!descriptors[i].attr.vc->IsNullsPossible()) {
for (int j = 0; j do not materialize multiindex
// (just counts tuples). WARNING: in this case cannot use multiindex for
// any operations other than NoTuples().
if (count_only) join_tips.count_only = true;
join_tips.limit = limit;
// only one dim used in distinct context?
int distinct_dim = table->DimInDistinctContext();
int dims_in_output = 0;
for (int dim = 0; dim NoDimensions(); dim++)
if (mind->IsUsedInOutput(dim)) dims_in_output++;
if (distinct_dim != -1 && dims_in_output == 1) join_tips.distinct_only[distinct_dim] = true;
}
// Optimization: Check whether all dimensions are really used
DimensionVector dims_used(mind->NoDimensions());
for (uint jj = 0; jj NoDimensions(); dim++)
if (!mind->IsUsedInOutput(dim) && dims_used[dim] == false) join_tips.forget_now[dim] = true;
}
// Joining itself
UpdateJoinCondition(join_desc, join_tips);
}
}
测试数据:
优化exists判断后的子查询耗时:
mysql> select
-> o_orderpriority,
-> count(*) as order_count
-> from
-> orders
-> where
-> o_orderdate >= date '1993-07-01'
-> and o_orderdate and exists (
-> select
-> *
-> from
-> lineitem
-> where
-> l_orderkey = o_orderkey
-> and l_commitdate )
-> group by
-> o_orderpriority
-> order by
-> o_orderpriority ;
+-----------------+-------------+
| o_orderpriority | order_count |
+-----------------+-------------+
| 1-URGENT | 114839 |
| 2-HIGH | 114276 |
| 3-MEDIUM | 114716 |
| 4-NOT SPECIFIED | 114913 |
| 5-LOW | 114927 |
+-----------------+-------------+
5 rows in set (4.04 sec)
对比未优化前的子查询耗时:

优化后的火焰图:
现在时间损耗已经全部落在了解压缩上


优化exists判断后的处理逻辑:
(gdb) bt
#0 stonedb::core::TwoDimensionalJoiner::CreateJoiner (join_alg_type=stonedb::core::JTYPE_GENERAL, mind=..., tips=..., table=0x7fc2f89fb510)
at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/stonedb/core/joiner.cpp:79
#1 0x00000000030afe84 in stonedb::core::ParameterizedFilter::UpdateJoinCondition (this=0x7fc2f89fb660, cond=..., tips=...)
at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/stonedb/core/parameterized_filter.cpp:579
#2 0x00000000030b3aec in stonedb::core::ParameterizedFilter::UpdateMultiIndex (this=0x7fc2f89fb660, count_only=false, limit=-1)
at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/stonedb/core/parameterized_filter.cpp:1203
#3 0x0000000002d73237 in stonedb::core::Query::Preexecute (this=0x7fc70415e800, qu=..., sender=0x7fc2f89d05f0, display_now=true)
at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/stonedb/core/query.cpp:777
#4 0x0000000002d44a6e in stonedb::core::Engine::Execute (this=0x75a5df0, thd=0x7fc2f8000b70, lex=0x7fc2f8002e98, result_output=0x7fc2f8010320, unit_for_union=0x0)
at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/stonedb/core/engine_execute.cpp:421
#5 0x0000000002d43d22 in stonedb::core::Engine::HandleSelect (this=0x75a5df0, thd=0x7fc2f8000b70, lex=0x7fc2f8002e98, result=@0x7fc70415ed18: 0x7fc2f8010320, setup_tables_done_option=0,
res=@0x7fc70415ed14: 0, optimize_after_sdb=@0x7fc70415ed0c: 1, sdb_free_join=@0x7fc70415ed10: 1, with_insert=0)
at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/stonedb/core/engine_execute.cpp:232
#6 0x0000000002e2c501 in stonedb::dbhandler::SDB_HandleSelect (thd=0x7fc2f8000b70, lex=0x7fc2f8002e98, result=@0x7fc70415ed18: 0x7fc2f8010320, setup_tables_done_option=0,
res=@0x7fc70415ed14: 0, optimize_after_sdb=@0x7fc70415ed0c: 1, sdb_free_join=@0x7fc70415ed10: 1, with_insert=0)
at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/stonedb/handler/ha_rcengine.cpp:82
#7 0x000000000246fe9a in execute_sqlcom_select (thd=0x7fc2f8000b70, all_tables=0x7fc2f801b530) at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/sql/sql_parse.cc:5182
#8 0x000000000246921e in mysql_execute_command (thd=0x7fc2f8000b70, first_level=true) at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/sql/sql_parse.cc:2831
#9 0x0000000002470e63 in mysql_parse (thd=0x7fc2f8000b70, parser_state=0x7fc70415feb0) at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/sql/sql_parse.cc:5621
#10 0x00000000024660fb in dispatch_command (thd=0x7fc2f8000b70, com_data=0x7fc704160650, command=COM_QUERY) at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/sql/sql_parse.cc:1495
#11 0x0000000002465027 in do_command (thd=0x7fc2f8000b70) at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/sql/sql_parse.cc:1034
#12 0x0000000002597c43 in handle_connection (arg=0x756c0b0) at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/sql/conn_handler/connection_handler_per_thread.cc:313
#13 0x0000000002c7b874 in pfs_spawn_thread (arg=0x758e640) at /data/jenkins/workspace/stonedb5.7-zsl-centos7.9/storage/perfschema/pfs.cc:2197
#14 0x00007fc70f86fea5 in start_thread (arg=0x7fc704161700) at pthread_create.c:307
#15 0x00007fc70dca6b0d in clone () at ../sysdeps/unix/sysv/linux/x86_64/clone.S:111

今天关于《【StoneDB子查询优化】subquery子查询-exists子查询的剔除遍历处理》的内容就介绍到这里了,是不是学起来一目了然!想要了解更多关于mysql的内容请关注golang学习网公众号!
版本声明
本文转载于:SegmentFault 如有侵犯,请联系study_golang@163.com删除
【StoneDB子查询优化】subquery子查询-优化exists场景的join查询优化器的处理-100GB数据量测试
- 上一篇
- 【StoneDB子查询优化】subquery子查询-优化exists场景的join查询优化器的处理-100GB数据量测试
- 下一篇
- 【StoneDB子查询优化】subquery子查询-内存拷贝分析及优化
查看更多
最新文章
-
- 数据库 · MySQL | 1星期前 | MySQL · 慢查询 · 索引优化 · COUNT查询 · 汇总表 · 联合索引 覆盖索引 汇总表 MySQL COUNT慢 COUNT(*)优化
- MySQL COUNT(*) 总数查询变慢怎么办:从扫描行数到汇总表的完整治理流程
- 329浏览 收藏
查看更多
课程推荐
-
- 前端进阶之JavaScript设计模式
- 设计模式是开发人员在软件开发过程中面临一般问题时的解决方案,代表了最佳的实践。本课程的主打内容包括JS常见设计模式以及具体应用场景,打造一站式知识长龙服务,适合有JS基础的同学学习。
- 543次学习
-
- GO语言核心编程课程
- 本课程采用真实案例,全面具体可落地,从理论到实践,一步一步将GO核心编程技术、编程思想、底层实现融会贯通,使学习者贴近时代脉搏,做IT互联网时代的弄潮儿。
- 516次学习
-
- 简单聊聊mysql8与网络通信
- 如有问题加微信:Le-studyg;在课程中,我们将首先介绍MySQL8的新特性,包括性能优化、安全增强、新数据类型等,帮助学生快速熟悉MySQL8的最新功能。接着,我们将深入解析MySQL的网络通信机制,包括协议、连接管理、数据传输等,让
- 500次学习
-
- JavaScript正则表达式基础与实战
- 在任何一门编程语言中,正则表达式,都是一项重要的知识,它提供了高效的字符串匹配与捕获机制,可以极大的简化程序设计。
- 487次学习
-
- 从零制作响应式网站—Grid布局
- 本系列教程将展示从零制作一个假想的网络科技公司官网,分为导航,轮播,关于我们,成功案例,服务流程,团队介绍,数据部分,公司动态,底部信息等内容区块。网站整体采用CSSGrid布局,支持响应式,有流畅过渡和展现动画。
- 485次学习
查看更多
AI推荐
-
- ljg-skills
- ljg-skills 是李继刚开源的 AI 技能与提示词集合,面向大模型使用者整理了一批可复用的 prompt、角色设定和任务技能模板,适合用于学习提示词设计、搭建个人 AI 工作流和沉淀团队常用智能体能力。
- 1661次使用
-
- MELO音乐
- MELO音乐是一站式AI视频与音乐制作助手,对标suno, udio的高品质体验。提供伴奏生成、原创写词、无损导出、哼唱识曲、混音变声等全套音频与短视频编辑工具。无论是流行Kpop、电音说唱、民谣古风、摇滚儿歌还是商用轻音乐,MELO为你免费谱曲,轻松做同款!
- 1609次使用
-
- UniScribe
- UniScribe 是一款 AI 音视频转文字与内容整理工具,支持上传音频、视频文件或粘贴 YouTube 链接,自动生成转写文本、摘要、思维导图和关键问题,并支持多格式导出,适合会议记录、课程学习、访谈整理和内容创作复盘。
- 1538次使用
-
- 剧云
- 剧云是专业中文剧本创作平台,安全稳定运行十余年,集成AI编剧、剧本医生审核、人物小传、剧情关系图、大纲编写、多人协作、Word导入导出、版权管控功能,数据安全防护,轻松高效创作剧本。
- 1736次使用
-
- 万象有声
- 万象有声,一个专为有声创作者打造的新一代智能有声内容创作平台。平台提供专业的智能拆章、智能画本编辑、AI配音、AI生成音效、后期制作、智能对轨、智能审听等有声创作全流程工具,可以帮助创作者高效、低成本创作出引人入胜的有声作品。立即体验,让有声书制作更简单!
- 1725次使用
查看更多
相关文章
-
- Linux系统下如何安装Mysql(centOS7以上不支持Mysql)
- 2023-01-16 100浏览
-
- 在windows上用docker desktop安装StoneDB
- 2023-01-20 100浏览
-
- 总结 mysql 一些小技巧
- 2023-01-21 100浏览
-
- MySQL如何给大表加索引
- 2023-01-26 100浏览
-
- 积分商城简要设计
- 2023-02-17 100浏览

