Common Scenarios to avoid with DataWarehousing

Database Design

Rule	Description	Value	Source	Problem Description
1	Excessive sorting and RID lookup operations should be reduced with covered indexes.		Sys.dm_exec_sql_text Sys.dm_exec_cached_plans	Large data warehouse can benefit from more indexes. Indexes can be used to cover queries and avoid sorting. The cost of index overhead is only paid when data is loaded.
2	Excessive fragmentation: Average fragmentation_in_percent should be <25%	>25%	sys.dm_db _index_physical_stats	Reducing index fragmentation through index rebuilds can benefit big range scans, common in data warehouse and Reporting scenarios.
3	Scans and ranges are common. Look for missing indexes	>= 1	Perfmon object SQL Server Access Methods Sys.dm_db_missing_index_group_stats Sys.dm_db_missing_index_groups Sys.dm_db_missing_index_details	A missing index flushes the cache.
4	Unused Indexes should be avoided	If an index is NEVER used, it will not appear in the DMV sys.dm_db_index_usage_stats		Index maintenance for unused indexes should be avoided.

Resource issue: CPU

Rule	Description	Value	Source	Problem Description
1	Signal Waits	> 25%	Sys.dm_os_wait_stats	Time in runnable queue is pure CPU wait.
2	Avoid plan reuse	> 25%	Perfmon object SQL Server Statistics	Data warehouse has fewer transactions than OLTP, each with significantly bigger IO. Therefore, having the correct plan is more important than reusing a plan. Unlike OLTP, data warehouse queries are not identical.
3	Parallelism: Cxpacket waits	<10%	Sys.dm_os_wait_stats	Parallelism is desirable in data warehouse or reporting workloads.

Resource issue: Memory

Rule

Description

Value

Source

Problem Description

Memory grants pending

Perfmon object

SQL Server Memory Manager

Memory grant not available for query to run. Check for

Sufficient memory and page life expectancy.

Page life expectancy

Drops by 50%

Perfmon object

SQL Server Buffer Manager

Page life expectancy is the average number of seconds a data page stays in cache. Low values could indicate a cache flush that is caused by a big read.

Look for possible missing index.

Resource issue: IO

Rule	Description	Value	Source	Problem Description
1	Average Disk sec/read	>20 ms	Perfmon object Physical Disk	Reads should take 4-8ms without any IO pressure.
2	Average Disk sec/write	>20 ms	Perfmon object Physical Disk	Writes (sequential) can be as fast as 1 ms for transaction log.
3	Big scans	>1	Perfmon object SQL Server Access Methods	A missing index flushes the cache.
4	If Top 2 values for wait stats are any of the following: ASYNCH_IO_COMPLETION IO_COMPLETION LOGMGR WRITELOG PAGEIOLATCH_x	Top 2	Sys.dm_os_wait_stats	If top 2 wait_stats values include IO, there is an IO bottleneck

Resource issue: Blocking

Rule	Description	Value	Source	Problem Description
1	Block percentage	>2%	Sys.dm_db_index_operational_stats	Frequency of blocks.
2	Block process report	30 sec	Sp_configure, profiler	Report of statements.
3	Average Row Lock Waits	>100ms	Sys.dm_db_index_operational_stats	Duration of blocks.
4	If Top 2 values for wait stats are any of the following: LCK_M_BU LCK_M_IS LCK_M_IU LCK_M_IX LCK_M_RIn_NL LCK_M_RIn_S LCK_M_RIn_U LCK_M_RIn_X LCK_M_RS_S LCK_M_RS_U LCK_M_RX_S LCK_M_RX_U LCK_M_RX_X LCK_M_S LCK_M_SCH_M LCK_M_SCH_S LCK_M_SIU LCK_M_SIX LCK_M_U LCK_M_UIX LCK_M_X	Top 2	Sys.dm_os_wait_stats	If top 2 wait_stats values include IO, there is a blocking bottleneck. Consider using row versioning to minimize shared locking blocks.

Exactly the opposite of OLTP applications, reporting or relational data warehouse applications are characterized by small numbers of (different) big transactions. These are frequently SELECT intensive operations. The implications are significant for database design, resource usage, and system performance.

Reporting and data warehouse performance objectives are as follows:

Data warehouse and relational data warehouse designs can have more indexes as the cost of index maintenance is paid only one time, during the batch update process.
Plan reuse should generally be avoided. Plan reuse may result in picking up a plan that was good for some other query (with different data distribution), but may not be good for this query. The time taken for plan generation of a large DataWarehouse query is not nearly as important as having the right plan.
Sorts can and should be minimized with correct index usage.
Missing index situations should be investigated and corrected.
Large IOs such as range scans benefits from on disk contiguity. Index fragmentation should be frequently monitored and kept to a minimum with index rebuilds.
Blocking is generally uncommon as most data warehouse transactions are read operations.
Parallelism is generally desirable for data warehouse applications.

Common Scenarios to avoid with DataWarehousing的更多相关文章

Common scenarios to avoid in OLTP
Database Design Rule Description Value Source Problem Description 1 High Frequency queries having a ...
8 Mistakes to Avoid while Using RxSwift. Part 1
Part 1: not disposing a subscription Judging by the number of talks, articles and discussions relate ...
Android Lint Checks
Android Lint Checks Here are the current list of checks that lint performs as of Android Studio 2.3 ...
(WPF) 基本题
What is WPF? WPF (Windows Presentation foundation) is a graphical subsystem for displaying user inte ...
Processing Images
https://developer.apple.com/library/content/documentation/GraphicsImaging/Conceptual/CoreImaging/ci_ ...
IMS Global Learning Tools Interoperability™ Implementation Guide
Final Version 1.1 Date Issued: 13 March 2012 Latest version: http://www.imsglobal ...
9.Parameters
1.Optional and Named Parameters calls these methods can optionally not specify some of the arguments ...
C# Development 13 Things Every C# Developer Should Know
https://dzone.com/refcardz/csharp C#Development 13 Things Every C# Developer Should Know Written by ...
Introducing Microsoft Sync Framework: Sync Services for File Systems
https://msdn.microsoft.com/en-us/sync/bb887623 Introduction to Microsoft Sync Framework File Synchro ...

随机推荐

【转】Java八种基本数据类型的比较及其相互转化
java中有且仅有八种基本数据类型,记住就行,共分为四类: 第一类:整型-->byte short int long 第二类:浮点-->float doub ...
JS驗證兩位小數
function SizeCheck(Textdiv) { var fg = true; str = $("#" + T ...
mysql批量执行sql文件
1.待执行的sql文件为1.sql.2.sql.3.sql.4.sql等 2.写一个batch.sql文件: source .sql; source .sql; source .sql; source ...
delegate事件绑定
为了代码的健壮性,绑定事件之前先解绑再进行绑定. var _$div = $("#id");_$div.undelegate("click mouseover mouse ...
JavaWEB域对象
PageContext: ServletRequest: HttpSession: ServletContext: void setAttribute(String name, Object valu ...
底部tab的返回退出和对话框
第一种: private long exitTime = 0; @Override public boolean dispatchKeyEvent(KeyEvent event) { if (even ...
Windows下MongoDB环境搭建
MongoDB下载登录MongoDB官网:www.mongodb.org:点击[Download MongoDB]按钮,进入如下所示界面选择目标操作系统及其版本,比如这里选择的是64位的Windo ...
Java核心知识点学习----多线程并发之线程间的通信,notify,wait
1.需求: 子线程循环10次,主线程循环100次,这样间隔循环50次. 2.实现: package com.amos.concurrent; /** * @ClassName: ThreadSynch ...
oracle 驱动安装备忘
ubuntu 从oracle官网下载两个必须的rpm包(这里选择的是version12.1.0.2.0, 64位操作系统) oracle-instantclient12.1-basic-12.1.0. ...
在sqlserver存储过程中给in参数传带逗号值的办法，如传'1','2','3'这样的
最近在一项目修改中,要在存储过程中给in参数传值,语句写的也对,但怎么执行都得不出结果,如果把这语句直接赋值.执行,却能得出结果,很是奇怪,如: 直接执行select schoolname from ...

Common Scenarios to avoid with DataWarehousing

Common Scenarios to avoid with DataWarehousing的更多相关文章

随机推荐

热门专题