DATASTAGE QUESTIONS
1. WHAT IS DIFFERENCE BETWEEN SEQUENTAIL FILE AND DATA SET?
Sequentail file--- reads the data sequentially. cannot handle nullability. memory linit is 2gb. file extension is .txt/.csv
Dataset-----------reads the data parallel can handle nulls. extension of file is .ds
2.what is APT configuration file?
It is the environment variable which is used to recognize the *.apt file in DataStage. It is also used to keep the node information, scratch information and disk storage information.
3. what is diffenrece between datasatge 11.x version and previous?
DataStage 11.5 is the first version that promises to run anywhere - it can run on a dedicated DataStage engine (Windows, Linux, AIX or Z). It can push processing down into databases via Balanced Optimizer. It can run natively on Hadoop.
Tell any one of these:
Dynamic RDBMS Stage: replaced by DRS Connector
Oracle OCI, Oracle OCI Load: replaced by Oracle Connector
Teradata API: replaced by Teradata Connector
DB2 UDB Load, UBD API, DB2 Z: replaced by DB2 Connector
4. What is descriptor file and data file in dataset?
As the name says, data files contains the data and the descriptor file contains the information about the data in the data files.
5. How to remove duplicates in datastage?
Remove Duplicate Stage:
Duplicates can be detached by using Sort stage. We can use the opportunity, as allow duplicate = false.
6. Differnce between Join,Merge,Looup?
Join: linknames--Left,Right,Intermidiate
Sorting is mandatory
supports- left join,inner join,right outer join,full outer join
Lookup: lonk names--- master,refernce
sorting is optional
supports-- drop(inner join) continue- left ouer join
All the three are dissimilar from each other in the way they use the memory storage, compare input necessities and how they treat various data . Join and Merge needs minimum memory as compared to the Lookup stage.
7.What are differnt types of Lookups in datasatge?
There are two types of Lookups in DataStage i.e. Normal lookup and Sparse lookup.
Normal lookup--- volume of primary is high and refernce is less
Sparse lookup--- Volume of refernce is high and primary less.
Default is Normal lookup
8.what are differnt types of partioning techniques?
Key based-Hash,Range
Key less--same,auto
9.Components in Datastage?
Designer- Design parllall and sequnce jobs,compile jobs,
Director-- see the logs of jobs, schdule jobs
Adminstrataor--- create project,set permission project
manager- import/export the prject
10 How to import/export jobs from one env to another env?
using export--- job design without executables--.dsx componenets
import----- .dsx format
Q #11) What are the primary usages of Datastage tool?
Datastage is an ETL tool which is primarily used for extracting data from source systems, transforming that data and finally loading it to target systems.
Q #12) What is a DataStage job?
The Datastage job is simply a DataStage code that we create as a developer. It contains different stages linked together to define data and process flow.
Stages are nothing but the functionalities that get implemented.
For example: Let’s assume that I want to do a sum of the sales amount. This can be a ‘group by’ operation that will be performed by one stage.
Now, I want to write the result to a target file. So, this operation will be performed by another stage. Once, I have defined both the stages, I need to define the data flow from my ‘group by’ stage to the target file stage. This data flow is defined by DataStage links.
Once, I have defined both the stages, I need to define the data flow from my ‘group by’ stage to the target file stage. This data flow is defined by DataStage links.
Q #13) What are DataStage sequences?
Datastage sequence connects the DataStage jobs in a logical flow.
Q #14) Where do the Datastage jobs get stored?
The Datastage jobs get stored in the repository. We have various folders in which we can store the Datastage jobs.
No comments:
Post a Comment