Link Partitioning#
Link partitioning controls how data is distributed across parallel processing nodes in DataStage batch flows. The SDK provides methods to configure partitioning on Link objects between stages.
Note
If a stage is configured to run in sequential execution mode (execmode='seq'), the available partitioning options are different. See Sequential Execution Mode.
Setting Partitioning#
To configure partitioning on a link, use the Link.set_partitioning() method. You must specify a part_type parameter, and can optionally provide perform_sort, part_stable, and part_unique parameters. This method returns the Link object for method chaining.
Valid part_type values are: 'auto', 'hash', 'modulus', 'range', 'roundrobin', 'entire', 'same', and 'random'.
Note
Make sure the destination stage has execmode='par' set (some stages use execution_mode instead).
>>> batch_flow = project.create_flow(name='Partitioning Example', flow_type='batch')
>>> source = batch_flow.add_stage('Row Generator', 'Source')
>>> target = batch_flow.add_stage('Peek', 'Target')
>>> target.configuration.execmode = 'par' # not needed if stage defaults to it
>>> link = source.connect_output_to(target)
>>> link.name = 'Link_1'
>>> link.set_partitioning(part_type='hash', perform_sort=True)
Link_1 (src='Source', dest='Target')
The perform_sort parameter is only valid for 'hash', 'range', and 'modulus' partitioning types.
Note
When using modulus partitioning with perform_sort=False, only one partition key is allowed. The SDK automatically sets the key_col_select attribute to the partition key column name. When perform_sort=True, multiple keys are allowed and key_col_select is set to 'default'.
Adding Partition Keys#
After setting the partitioning type, use the Link.add_partition_key() method to specify which columns to use for partitioning. This method requires a key_col parameter and accepts optional parameters for sorting configuration. This method returns the Link object for method chaining.
>>> link.add_partition_key('CUSTOMER_ID', sorting=True, sort_order='asc')
Link_1 (src='Source', dest='Target')
>>> link.add_partition_key('ORDER_DATE', sorting=True, sort_order='desc')
Link_1 (src='Source', dest='Target')
>>> link.add_partition_key('REGION', sorting=False, case_sensitive=False)
Link_1 (src='Source', dest='Target')
You can also chain these method calls together:
>>> link.set_partitioning(part_type='hash', perform_sort=True).add_partition_key('CUSTOMER_ID', sorting=True).add_partition_key('ORDER_DATE', sorting=True)
Link_1 (src='Source', dest='Target')
Removing Partition Keys#
To remove a partition key from a link, use the Link.remove_partition_key() method with the column name. This method also returns the Link object for method chaining.
>>> link.remove_partition_key('REGION')
Link_1 (src='Source', dest='Target')
Inspecting Partitioning Configuration#
You can inspect the partitioning configuration of a link by accessing its part_type and key_cols_part attributes.
>>> link.part_type
'hash'
>>> link.perform_sort
True
>>> len(link.key_cols_part)
2
>>> link.key_cols_part[0]
{'keyCol': 'CUSTOMER_ID', 'partitioning': True, 'sorting': True, 'ci-cs': 'cs', 'asc-desc': 'asc'}
Sequential Execution Mode#
When a stage runs in sequential execution mode (execmode='seq'), it processes all data on a single node — no partitioning is applied. However, you can still control how incoming data arrives by setting part_type='sequential' together with sort keys, giving you ordering guarantees without distributing the data. Adding and removing partition keys works the same way as in Adding Partition Keys and Removing Partition Keys.
>>> batch_flow = project.create_flow(name='Sequential Example', flow_type='batch')
>>> row_gen = batch_flow.add_stage('Row Generator', 'Row_Gen')
>>> remove_dups = batch_flow.add_stage('Remove Duplicates', 'Remove_Dups')
>>> peek = batch_flow.add_stage('Peek', 'Peek')
>>> remove_dups.configuration.execmode = 'seq' # some stages use execution_mode instead
>>> link_1 = row_gen.connect_output_to(remove_dups)
>>> link_1.name = 'Link_1'
>>> link_2 = remove_dups.connect_output_to(peek)
>>> link_1.set_partitioning(part_type='sequential', perform_sort=True, part_stable=False, part_unique=False)
Link_1 (src='Row_Gen', dest='Remove_Dups')
>>> link_1.add_partition_key('CUSTOMER_ID', sort_order='asc')
Link_1 (src='Row_Gen', dest='Remove_Dups')