Make Spark know the partitioning of the read data

The connector partitions the graph to allow Spark to read it in parallel. But Spark does not know anything about the partitioning. Say the connector partitions the graph by predicate and uid range, Spark would not know that and repartition / shuffle the read data if it wanted to join on partition or uid. If Spark would know the exact partitioning scheme, it could avoid un-needed shuffle steps.

Check to what extend Spark allows data sources to tell it about its partitioning.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Make Spark know the partitioning of the read data #153

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Make Spark know the partitioning of the read data #153

Description

Metadata

Metadata

Assignees

Labels

Type

Projects

Milestone

Relationships

Development

Issue actions