Java Spark Engineer

@ 2T Consulting
2T Consulting2tconsulting.com

Java Spark Engineer

South Bound Brook, New Jersey
Posted today

About the job

The company specializes in data infrastructure and engineering. The role involves architecting and building fault-tolerant data pipelines, leading system design, and mentoring engineers to ensure reliable and efficient big data solutions.

Requirements

  • 7+ years Java development
  • 5+ years Spark experience
  • Strong SQL and storage format skills
  • Experience with Kafka and cloud Spark
  • Knowledge of distributed systems

Qualifications

  • Bachelor’s or Master’s in CS or related
  • Expert in distributed systems
  • Proven system design experience

Full job description

Primary Responsibilities

Architect and build scalable, fault-tolerant data pipelines using Apache Spark (Java)

Lead design of batch and streaming ETL/ELT systems handling large data volumes

Deep-dive performance tuning: partitioning strategy, memory management, shuffle/skew optimization, job cost reduction

Set coding standards and lead code/design reviews across the team

Drive technical decisions on data architecture, storage formats, and pipeline orchestration

Mentor mid-level and junior engineers; act as a technical escalation point

Partner with product, analytics, and platform teams to translate requirements into scalable systems

Own production reliability — on-call ownership, incident response, root-cause analysis for pipeline failures

Evaluate and introduce new tools/frameworks where they improve the system

Contribute to capacity planning and cost optimization for cluster infrastructure

Required Qualifications

Bachelor’s or Master’s degree in Computer Science, Engineering, or related field

7+ years of professional Java development experience

5+ years hands-on experience with Apache Spark in production environments

Expert-level understanding of distributed systems: fault tolerance, data locality, shuffle mechanics, resource management

Proven track record designing systems processing terabyte+ scale data

Strong SQL skills and deep familiarity with columnar storage formats (Parquet, ORC, Avro, Delta Lake/Iceberg)

Experience with cluster managers (YARN, Kubernetes) and cloud-managed Spark

Proficiency with Kafka

Strong grasp of CI/CD, containerization, and infrastructure-as-code practices

Preferred Qualifications

Experience with Flink or other stream-processing frameworks

Familiarity with data governance, lineage, and quality frameworks

Experience with workflow orchestration at scale

Background in system design for multi-tenant or multi-region data platforms

Prior experience leading a team or acting as a technical lead

Show full description