Job Description : ¿ Develop and maintain data pipelines, ELT processes, and workflow orchestration using Apache Airflow, Python and PySpark to ensure the efficient and reliable delivery of data. ¿ Design and implement custom connectors to facilitate the ingestion of diverse data sources into our platform, including structured and unstructured data from various document formats . ¿ Collaborate closely with cross functional teams to gather requirements, understand data needs, and translate them into technical solutions. ¿ Implement DataOps principles and best practices to ensure robust data operations and efficient data delivery. ¿ Design and implement data CI/CD pipelines to enable automated and efficient data integration, transformation, and deployment processes. ¿ Monitor and troubleshoot data pipelines, proactively identifying and resolving issues related to data ingestion, transformation, and loading. ¿ Conduct data validation and testing to ensure the accuracy, consistency, and compliance of data. ¿ Document data workflows, processes, and technical specifications to facilitate knowledge sharing and ensure data governance. Responsibilities: ¿ Proficiency in using Apache Airflow and Spark for data transformation, data integration, and data management. ¿ Experience implementing workflow orchestration using tools like Apache Airflow or similar platforms. ¿ Demonstrated experience in developing custom connectors for data ingestion from various sources. ¿ Strong understanding of SQL and database concepts, with the ability to write efficient queries and optimize performance. ¿ Strong programming skills, particularly in Python, with experience in web scraping and implementing OCR based data extraction. ¿ Experience implementing DataOps principles and practices, including data CI/CD pipelines. ¿ Understanding of NLP techniques to analyze text data and derive valuable insights for compliance and business intelligence purposes. ¿ Familiarity with NLP techniques and libraries for text data analysis. ¿ Familiarity with data visualization tools Apache SuperSet and dashboard development. ¿ Knowledge of data streaming and real time data processing technologies (e.g., Apache Kafka). ¿ Strong understanding of software development principles and practices, including version control (e.g., Git) and code review processes. Required Skills ¿ Python, Pyspark, Sql,Airflow, Trino, Hive, Snowflake, Agile Scrum Optional Skill ¿ Linux,Openshift, Kubernentes, Superset