Forward Deployed Reliability Engineer
At a glance
Mid-level Solutions & Sales Engineering role at Palantir. London · full-time.
Pay not stated
Growth Roles summary, based on the employer's posting.
What you'll do
- Respond during assigned on-call weeks to critical outages before they affect customers
- Investigate field issues, restore affected workflows, and prevent recurring reliability problems
- Work with internal stakeholders to scale and strengthen Foundry workflows
- Spot repeated manual work, then automate or simplify it with scripts and process changes
- Turn operational lessons into product improvements, documentation, and shared reliability practices
What you bring
- A background in computer science, engineering, information systems, or another technical discipline
- Working ability in Python, Java, and SQL, plus familiarity with Spark optimisation and parallel data processing
- Experience analysing root causes and recording solutions that others can apply
- Clear communication with both technical and non-technical stakeholders, alongside strong prioritisation skills
Who this fits
This role suits a hands-on reliability engineer who can work independently and with others on unclear technical and operational problems. You should be comfortable with rapid troubleshooting, automation, continuous improvement, and knowledge sharing. The job is hybrid in London, with assigned on-call weeks requiring responses to critical outages, though routine weekend or after-hours work is not required.
From the employer
The Role
As a Forward Deployed Reliability Engineer (FDRE), you ensure the stability and reliability of mission-critical workflows built on Palantir software. You gather signal by going on call — resolving problems before the customer is impacted — and use those learnings to drive product change, shape our internal tooling, and refine our operational processes so that we provide an increasing quality of service to more and more customers.
Your approach is hands-on and pragmatic: you’ll rapidly address issues as they arise with quick and effective solutions, and advocate for workflow or product improvements once the immediate issue is resolved. You are energised by engaging directly with problems, from writing a script to automate a manual task, to finding creative workarounds, or building a case for a product enhancement. You don’t just fix issues — you look for opportunities to simplify, automate, and make the entire system more resilient.
An FDRE synthesises learnings from support into best practices for others to follow. These are captured in documentation and shared with the team and the wider organisation. In this way, you raise the bar for reliability and efficiency across Palantir.
Core Responsibilities
- Develop a deep understanding of Palantir’s products and operational processes
- Go on-call, responding quickly and effectively to mission-critical incidents
- Diagnose, resolve, and proactively prevent issues encountered in the field
- Collaborate with internal stakeholders to increase the scalability and reliability of Foundry workflows for our customers
- Identify recurring pain points and inefficiencies, and take initiative to automate or streamline workflows
- Advocate for and implement product enhancements based on insights gleaned from the field
- Create clear, actionable documentation and share best practices to elevate team and company-wide reliability
Note: While active work is not required on weekends or outside business hours, you must be available to respond to critical outages during assigned on-call weeks.
What We Value
- Ability to work independently and collaboratively to solve ambiguous technical and operational challenges
- Excellent written and verbal communication skills, capable of interacting effectively with both technical and non-technical stakeholders
- Proficiency in Python, Java, and SQL
- Familiarity with parallel data processing and Spark job optimisation
- Strong organisational skills and attention to detail, with the ability to prioritise effectively
- Resourcefulness and creativity in fast-paced, dynamic environments
- Experience with root cause analysis and documenting solutions for broader impact
- Enthusiasm for hands-on problem-solving, continuous improvement, and knowledge sharing
What We Require
- Background in Computer Science, Engineering, Information Systems, or other technical field
Not this one either?
Claude or ChatGPT reads the other 90,537 for you.
Get better matches →