Intervu is in beta — feedback welcome at support@intervu.io

LinkedIn DevOps / SRE Interview Questions

30 real practice questions for the mid-level DevOps / SRE role at LinkedIn (Social/Technology), spanning behavioral, technical, system design, leadership, and problem solving. Build and maintain infrastructure, CI/CD pipelines, and ensure system reliability. The first 3 questions below include what LinkedIn interviewers actually listen for, plus likely follow-ups.

Questions
30
Categories
Behavioral (6), Technical (6), System Design (6), Leadership (6), Problem Solving (6)
Difficulty mix
10 easy · 10 medium · 10 hard
Avg. answer time
~4 min

Behavioral Questions (6)

  1. 1.Tell me about a recent time when you gave or received difficult feedback about your infrastructure work that significantly changed how you approached a problem. Walk me through what happened.

    easy~3 min

    What interviewers look for

    • Shows openness to receiving critical feedback without defensiveness
    • Demonstrates ability to translate feedback into concrete behavioral or technical changes
    • Describes the feedback exchange in a way that shows respect for the feedback giver
    • Shows follow-through and measurement of improvement after receiving feedback
    • If giving feedback, demonstrates LinkedIn's 'direct but kind' communication style

    Likely follow-ups

    • How did you initially react when you received this feedback?
    • What specific changes did you make, and how did you measure whether they were working?
    • Have you since given similar feedback to others on your team?

    Company context

    LinkedIn's Be Open, Honest, and Constructive principle requires direct, kind feedback and transparent communication at every level. In infrastructure roles, this is especially critical given the cross-team dependencies and the need for honest assessment of system reliability and technical debt. LinkedIn expects engineers to both give and receive feedback that improves technical outcomes.

  2. 2.Walk me through a time when you took ownership of an infrastructure problem that was clearly outside your team's direct scope. Why did you step in and how did you drive it to resolution?

    easy~4 min

    What interviewers look for

    • Shows proactive problem identification and willingness to step outside defined responsibilities
    • Demonstrates end-to-end accountability for driving resolution rather than just flagging the issue
    • Describes coordination with other teams while maintaining ownership of the overall outcome
    • Shows consideration of long-term solutions rather than just short-term fixes
    • Illustrates how taking ownership benefited the broader platform or member experience

    Likely follow-ups

    • How did you decide this was worth your time given your existing responsibilities?
    • What was the reaction from the team that officially owned this area?
    • Did this lead to any permanent changes in ownership or processes?

    Company context

    LinkedIn's Act Like an Owner principle expects engineers to think and act like owners, taking end-to-end accountability for systems and outcomes beyond their immediate scope. In infrastructure roles, this often means stepping in when platform reliability is at risk, even when the root cause lies in another team's domain. LinkedIn values engineers who prioritize platform health over organizational boundaries.

  3. 3.Tell me about a time you pushed back on an infrastructure change or feature request because you believed it would negatively impact member experience. What was the situation and how did you handle it?

    medium~4 min

    What interviewers look for

    • Demonstrates clear understanding that infrastructure decisions directly impact member experience at LinkedIn's scale
    • Shows ability to articulate technical risks in terms of member impact (latency, reliability, data privacy)
    • Provides specific metrics or evidence used to support their position
    • Describes how they influenced stakeholders while maintaining collaborative relationships
    • Shows follow-through to ensure member impact was actually measured and validated

    Likely follow-ups

    • How did you quantify the potential member impact - what metrics did you use?
    • What was the reaction from the product or engineering team when you pushed back?
    • Looking back, do you think you made the right call? What would you do differently?

    Company context

    LinkedIn's Members First principle requires that every engineering decision starts with member impact. For DevOps/SRE roles, this means understanding how infrastructure changes affect the 900+ million members who depend on LinkedIn's platform reliability and performance. The company expects SREs to be member advocates, not just technical implementers.

  4. 4.Describe a situation where you raised the engineering bar for your team - maybe around monitoring, deployment practices, or infrastructure standards. What was broken and how did you drive the change?

    medium~5 min
  5. 5.Describe a time when you had to build a working relationship with a difficult product engineering team or stakeholder to deliver a critical infrastructure project. How did you approach it?

    hard~5 min
  6. 6.Tell me about a calculated infrastructure risk you took that didn't work out as expected. What was your reasoning going in, what went wrong, and what did you learn?

    hard~5 min

Technical Questions (6)

  1. 7.LinkedIn's canary deployment process automatically promotes changes from 1% to 100% traffic if metrics look good. What key metrics would you monitor during a canary for a service that handles member profile updates?

    easy~2 min
  2. 8.You notice LinkedIn's Java microservices are experiencing increased garbage collection pauses during peak traffic. The services use LinkedIn's JVM auto-tuning, but GC times are still impacting P99 latency. How do you investigate and resolve this?

    easy~3 min
  3. 9.LinkedIn deploys changes to production within 30 minutes via trunk-based development. A critical microservice starts failing intermittently after your deployment, but only affecting 2% of traffic. Walk me through your incident response approach.

    medium~4 min
  4. 10.Write a Python script that monitors Kubernetes pod restart rates across our microservices and alerts when any service exceeds normal restart patterns. Focus on the alerting logic.

    medium~3 min
  5. 11.You're designing monitoring for LinkedIn's new machine learning inference service that powers job recommendations. It needs to handle 100k QPS with strict P99 latency SLAs. What observability strategy would you implement?

    hard~5 min
  6. 12.LinkedIn's feed ranking service uses Kafka for real-time updates and batch jobs for model training. How would you design the infrastructure to handle both streaming and batch workloads efficiently?

    hard~5 min

System Design Questions (6)

  1. 13.LinkedIn Learning serves millions of video streams globally. Design the content delivery and caching infrastructure to ensure consistent playback quality for learners worldwide, especially in regions with limited bandwidth.

    easy~3 min
  2. 14.LinkedIn's microservices architecture has 3000+ services communicating via Rest.li and gRPC. Design a service mesh observability platform that helps engineers debug cross-service issues and understand the blast radius of service failures in real-time.

    easy~3 min
  3. 15.Sales Navigator processes millions of search queries daily from sales professionals. Design the search infrastructure to handle peak usage while maintaining sub-200ms response times, considering that search patterns are highly skewed toward popular companies and titles.

    medium~4 min
  4. 16.LinkedIn Jobs recommendation engine needs to process member activity data in real-time to update job suggestions. Design the data pipeline that can handle 50k events per second while ensuring recommendations reflect the latest member interactions within 5 minutes.

    medium~5 min
  5. 17.LinkedIn Feed serves billions of personalized posts daily. Design the infrastructure to pre-compute and serve feed rankings while handling the thundering herd problem when a viral post from an influencer with millions of followers gets published.

    hard~5 min
  6. 18.LinkedIn Recruiter handles sensitive candidate data across different regions with varying privacy regulations. Design a data storage and access control system that ensures compliance while allowing recruiters to efficiently search and manage candidate pipelines globally.

    hard~5 min

Leadership Questions (6)

  1. 19.Tell me about a time you had to influence engineers across multiple product teams to adopt a new infrastructure standard or tool. How did you build consensus?

    easy~3 min
  2. 20.LinkedIn's observability team is asking all service owners to adopt their new distributed tracing solution. Several teams on your platform are resistant because it requires code changes. How do you drive adoption?

    easy~3 min
  3. 21.LinkedIn's engineering culture emphasizes fast iteration and continuous deployment. Describe a time when you had to slow down or block a deployment to maintain system reliability. How did you handle the pushback?

    medium~4 min
  4. 22.A junior engineer on your team deployed a configuration change that caused a service outage affecting Sales Navigator users. They're feeling terrible about it. How do you handle the situation with the engineer and with stakeholders?

    medium~4 min
  5. 23.You discover that LinkedIn's Kafka infrastructure is experiencing message lag during peak hours, potentially affecting feed updates for millions of members. Multiple product teams are asking for updates. How do you lead the incident response?

    hard~5 min
  6. 24.You need to migrate LinkedIn's legacy monitoring system to a new platform, but it will require temporary reduced visibility during the transition. Three different VP-level stakeholders have conflicting priorities about timing. How do you build alignment?

    hard~5 min

Problem Solving Questions (6)

  1. 25.LinkedIn's member profiles are updated 200 million times per day globally. Estimate the infrastructure cost impact if we reduced our profile cache TTL from 1 hour to 15 minutes.

    easy~3 min
  2. 26.You notice that LinkedIn Jobs search latency spikes every Monday morning at 9 AM PST but not on other days. Walk me through how you'd investigate this pattern.

    easy~4 min
  3. 27.LinkedIn Learning serves 25 million monthly learners globally. Estimate the storage and bandwidth costs if we increased video quality from 720p to 1080p for all courses.

    medium~5 min
  4. 28.Sales Navigator's prospect search needs to handle 10x traffic during a major industry conference week. Current capacity is 50k searches per hour. Design your scaling approach.

    medium~5 min
  5. 29.LinkedIn's notification system processes 500M daily notifications across email, mobile, and web. Estimate the cost impact if iOS notification delivery rates dropped from 95% to 85% due to Apple policy changes.

    hard~5 min
  6. 30.LinkedIn's Feed ranking model serves 15B feed impressions daily. The ML inference service is experiencing 99.5% availability but we need 99.9%. Estimate the infrastructure investment required.

    hard~5 min

More LinkedIn interview questions