The Databricks Data Engineer

The Databricks Data Engineer

Jakub Lasak
País Estados Unidos
Géneros Tecnologia
Idioma EN
Episódios 13
Último 14.09.2026

Helping 18k+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.

Episódios

  • How Databricks engineers lose the final round (and the 4 places it happens) 14.09.2026 10min
    You cleared the recruiter screen. You cleared the Databricks technical round, and cleared it well. Then the final round, then the email that says strong candidate, tough decision. You go back looking for the mistake and there isn't one. You were good. Somebody was just easier to picture.That sentence never makes it into the feedback. And it isn't a bad panel or a bias against you. It's a belief you carry into the room: that the final round is the last exam, the hardest one. If your last stage is still a coding or SQL round, that one is an exam, go and study. But most of these emails come out of a different loop, where the technical question was settled and the form closed well before you walked in.In this episode:- What that last round is actually deciding, and why it rarely shows up on a rubric- The four places a strong candidate loses it, ordered from the most common to the one that stays invisible longest- How an unnarrated trade-off reads as taste instead of judgment, and the one sentence that fixes it- The single behavior to take into your next final round, and why trying to fix all four is the wrong moveThis episode is for Databricks data engineers who keep reaching final rounds and keep getting the same email. Whether you're three no-offers into a year of interviewing or prepping for a loop that starts next week, you'll walk away with one concrete thing to change in the room, not another weekend of Spark internals.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every week.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The real reason your Databricks week never got simpler - it isn't the lakehouse 07.09.2026 11min
    Your org finished the migration. Two systems became one. The nightly copy nobody wanted to own is gone, and so is the quarterly meeting where finance and the data science team each showed up with a real revenue number from tables in different buildings. That pain was real, and it is genuinely dead. So why does your week feel heavier than the deck promised?Because most of what got sold as simplification was a relocation. The copy didn't die, it turned into an argument about which schema in which catalog is the real one. The sync job didn't die, it became a maintenance calendar that pages nobody as an alert and pages you months later as the job that got a little slower every week.In this episode:- Why every architecture in this field gets sold on removal, and what the slide never names- How to tell the work that genuinely died in the lakehouse from the work that only changed address- Why a heavier week after a clean migration is not a verdict on you- The two questions to run on the next architecture pitched to your org, and why the second one is about names, not tasksThis episode is for Databricks data engineers who got the win the lakehouse promised and still can't explain why the calendar filled back up. Whether you ran the migration or inherited it, you'll walk away with a lens you can run on any architecture sold as the thing that finally makes data simple.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every week.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The Databricks incident where the most senior engineer never touched the keyboard 31.08.2026 11min
    Thursday, just after two in the afternoon. The first message doesn't come from an alert, it comes from a person in revenue operations: the biggest region jumped overnight and nobody sold anything new. Nobody answers.The engineer who owns the nightly Lakeflow job finds it in the run history. An overnight failure, a restart before anyone was awake, a window of orders written a second time under every executive dashboard in the company.So the fixing starts immediately. Head down, no message, because typing in the channel feels like stealing time from the fix. Which is how a second engineer joins a silent channel, starts the same correction, and lands the same window twice again. Effort was never the scarce thing in that channel.In this episode:- Why every outage is really two problems running in parallel, and only one of them has an owner- What a staff engineer did in ninety seconds without opening a notebook, and why the slower fix was the point- How to state an impact boundary a non engineer can repeat, including the bucket everyone skips- The one sentence that takes the most underpriced role in an incident, at any levelThis episode is for Databricks data engineers who have sat in a chaotic incident channel watching questions pile up faster than anyone can answer them. Whether you own the broken job or you're wondering if it's your place to speak, you'll know which seat in that channel is empty and how to take it.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every week.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • 9 Databricks decisions that reveal your real seniority (and one that only looks senior) 24.08.2026 11min
    Three in the morning. The nightly job that loads the orders table died halfway. Half the tasks are green, and the rerun button is right there. What you do in the next ten seconds says more about your level than anything on your resume.Nobody places you by your title or your years. They place you by a pattern of routine calls made fast, in the dark, and everyone around you can read it except you.In this episode:- The ten-second question to ask before you rerun anything that writes- Why the fastest fix for a slow Databricks job is often the one that renews its cost every month- How the flaky job you quietly restart every week stops being a defect and starts being weather- What it means that most data engineers cannot price their own job, and what changes when they can- The one call that reads as the most senior contribution in the room and almost never isThis episode is for Databricks data engineers who have the title, or want it, and want to know how hiring managers actually place them. Whether you are prepping for a senior interview or wondering why your level has not moved in two review cycles, you will walk away with a lens for your very next decision.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The Databricks job market in 2026: there are two of them, and only one answers 17.08.2026 11min
    One engineer tells you hiring is back and recruiters are warm again. The next one is sixty applications deep and hearing nothing. Same month, same platform, sometimes the same city.That sounds like noise. It isn't. It's a split, and what sorts you onto one side of it has almost nothing to do with how good you are.In this episode:- Why forty applications in a week returns two dead screens, while a smaller and pickier search gets answered- The three things that changed in the last year and turned the generic Databricks data engineer posting into a stack of a hundred before lunch- Why fewer postings can be the better market for you, and the one condition that has to hold for it- The honest case against specializing, plus the test for which second surface is actually safe to bet on- The one resume sentence to rewrite this week, using work you already did and never wrote downThis episode is for Databricks data engineers hearing two contradictory versions of the 2026 market and unsure which one is theirs. Whether you're mid-level and stuck in the application pile, senior and getting less traction than your years should buy, or early career and facing the most crowded entry market in a while, you'll walk away knowing which market your resume is sitting in and what it costs to move.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • Databricks broadcast joins: when the memo beats the meeting (and when it kills a task) 10.08.2026 10min
    Two engineers on the same team join the same big orders table to the same small lookup table. Same cluster, same data, one line of code different. Sarah's finishes in the time it takes to get a coffee. Mike's dies, and the error isn't about the data at all. A task ran out of memory building a hash map.Broadcast joins get passed around as a tip instead of a model: small table equals fast, flip the switch when a join runs long. But the mechanism in almost every write-up is years out of date, and the size Spark checks is not the size that has to fit in memory.In this episode:- How to explain broadcast versus shuffle joins at standup, in one sentence, without a whiteboard- What actually happens on a current Databricks runtime when a join broadcasts, and why the popular explanation stopped being true- Why the size estimate Spark trusts is not the size that lands, and where to read the real payload- What adaptive query execution rescues you from, and the two places its hands are tied- The question to ask before you add a broadcast hint, and the failure signature that tells you a broadcast is what brokeThis episode is for Databricks data engineers who write joins every week and treat the broadcast hint as a speed switch. Whether you're mid-level and tired of guessing why one join flies and an identical one falls over, or senior and about to be asked in an interview why Spark chose a sort merge join, you'll walk away able to predict the call before you run the query.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every week.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake1
  • Build the platform or use the platform: the two Databricks data engineer tracks nobody names 04.08.2026 11min
    Two Databricks engineers sit at adjacent desks. Same title, same pay, same words in the ladder document. Watch a year go by, and they are not doing the same job.Both get rated strong. Both get told they're ready for more scope. Nobody says the useful thing: the evidence those two are stacking isn't interchangeable, and one of those years makes no sense on the other one's promotion packet.In this episode:- Why the data engineering ladder quietly forks, and why that fork is missing from every ladder document- The five-minute question that tells you which track your last year of Databricks work actually built toward- How each track reaches staff, and the very different proof each demands- Where each track stalls a good year, and what to say in your next one on one to get moved off itThis episode is for Databricks data engineers past the junior stage who keep getting strong reviews and vague answers about what's next. Whether you write the cluster policies and ingestion frameworks or ship the tables the business argues from, you'll walk away able to name your track and ask for the work that gets you to the next rung.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • Why Unity Catalog exists: the Databricks governance chaos it was built to end 27.07.2026 10min
    You inherited Unity Catalog already switched on. You learned the catalogs, the schemas, the grant statements, and never once saw the problem all of it was built to solve. Then someone in a review asks why the company suffered through that migration, and the best you've got is one word: governance.Here's what that word hides. Before Unity Catalog, one company ran five separate Databricks workspaces, each walled off - its own tables, its own users, its own rules, and nothing connecting them. A new analyst couldn't find where a table lived without three Slack threads. An auditor's one question - who touched customer data last quarter - cost two weeks of digging through old notebooks.In this episode:- Why every Unity Catalog design choice is scar tissue over a specific pain from the workspace-per-team era- How to explain what the migration actually bought your company in sixty seconds, to a product manager or a staff engineer- The difference between the engineer who says "governance" and the one who gets handed the next platform decision- Why turning Unity Catalog on does not automatically give you what it promises, and what actually doesThis episode is for Databricks data engineers who use Unity Catalog every day but couldn't explain to a product manager why it exists. Whether you joined after the migration and only ever saw the dropdowns, or you're heading into a design conversation where you'll need to justify the work, you'll walk away able to name the world that came before and read every feature as an answer to it.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • How to get credit for Databricks platform work without becoming the on-call martyr 20.07.2026 11min
    Your pipeline hasn't failed in months, and nobody noticed. Then it breaks at 2am, you fix it in twenty minutes, and Slack fills with fire emojis and thank-yous. You're the backbone of the team. You've also been at the same level for three years, and nobody can quite explain why.It's not bad luck, and it's not a skill gap. The exact thing everyone praises you for, being the one who can always fix it, is the reason they can't afford to move you up.In this episode:- Why being the indispensable on-call hero quietly caps the level you can reach- How one engineer systematized a whole class of Databricks incidents out of existence and got promoted for it- The difference between measuring your work in saves and measuring it in risk retired- How to make yourself removable from the hero seat without handing away your job security- The one self-test that tells you whether your recognition depends on things staying brokenThis episode is for Databricks data engineers who carry the pager and keep the platform standing, but keep watching less essential peers get promoted first. Whether you're the war-room legend or the quiet engineer whose prevention work goes uncredited, you'll walk away with three concrete moves to turn indispensability into a promotion case.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • Why Databricks Built Delta Lake When Parquet Was Already Good Enough 13.07.2026 9min
    A nightly job is halfway through writing a batch when the cluster dies. Nobody runs a rollback, because there's nothing to roll back to. The next morning a dashboard is quietly serving half-written garbage, and no one can tell the good files from the wreckage.That's not a bug in Parquet. Every one of those files is valid, beautifully compressed, doing its job perfectly. The problem lives one level up, in the one thing a folder of Parquet files has never had and a real table cannot live without.In this episode:- Why a folder of Parquet files was never actually a table, and what a crashed write silently does to it- The single idea that turns ACID, time travel, and concurrent writes from three separate Delta features into one- How Delta really protects you when two pipelines write at once, and why your code should expect a "no" and retry- A portable question you can point at any storage system to know in one sentence whether you can trust it- The difference between how junior and senior engineers explain why their team runs on DeltaThis episode is for Databricks data engineers who write to Delta tables every day and have never had to explain why it exists. Whether you're prepping for an interview or sitting in a migration meeting, you'll walk away able to explain in sixty seconds what Delta actually bought you, and why the thing that came before it was quietly lying to everyone.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The Databricks pipeline that ran green for a year and was wrong the whole time 06.07.2026 8min
    A revenue-attribution pipeline runs overnight, bronze to silver to gold, and lands a clean set of numbers the whole company uses to decide where the money goes. Every morning: green. Every morning, everyone moves on with their day. For a year, nobody looked any closer.The job succeeded every single night, so everyone trusted the data. But "finished" and "correct" are two different words, and the space between them had been quietly swallowing a slice of revenue for eleven months, with no error, no warning, and not one alert.In this episode:- Why a Databricks job can run green every day and still feed wrong numbers into decisions that matter- The one question that separates "did it run" from "is it right" and cracks silent failures wide open- How an ordinary upstream schema change and an inner join can drop rows for months without tripping anything- The single correctness check that would have caught this on night one instead of month eleven- What to say in the room after an incident like this if you want leadership to trust you with the next hard problemThis episode is for Databricks data engineers who own pipelines whose numbers someone actually makes decisions on. Whether you've ever said "the pipeline's fine, it's probably the dashboard" or you're staring at an all-green run history right now, you'll walk away knowing exactly what to check before you trust it again.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • How Photon actually makes your Databricks queries faster (and when it silently doesn't) 29.06.2026 10min
    Two engineers run the same SQL on the same Delta table. Same data, same cluster size, copy-pasted code. Alex goes to make a coffee and comes back to a query still running. Sam's is done before they finish reading the first Slack message. The only difference is one checkbox on the cluster called Photon.Most Databricks data engineers have that box ticked, pay a premium for it on every DBU, and could not explain in plain English what it actually does. "Photon makes it faster" isn't an answer - it's just saying the box does what the box does. And there's one thing it quietly stops doing while you keep paying for it.In this episode:- The single architectural change underneath Photon, explained without one line of code or config- A mental model that lets you explain Photon to a PM in 60 seconds flat- Why Photon flies on scans, filters, joins, and aggregations but does nothing for certain custom code- The silent fallback that charges you the fast-engine premium for slow-engine work, with nothing on screen to warn you- The one query to open in the query profile this week, and the single question to ask of itThis episode is for Databricks data engineers who run Photon every day and treat it as a speed toggle they never think about. Whether you're a mid-level engineer who clicks the box and moves on, or a senior who's never checked whether the premium is paying off, you'll walk away knowing exactly how to tell when Photon is doing the work you're being billed for - and when it isn't.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The Databricks interview round nobody studies for (and almost everybody fails) 22.06.2026 11min
    Picture the debrief room after a Databricks loop. Two candidates went through that day. On paper, a coin flip: SQL tied, Spark internals solid, system design clean for both. Score only the rounds with a rubric and you cannot separate them. And yet the room isn't split. One gets the offer, and the thing that decided it wasn't any of the rounds they studied for.It was the conversation everyone treats as filler. There's a reason your strongest technical answers can't win it for you, and it's not the one you'd guess.In this episode:- Why the round with no whiteboard and no visible rubric is the one that decides close calls- What it actually measures, and why your best technical round can't test it- The three ways strong engineers fail it without ever noticing- Why "tell me about a decision you'd make differently" is a trap baited with your own best work- The one prep move that needs zero new facts, just a few honest walksThis episode is for Databricks data engineers cleared past the technical bar who keep losing the close calls. Whether you're prepping for your next loop or wondering why a strong interview still ended in a no, you'll walk away with a specific way to rehearse the conversation everyone treats as a throwaway.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The Spark Shuffle is baggage claim: why your job waits instead of computes (and more workers won't fix it) 15.06.2026 11min
    Your Spark job has been running for forty minutes. The dashboard shows your cluster isn't even busy. So you do the obvious thing: add more workers. And it changes nothing.Here's why. During a shuffle, Spark is barely computing at all. It's tagging every row by destination, piling rows together, spilling the overflow to disk, and hauling data across the network between executors. It's an airport rerouting every passenger's bag to a new carousel, and more baggage handlers can't speed up a single overloaded belt.In this episode:- Why your slowest wide transformation spends most of its time on logistics, not computing- The four-step model that lets you explain the shuffle to a teammate in sixty seconds- Why adding workers can make a skewed job slower, not faster- The two numbers in the Spark UI that tell you whether it's skew, partition count, or spill- The one diagnostic to run before you ever resize the cluster againThis episode is for Databricks data engineers whose joins and aggregations crawl for reasons the cluster size never seems to fix. Whether you're mid-level and tired of guessing, or senior and tired of paying for compute that doesn't help, you'll walk away able to read a slow shuffle instead of throwing hardware at it.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • Your Databricks data quality framework is a Yeti: everyone talks about it, nobody has seen it work 08.06.2026 11min
    An architecture review. A platform team is presenting their data quality setup, and honestly, it's impressive. Expectations on every ingestion table. Drift metrics on the dashboard. A dedicated alerts channel. Then a finance engineer asks the only question that counts: when did this last catch something before one of us did? Silence.That silence is the whole problem. The decks, the suites, the dashboards are everywhere. The proof that any of it actually works is somewhere else entirely.In this episode:- Why the discipline every Databricks team talks about is the one with the fewest confirmed wins- A 30-second test that tells you if your data quality framework is alive or just well documented- Why checks chosen by what's easy to write miss the incidents that actually break you- The one deadline-night decision that quietly kills more frameworks than any outage- The reframe that separates teams who get caught off guard from teams who don'tThis episode is for Databricks data engineers who own pipelines feeding the numbers their company argues about. Whether you've got expectations in warn mode you never reverted, or a monitoring dashboard nobody reads, you'll walk away with a way to tell whether your framework is real and what to do if it isn't.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #DataQuality #ApacheSpark #DeltaLake
  • Why senior Databricks engineers write less code than mid-level ones 02.06.2026 11min
    Two engineers, same team, both five years in. Last quarter Mark shipped forty-seven pull requests across three pipelines. Sam shipped nine. On any dashboard, Mark wins by a mile. Sam got the staff offer. Mark got a kind note about continuing to demonstrate impact.This isn't politics, and it isn't luck. It's a pattern that specifically catches the engineers who are best at shipping, because the most valuable work a senior Databricks data engineer does is invisible by construction. You can't put a ticket number on a problem that never happened.In this episode:- Why the exact behavior that makes you great at mid-level is the behavior that keeps you stuck there- The three categories of senior work that produce zero lines of code but move the entire platform- How to tell leveraged work apart from work that just feels safe to ship- Why "less code" is a symptom and not a goal, and the failure mode of engineers who get that backwards- The one question to ask before your hands hit the keyboard that changes what you volunteer for next sprintThis episode is for Databricks data engineers who ship more than anyone on the team and quietly wonder why it isn't landing at review time. Whether you're a mid-level engineer optimizing the wrong line on the chart, or a senior tired of watching your highest-leverage work go uncounted, you'll walk away with language to make prevented work legible and a lens for spending your hours where they actually compound.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • 4 habits that quietly turn your Databricks Delta Lake into a swamp 26.05.2026 11min
    You built the table right. Well-partitioned, documented, fast enough that the row count came back before you finished reading your own Slack. Six months later it takes four minutes to return that same count, and nobody on your team ever decided to make it that way. There was no meeting, no design doc, no ticket titled "let's make this unqueryable by Q3."A swamp is not a decision. It's the sum of a few dozen reasonable shortcuts that compound into something nobody would have signed off on if you'd proposed it all at once. Which is why telling people to "be more careful" never fixes it. They were already careful.In this episode:- Why your slowest Delta table isn't slow because the data is big, and what it's actually choking on- The storage-bill surprise that's invisible in every query until the invoice lands- How the most generous thing you do for a blocked teammate quietly destroys whether anyone can trust the table- Why nobody can clean up a swamp where nobody knows what's load-bearing, and the cheapest fix in the whole estate- When you should ignore all of this advice, because over-governing a throwaway table is just a different swampThis episode is for Databricks data engineers staring at the one table everyone groans about, the one that actually matters, wondering how it got like this. Whether you run batch, streaming, or DLT, you'll walk away able to name exactly which kind of rot is filling your worst table and the specific senior counter-move that reverses it.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • Liquid Clustering vs Z-Ordering: 4 questions that decide 18.05.2026 18min
    You open your Databricks workspace. Two Delta tables. Same size, same downstream BI workload. Table A was partitioned and z-ordered in 2023, runs fine. Table B is greenfield this quarter, liquid clustering by default. Your tech lead asks how aggressive you want to be with migration tickets. Whatever you type back is probably wrong.This is not a feature swap. It's a paradigm shift, and the migration math only makes sense once you can name what actually moved underneath you. Migrate-everything is wrong. Migrate-nothing is wrong. The right answer is per-table, with named criteria.In this episode:- What actually changed when liquid clustering shipped, and the one phrase that simplifies every migration debate you'll have for the next two years- The four-question filter to run table by table, in order, before you commit to a layout decision- The surviving cases where the old paradigm still wins, including the one the evangelism crowd never names- Why liquid clustering and partitioning on a Delta table are mutually exclusive, and the operational property you give up if you migrate the wrong tables- The named audit that turns six hundred legacy tables into three buckets in an afternoon- What kind of senior engineer your tech lead remembers when the promotion conversation happensThis episode is for Databricks data engineers staring at a migration backlog, defending a greenfield default, or trying to explain to a platform team why some tables shouldn't be touched. Whether you're a mid-level engineer running your first migration, or a senior engineer setting the standard for the next two years of greenfield Delta tables, you'll walk away with a defended per-table answer and the vocabulary to back it up.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jrlasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The compounding curve: why some Databricks engineers' salaries grow 5x faster than others 11.05.2026 22min
    Year one. Two new juniors join the same Databricks platform org. Same starting salary, same skills, same desk. Year three, five thousand bucks apart. Year eight, household-car-and-a-half apart. Every year. Forever.Both worked hard. Both stayed technical. Both got positive reviews. Neither did anything wrong. So what happened? Salary in this field isn't one curve. It's two that look identical for the first three years, then peel apart. The choice between them gets made on a handful of small Tuesdays most engineers don't even remember.In this episode:- Why skill is the floor and leverage is the ceiling, and why the better technician is often the worse-paid engineer- The four small Tuesday choices that decide which curve a Databricks data engineer walks up- The difference between expanding what you ship and expanding what you own, and why your manager only fights for one of them- How a junior with twelve hours of writing across four years out-leveraged engineers with twice her tenure- The compass question to run on every career fork before the curve runs youThis episode is for Databricks data engineers who suspect their salary trajectory isn't matching their effort, and who want to know what the highest-paid engineers on their team are doing differently. Whether you're a mid-level wondering why peers at the same level make fifty grand more, or a senior trying to understand why your raises keep shrinking, you'll walk away with a four-part audit you can run on your last six months and your next decision.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jakublasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake
  • The 90/9/1 rule of Databricks performance work - how to triage Spark optimization in 60 seconds 04.05.2026 17min
    Your team is three weeks into a Databricks performance push. Broadcast hints in PRs. AQE flags toggled like christmas lights. Partition counts re-tuned for the third time. The manager is asking, gently, when the gains are showing up in the bill.The staff DE on the next team finished theirs in two afternoons. Same workloads, bigger drop. They were running a triage you have never been taught.In this episode:- Why most of what your team calls Spark optimization is cosmetic and will never move the bill, no matter how clean the PR- The two named tests senior Databricks engineers run on every workload before they touch a config- Why the same change (caching, salted joins, skew handling) can be cosmetic on one workload and structural on the one next to it- Where the real leverage in a Spark workload actually lives, and why it is almost always visible from outside the codeFor Databricks data engineers stuck in a performance push that is not converting effort into runtime or bill drops. Whether you are mid-level drowning in config tweaks, or senior watching the bill refuse to move, you will walk away with a one-minute triage you can run on any Spark workload tomorrow morning.---Helping 18,000+ Databricks data engineers become seniors: interview like seniors, execute like seniors, think like seniors.Follow The Databricks Data Engineer for new episodes every Monday, Wednesday, and Friday.LinkedIn: linkedin.com/in/jakublasakNewsletter: dataengineer.wiki#DataEngineering #Databricks #DataEngineer #CareerGrowth #ApacheSpark #DeltaLake

Popular em

Este podcast também aparece nas paradas de podcasts destes países.