Founding ML Engineer — Live Music Intelligence
Gotobeat – The operating system for live music
Location: Remote-first (UTC to UTC+3 preferred) | Type: Full-time, Permanent | Reports to: Head of Tech
Salary: £80,000 – £130,000 + meaningful equity
How to apply: send your CV to alfredo@gotobeat.com
The question you will answer for a living
Can this artist fill this room, in this city, on this Thursday? Every promoter on earth bets their own money on that question, and almost all of them answer it on gut feel, a Spotify listener count and a phone call to a friend. Get it right and a 300-cap room in Leeds hosts the act that headlines a festival two summers later. Get it wrong and someone loses the deposit on a venue that is half empty on the night.
We built a machine that answers it better than gut feel. It is called Antonello, and it is looking for the person who will make it great.
About Gotobeat
Gotobeat uses AI-driven, event-driven software to make live music fairer and more profitable for everyone involved. We help emerging artists earn more, venues fill more seats, and promoters spend more time building real relationships—not wrestling with admin. Our platform spans planning, promotion, ticketing and post-show settlement, backed by an AWS serverless stack, data pipelines in Python, and agentic LLM tooling.
Why this role matters
Most music data companies see one slice: streams, or social, or ticket sales. Gotobeat is a ticketing platform and a promoter tool, so we hold the one dataset that ties them together—what an artist streamed, what they posted, where they played, and how many people actually paid at the door. That join is the most valuable thing we own, and today it is held together by one founder and a lot of late nights.
You are the first hire whose whole job is that data. You inherit the models, the pipelines and the storage, and you own them end to end: from the notebook where a feature is born to the number a promoter reads before they sign a deal. When your forecast moves, real shows get booked or not. There is no more direct feedback loop in machine learning than an empty room or a sold-out one.
What you inherit on day one
- Antonello, the capacity engine. A per-city venue-capacity forecast: an XGBoost ensemble over an event-history prior, calibrated ranges, confidence gates that refuse to over-claim, and an LLM-driven evidence-repair loop that hunts down the shows our scrapers missed.
- The discovery loop. A daily pipeline that pulls new artists from Songstats, Spotify playlists and related-artist graphs, runs them through fit gates, genre families and liveness checks, and hands a promoter a shortlist that matches their taste and their rooms.
- The data estate. DynamoDB as the system of record, a real-time change-data-capture mirror into a normalised Postgres schema on Supabase, S3 snapshots, and the scrapers and API clients behind Songstats, Bandsintown, Songkick, Spotify and Instagram.
- The analytics. Sell-through and price suggestions, the hype and momentum scores, ad attribution across Meta and Google, and a rule engine that second-guesses the model when the evidence is thin.
What you'll do
- 35% Own the capacity model. Retrain it, re-calibrate it, and score it against what actually happened at the door. Decide when it should stay silent. Ship every improvement to production yourself.
- 25% Own the data pipeline and the storage. Make every source complete, fresh and cheap, and make the join between streaming, social, touring and ticket sales something a colleague can query in one line.
- 15% Turn artist discovery into a ranker that beats a good A&R. Find the act that is about to break in a city before the rest of the industry notices.
- 15% Build the analytics a promoter bets on: expected sell-through, the right ticket price, which ad spend actually sold tickets, and the honest confidence behind each number.
- 10% Expose the models as tools our LLM agents call, with the evaluation and the guardrails that let an agent draft a deal from a forecast without embarrassing anyone.
The stack, honestly
- Python 3.13 in Lambda containers for the models and the scrapers: XGBoost, numpy, Playwright.
- TypeScript on Node for the orchestration, the API and the promoter product: AWS Lambda, Step Functions, EventBridge, SST v3 as infrastructure-as-code.
- DynamoDB with ElectroDB, DynamoDB Streams into Supabase Postgres, S3 for snapshots and artefacts.
- Claude, Gemini and Perplexity for the agentic parts, called through MCP tool surfaces.
- No Kubernetes, no Spark cluster, no GPU farm. Our data is small and precious, not big and cheap: tens of thousands of artists, hundreds of thousands of shows, and the ticket sales behind them. Judgement beats compute here.
- No pager. The crons run overnight, the alarms land in Slack, and if something broke you read about it at nine with a coffee.
You'll be successful here if you have
- 5+ years shipping machine learning that real people relied on, and the scars from at least one model that was confidently wrong in production.
- Real depth in tabular ML and forecasting: gradient boosting, calibration, evaluation on small and noisy data, and the discipline to prefer a simpler model that you can explain to a promoter.
- Data engineering as a craft, not a chore: you have designed a schema, owned a pipeline end to end, and made an expensive source cheap.
- Fluent Python, and enough TypeScript to ship the glue yourself instead of waiting for someone else.
- Practical LLM engineering: tool use, retrieval, evaluation, and a clear view of where a model must ask before it acts.
- A track record of building autonomously. You have taken a problem from a blank repo to something people use, without a committee writing the tickets.
- Genuine love for live music. You have stood in a half-empty room and known why, and you want the next one to be full.
Nice to have
- You work now, or worked recently, on music or entertainment data—Spotify, Songstats, Chartmetric, DICE, Bandsintown, a label, a booking agency or similar. This is a strong signal for us.
- Experience with streaming, social or touring APIs, and with scraping that survives a site redesign.
- Experience with AWS serverless data patterns and change-data-capture pipelines.
- Experience with the MCP tool protocol and agent evaluation.
How we work
- Remote-first culture with quarterly in-person hack weeks in London.
- Small, senior team—no heavy management layers, lots of autonomy.
- One-week sprints, async-friendly comms, and fast decision cycles.
- You get the founder who built the current models as a partner for the first months, then the keys.
- We value clean code, continuous learning, and doing right by artists.
Benefits
- £80,000 – £130,000 base salary, set by experience, plus meaningful founding-team equity.
- Fully remote, with the hours that suit you.
- Infinite paid holiday (+ local public holidays).
- Complimentary gig tickets and backstage passes to all Gotobeat shows.
- A budget for the hardware, the data and the conferences you need.
Hiring process
- Intro chat (30 min) – the role, the music and whether we make sense to each other.
- Technical deep dive (90 min) – you walk us through a model you shipped, then we pair on a real prediction from our data that went wrong. No LeetCode, no take-home.
- Founders' chat (30 min) – vision, values, equity, and any remaining questions.
Total time from first call to offer: ~2 weeks. We move fast and we reply.
Ready to join Gotobeat?
Send your résumé to alfredo@gotobeat.com with the subject line "Founding ML Engineer – [Your Name]". Tell us about a model you shipped and what it got wrong, and name one artist you think is about to break. We review every application and reply within five business days.
Gotobeat is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all team members.
Explore other open positions at: Gotobeat Hiring Page