Founding AI Engineer

The role

Founding AI Engineer at CrowdCut. You own the model layer: video understanding, audio understanding (speech, music, ambient sound), highlight detection, multimodal ranking, music-aware editing, and the AI components that turn raw UGC footage into a share-worthy memory. Audio is the binding element of this product, and you are accountable for getting it right.

This is an IC role. No reports. You work in close partnership with our Founding Lead Engineer — they own the rest of the stack, you own the model layer, and you build the product together.

What you’ll do

  • Own the media intelligence pipeline end-to-end. Choose the models, train or fine-tune what needs it, and ship the inference path that runs at scale.

  • Build the multimodal understanding layer that ranks crowd footage by relevance, quality, and narrative value.

  • Build the auto-edit system invoked after analysis — shot selection, sequencing, pacing, music sync, transitions, intros and outros.

  • Push the line on what runs on-device versus in the cloud. Cost and latency matter from day one.

  • Partner closely with the Founding Engineer on infrastructure, the camera and upload path, and the product surface.

What we’re looking for

What you believe matters more than what’s on your résumé. We’re looking for:

  • You are a builder. You want to ship things people use. You’re not a researcher who writes papers without shipping, and you’re not a manager.

  • You want to make a dent in the universe and work on something that matters. CrowdCut, in the end, is a product about how people, families and communities remember themselves. This work has meaning to you.

  • You are AI-native in what you build AND how you build. Working with frontier models regularly is table stakes. So are autonomous coding agents, continuous deployment, and rigorous evals. You have a point of view on proprietary vs. open-weight trade-offs for video and audio.

  • You care about the product. The model layer is in service of an experience real people will use, and you make decisions with that in mind.

  • You are motivated by hard technical problems and the upside of being early.

  • You are passionate about live events: sports, music, life moments that bring people together.

Experience we expect:

  • Track record of shipping generative media, multimodal, or large-scale recommendation systems in production — at a frontier lab, an AI-native startup, or a deep ML team at a big company.

  • Deep familiarity with the modern AI stack and the open-source ecosystem around it.

  • Based in the San Francisco Bay Area or willing to be. This role is in-person.

Bonus points:

  • Experience building photo and video editing and sharing tools — CapCut, Adobe Premiere Mobile, iMovie, VSCO, Splice, Meta Edits, or the equivalent creation teams inside larger consumer companies (TikTok creation, Instagram Reels editing, Snap camera, YouTube Shorts, Apple Memories, Google Photos).

  • Hands-on experience with audio and visual AI — speech understanding, music intelligence, video understanding, computer vision, or generative media. Audio is the binding element of CrowdCut, but the model layer spans both modalities. Tools like Roboflow are already in our workflow; engineers with shipped work on either side (Spotify, Shazam, Runway, or computer-vision-heavy consumer products) have a strong edge.

  • You’ve shaped how content travels at scale on TikTok, Snap, or Instagram — recommendations, ranking, feed dynamics, or the social graph itself. The model layer at CrowdCut sits in the same lineage, and that experience compounds quickly here.

What we offer

  • Founding team equity and market cash compensation.

  • One of the most interesting unsolved problems in consumer AI media, on a real product with a real user wedge.

  • A true working partnership with the Founding Lead Engineer.

  • Direct access to the founders and the advisor network.

How to apply

Email hello@crowdcut.ai. Tell us what you’re working on now and what would make you leave it.