AI Labs
Data

How We Contract With AI Labs: NDA, IP, Data

The single longest part of starting an engagement with a frontier AI lab isn't technical scoping - it's the vendor security review. Six weeks is normal. Eight is common. Twelve happens when the vendor discovers gaps mid-process. Here's how we structure NDAs, IP assignment, data handling, and security reviews so the paperwork doesn't outlast the pilot.

Sachin Rathor | CEO At Beyondlabs

Sachin Rathor

24 Aug 2026

7 min read

Modern corporate technology thumbnail showing a professional reviewing an AI lab contract with NDA, IP assignment, data handling, security, and cloud protection visuals in a peach, orange, and purple design.

NDA, IP Assignment, and Data Handling - How We Contract With AI Labs

The single longest part of starting an engagement with a frontier AI lab is not the technical scoping. It's the vendor security review. Six weeks of paperwork is normal. Eight weeks is common. Twelve weeks happens when the vendor has gaps in their contracting or security posture and discovers them in the middle of the process.

We've structured our contracting and security posture to make this review fast. Not by cutting corners - by having the answers ready before the questions get asked. This essay walks through the four pieces of that posture: NDAs and confidentiality, IP assignment, sub-processor and data handling, and how a typical vendor security review actually unfolds. If you want the full packet, you can request our vendor security packet.

This is the operational complement to how we structure pods. The pod is the delivery vehicle; the contract is what makes it legally and operationally safe to operate.

NDA structure - what we sign and what we don't

The default is a mutual NDA, signed before the first substantive call. The "mutual" part matters: a one-way NDA where only the vendor is bound but the lab is free signals an asymmetric relationship that we'd rather not start under. Labs serious about partnership sign mutual.

We sign the lab's standard NDA when available - most frontier labs and tier-2 labs have one and the language is reasonable. When the lab doesn't have a template, we provide our own clean mutual NDA. Most labs return signed within 24 to 72 hours.

A few clauses we pay attention to in any NDA. The definition of confidential information should be broad enough to cover task specs, gold sets, rubric details, and any information about the lab's training pipeline - but not so broad that it covers generic industry knowledge. The survival period is standard at three to five years post-termination for ordinary confidential information, and perpetual for trade secrets; we sign both. Standard carve-outs should exist for information that was publicly known, independently developed, or required to be disclosed by law. And the NDA should permit disclosure to our engineers under their own confidentiality obligations - without this, NDA compliance becomes administratively painful.

What we don't sign without negotiation: NDAs that require us to indemnify for indirect or consequential damages, or that include open-ended non-compete language disguised as confidentiality. Both are dealbreaker patterns we've seen in poorly-drafted templates and we push back.

IP assignment - what gets assigned, what doesn't

Every annotation, trace, evaluation, and other training data deliverable we produce is fully assigned to the client. Full IP, including all reproduction, modification, and distribution rights, with no carve-out for our future use except where explicitly negotiated.

This is the right default for AI training data work. Labs are training models on this data; they need clean ownership. Any vendor that tries to retain "training data rights" or similar should be evaluated carefully because it usually signals either misaligned incentives or inexperience with the category.

What does not get assigned, by default: pre-existing methodologies, processes, and tooling. Our 5-stage vetting process, our QA stack, our gold-set construction methodology, and our schema templates were developed before the engagement and aren't deliverables. We retain them. We also retain anonymized aggregate operational data - internal metrics like throughput rates, calibration cadence patterns, and reviewer-pool composition - in anonymized form for internal benchmarking.

We document the boundary between deliverable IP (assigned) and methodology IP (retained) in every Statement of Work. Surprises later in the engagement are the most common source of contracting friction; we put the boundary in writing up front.

Sub-processor and data handling - the structural piece

For any meaningful contract with an AI lab, the lab will ask: who has access to our data, where does it live, and how do we know it's protected? The answer needs to be specific.

Sub-processor list, disclosed up front. We maintain a documented list of every third-party processor that touches engagement data. Currently: a SOC 2-aligned background check provider used in our vetting process, our cloud storage provider, our identity management provider, our endpoint management provider, and our task tooling provider. Names, jurisdictions, and data categories are listed in the security packet. When a lab requires sub-processor approval rights - which is common - we accommodate. We notify on changes, and we'll switch providers if a specific provider is incompatible with the lab's risk posture. We've done this several times.

Data handling operates in four layers. Endpoint controls: all engineer devices are managed devices with full-disk encryption, MDM, automatic patching, and inactivity-locking - no personal devices touch engagement data. Isolated project environments: each engagement gets its own logically isolated environment with no cross-engagement data co-mingling, and access is provisioned per the principle of least privilege. Encryption in transit and at rest: a standard requirement, but documented and audited. Secrets management via enterprise-grade tooling: no engineer holds long-lived credentials directly.

Data residency, when required. For engagements with explicit data residency requirements (EU residency for certain GDPR scenarios; US-only for some defense-adjacent work), we configure regional environments. This adds one to two weeks to setup; otherwise it's transparent to delivery.

Deletion and audit. Deletion certificates are issued on engagement close or on request. Audit logs are maintained for the engagement lifecycle and made available to the client during their audit windows.

The framing for all of this is SOC 2-aligned. We're not currently SOC 2 certified, but our controls map to the SOC 2 trust services criteria and we can produce a controls summary in 24 hours. We're on a roadmap toward formal SOC 2 Type II certification and ISO/IEC 27001 readiness; the timing depends on engagement volume.

Vendor security review - what to expect and how we make it fast

The vendor security review is a structured assessment most labs run before they sign anything substantive. The pattern is predictable: a security questionnaire (often 100 to 200 questions), reference check, controls walkthrough, and possibly a video call with our security lead.

Here's how we make this fast.

Pre-built artifacts ready on request. Our security packet contains a one-page controls summary, the sub-processor list, sample data handling diagrams, our standard MSA terms, and an FAQ covering the questions we most often see in security questionnaires. We can send it within 24 hours of NDA execution. Labs that have run a few of these recognize the format and can process it quickly.

Reference clients available on request. We're early in the AI training data category, so the reference list is short, but we have it documented and ready. For engagements that haven't yet produced public references, we can offer anonymized references through partner relationships.

Single point of contact. Sachin runs vendor security reviews directly. No handoffs, no second-tier sales engineer who has to escalate every question. The trade-off is that he can only run a few of these in parallel; the benefit is that the lab is talking to the person who can actually answer.

Realistic timelines. A clean security review with us typically closes in two to four weeks from NDA to signed MSA. Labs with their own multi-week internal procurement processes will extend that, but the vendor-side delays are minimal. We avoid the "vendor takes eight weeks to respond to the questionnaire" failure mode by treating it as a delivery commitment.

Contracting structures we sign

We sign four contract structures, each matched to engagement scale.

The default for any expected production engagement is a Master Services Agreement plus Statements of Work. The MSA covers confidentiality, IP, security, indemnity, and termination terms; each SOW covers a specific pod and scope. This is what frontier labs and most tier-2 labs use.

Per-project fixed-fee contracts are used for pilot engagements and one-off evaluation projects: defined scope, fixed price, no monthly billing. This is the standard pattern for first engagements before an MSA is in place.

Per-task pricing with throughput SLAs is used inside MSAs for production pods: monthly invoicing, defined quality threshold, 30-day notice for scope changes.

Hourly engagements are used only for evaluation, red-teaming, and consulting work where deliverable scope is genuinely hard to fix up front. Capped by default to avoid runaway billing.

Different contract structures map onto different pod sizes; the pricing and pod structures page has the operational pairing. The wrong contract structure for the pod size is the most common cause of friction we see in vendor relationships, both with us and with peers in the category.

What we won't sign

A short list of contractual terms we negotiate or decline: unlimited indemnity, especially for consequential damages; open-ended non-compete language that would prevent us from working with adjacent labs; indemnity for content of training data (we can't take legal responsibility for whether a lab's task spec is itself lawful); auto-renewal without notice on multi-year contracts; and audit rights that would expose other clients' confidential information.

Most of these come up only in poorly drafted templates and get fixed in negotiation. None are dealbreakers in their reasonable forms.

How to start a conversation

If you're scoping an AI training data engagement and want the security and contracting picture before you commit time to a discovery call, you can request our vendor security packet. NDA-gated, sent within 24 hours, containing the controls summary, sub-processor list, sample MSA terms, and an FAQ that answers most security questionnaire items in advance.

For everything operational on the delivery side - pod structures, throughput, calibration, QA - start with the companion essays in this content series. The contracting side is half of why engagements work; the operational side is the other half. Both have to be in place.

Summarize with

1052 Antone Way Petaluma, CA 94952

Summarize with

Disclaimer:

Beyond Labs LLC provides the information on this website for general informational purposes only and nothing herein constitutes professional, legal, financial, investment, or contractual advice, nor does it create a client relationship; all services are governed exclusively by executed written agreements. While we strive for accuracy, we make no representations or warranties, express or implied, regarding the completeness, reliability, or results of any content, case studies, or materials presented, and past performance does not guarantee future outcomes. References to third-party brands, platforms, or technologies are for descriptive purposes only and do not imply partnership, endorsement, or affiliation unless expressly stated in writing. Beyond Labs operates as an independent consultancy and disclaims liability to the fullest extent permitted by law for any reliance placed on website content. We reserve the right to modify this Disclaimer at any time, and continued use of this website constitutes acceptance of the updated terms.

Beyond Labs is a registered trademark of Beyond Labs, LLC. All third-party names, logos, and brands mentioned on this site are the trademarks of their respective owners. Beyond Labs, LLC is an independent entity with no endorsement, sponsorship, or affiliation with these third parties. Any use of third-party names, logos, or brands is solely for identification purposes and does not imply endorsement or partnership.

© Beyond Labs, LLC 2026. All rights reserved.

Based in the USA, Supporting Teams Globally.