How SCET Works
Understanding the AI technology behind copyright analysis
System Architecture
SCET uses a multi-layered architecture that combines artificial intelligence, legal rule engines, and real-time data collection to provide accurate copyright status predictions. Here's how each component works together:
1. AI Semantic Search
When you search for a work, SCET uses semantic search algorithms to understand your intent. It doesn't just match keywordsโit understands context, variations, and related works. The AI calculates similarity scores using Jaccard algorithms to find the most relevant results.
2. Metadata Wrapper
SCET connects to multiple data sources simultaneously: Open Library (books), Wikipedia (general knowledge), MusicBrainz (music), US Copyright Office, and more. The metadata wrapper normalizes data from different formats into a unified structure for analysis.
3. Legal Rule Engine
The rule engine applies copyright laws based on jurisdiction. It considers publication year, creator death date (if available), work type, and jurisdiction-specific rules. For example, in the US, works published before 1928 are generally in the public domain.
4. Machine Learning Model
The ML model learns from each search to improve accuracy. It identifies patterns in copyright data, predicts missing information, and adjusts confidence scores based on data quality. The more the system is used, the smarter it becomes.
5. Confidence Scoring
Each result includes a confidence score (0-100%) based on data completeness, source reliability, and legal certainty. High confidence (>80%) means strong evidence. Low confidence (<50%) means professional legal review is recommended.
6. Dynamic Updates
Copyright status changes over time as works enter public domain. SCET's dynamic system automatically updates status based on current dates, ensuring you always get the most accurate information available.
Search Process Flow
- Query Processing: Your search query is analyzed and normalized
- Multi-Source Search: Simultaneous queries sent to all integrated data sources
- Text Similarity Matching: Results filtered using 30% similarity threshold
- Metadata Extraction: Title, creator, year, and type data extracted
- Legal Analysis: Copyright rules applied based on jurisdiction
- Confidence Calculation: Score computed based on data quality
- Results Ranking: Results sorted by similarity and confidence
- Display: Top 15 results presented with detailed metadata
Text Similarity Algorithm
SCET uses the Jaccard similarity coefficient to compare your search query with titles in the database. This algorithm measures the overlap between two sets of words:
Jaccard Similarity = (Common Words) / (Total Unique Words)
For example, searching "Harry Potter" against "Harry Potter and the Philosopher's Stone" yields approximately 67% similarity. Only results with โฅ30% similarity are shown to filter out irrelevant matches.
Copyright Status Logic
The system determines copyright status using these rules:
- PUBLIC_DOMAIN: Published before jurisdiction-specific cutoff year (e.g., pre-1928 in US)
- PROTECTED: Published after cutoff year and within copyright term (typically life + 70 years)
- UNKNOWN: Insufficient data to make determination (requires manual verification)
Data Quality & Reliability
SCET prioritizes data quality through:
- Multiple Source Verification: Cross-referencing data across sources
- Source Authority Ranking: Official copyright offices ranked higher than crowd-sourced data
- Transparent Confidence Scores: Users see how confident the system is in each prediction
- Source Attribution: Every result shows exactly where the data came from