The Messy Data Nobody Owns

Fileshares, SharePoint and OneDrive are constant pain points for organisations trying to secure and limit the data stored and exposed in these repositories. We’ve found everything from sensitive client data to domain admin credentials stored in clear text — even in organisations with mature cybersecurity teams, processes and controls. It is simply a beast of a problem to tackle at scale.

I recall one red team where we just simply could not find a way to escalate our privileges within Active Directory. The client had evidently invested significant time and effort into securing their domain. That was until we discovered a backup of a domain controller sitting on a file share that was readable by anyone in the organisation — completely undermining all of that work. This is the nightmare of unstructured data, particularly in large organisations.

Many tools have tackled this problem, beginning with more primitive ones such as SharpShares that simply map file shares and far more advanced ones like Snaffler, that tries to enumerate information stored on these shares as well. These are fantastic tools that really go a long way in helping identify sensitive data stored on file shares, but SharePoint and OneDrive have not received as much attention.

Introducing SharePoint Scavenger

SharePoint Scavenger, as the name suggests, is a scraper for SharePoint (and technically OneDrive – but adding this to the name would create an abomination). The tool is intended for both Pentesters and Blue Teamers to assist in finding credentials and other sensitive data that is unintentionally exposed in online file shares. It uses both SharePoint and Microsoft Graph APIs to iterate through KQL queries to search for potentially sensitive data and can then download these files for further inspection and processing.

How does it work?

tldr; Get an o365 access token, iterate through some KQL search queries and save the output as a CSV. The files that match can then be downloaded for further inspection

The longer version:

  • Two authentication flows are supported:
    • devicecode authentication (as most scripts and tools for o365 do, but Microsoft recently added default conditional access policies to block this flow)
    • Interactive authentication using Selenium (shoutout @dirkjan’s Roadtx)
  •  The resulting access tokens and refresh tokens are saved to disk, which can then be reused for subsequent executions of the tool. If the token is set to expire within 10 minutes, the refresh token is automatically used to obtain new tokens.
  •  The script then iterates through a list of KQL queries saved in queries.txt. These endpoints can be finicky, so if a 500 status code is returned, then the query is re-added to a queue for later iterations.
  • Two output files are generated <date_time>_output.csv and <date_time>_output_unqiue.csv. The unique output removes entries that are identical but excludes the query or hithighlighted summary attributes.
  • The download mode can then be used to save each of these files to disk
    • By default, any file over 50mb will prompt before downloading

Example Queries

The included queries.txt has a number of KQL queries that can be ammended based on your specific needs. The following examples are given as a quickstarter guide:

  • String literals are added simply in quotes. So lets say you want to search for SSH keys, use the following line in queries.txt:
    • “—–BEGIN RSA PRIVATE KEY—–“
  • Searches can be performed on various attributes, for example filename and file type:
    • filename: “SharePointScavenger.bat”
    • filetype:vmdk
  • Wildcards are supported with a “*”, but these must be in parentheses. For example the following will search for various Microsoft Word formats:
    • (filetype:doc*)
  • Boolean operators are supported and definitely help narrow down false positives when searching for strings like “Password”. The following query will search for the string “Password: ” in any .txt or various Microsoft Word formats:
    • ” Password: ” AND (filetype:txt OR filetype:doc*)

Future Plans

Once Microsoft documentation and I have had a healthy break, I would like to consolidate some of the queries to use more of Microsoft’s Graph API, possibly expanding to searching email attachments and Teams chat files.

Once this tool is used on a wider range of environments, I also imagine there will be some compatibility issues and nuances that need to be catered for.

Acknowledgements

@dirkjan’s Roadtx tool is used for the interactive authentication flow:

TeamsEnum was used for the device code auth flow, and pretty output formatting: